← Part of the full guide: Why AI Doesn’t Recommend Your Brand
Why the question set matters more than the tool
Every AI visibility number is downstream of the questions used to produce it. Pick questions your brand happens to be strong on and the number flatters you. Pick questions nobody actually asks and the number is meaningless. The question set is the single largest determinant of measurement quality, and the part almost nobody discloses.
We publish the full structure for the same reason we publish the AIVS rubric: a number whose method cannot be inspected cannot be trusted. If someone measures with this and gets a different result from us, explaining the difference is our job.
The eight tiers
| Tier | Count | What it tests |
|---|---|---|
| 1. Direct vendor questions | 18 | The only tier that produces an ordered list of named companies — this decides whether you place in the top five |
| 2. Vendor plus qualifier | 11 | Filters by budget, company size or pricing model, showing which sub-segments you place in |
| 3. Symptom to provider | 10 | Real buyers arrive with a symptom, not the name of a service category |
| 4. Learning questions | 18 | Whether your site is used as a source, which is different from being named as a provider |
| 5. Comparison and decision | 11 | Late funnel — answers usually give criteria plus example names, a second route into a shortlist |
| 6. Service-specific | 8 | Narrow questions matching individual services — lower volume, far thinner competition |
| 7. Vertical | 8 | Built from the industries your actual clients are in, where competition is much lighter |
| 8. Brand-name questions | 9 | Entity strength — does the system know you, and how does it assess your credibility |
Tiers 1 and 8 must always be read together, because they report different things. On ourselves, tier 8 passed comfortably — both systems described our services correctly — while tier 1 returned zero every time. That combination says the problem is shortlist inclusion, not recognition. The full record is in our self-test.
Why you need both hiring and learning questions
These two groups return different outputs and are fixed by different work.
Hiring questions — asking who provides a service — produce a list of company names. Getting into that list requires an entity that can be corroborated externally: a correctly categorised business listing, reviews, and mentions in sources the system trusts.
Learning questions — asking how to do something — produce an explanation with citations. Getting cited requires content that answers the question cleanly and contains something the other candidates do not have.
A business measuring only the first group never sees that its content is already being used. A business measuring only the second concludes it is doing well while never being put in front of anyone about to spend money.
What to record
Per question, per platform, log five values.
- Named or not — the brand appears in the body of the answer. Appearing only in a source list does not count here.
- Position — if the answer is ordered, record the rank. If it is not ordered, leave this blank.
- Cited as a source — tracked separately from the first value, because being used as a source and being recommended as a provider are different states.
- All sources cited — capture every domain the system used. This tells you where you need to be present to enter the retrieval pool at all.
- Competitors named, in order — this shows who currently holds the seat you want.
The figure worth acting on is the share of runs in which your brand is named, not the outcome of any single run. These systems vary with session state and inferred location, so reading a conclusion from one attempt is a reliable way to be wrong.
Mistakes that invalidate a question set
Writing questions in industry vocabulary
The words practitioners use for a service are usually not the words buyers type. Write the questions in the language customers use when describing the problem, not the name of the service category.
Changing wording between rounds
Editing even one word mid-programme breaks comparability between months. Freeze the set. If you want to add questions, add them as a separate set and start a fresh baseline for it.
Running the set inside one session
Earlier turns influence later answers. Running the whole set in a single chat biases the tail of the set toward the first responses. Open a fresh session for every question.
Leaving browser extensions enabled
Some extensions inject search results into your prompt without displaying it, so you end up measuring the extension rather than the system. We hit exactly this during our own test and had to disable extensions before collecting anything valid.
Using too few questions
Five to ten questions cannot separate normal variance from a real trend. We use at least 25 per business for AI Citation Presence scoring, and the full 93 when a more granular picture is needed.
Build your own set
- Start with 15 to 20 tier-one questions. Write the sentences a buyer types when looking for a provider, in every language your customers use — the lists returned differ by language.
- Add symptom questions. Take the wording from what clients say in their first message to you. That is the language people use before they know the category name.
- Add vertical questions from your existing clients. Competition is much lighter here, and it is usually where a smaller brand first gets named.
- Finish with brand-name questions. These separate “the system does not know us” from “it knows us but does not put us forward” — different problems with different fixes.
- Freeze the set and schedule re-measurement. Monthly is about right. More frequent than that and you mostly observe variance.
CiteLogics — measured with a question set we will show you
The AI Visibility Audit uses at least 25 questions per business across ChatGPT, Claude, Perplexity and Gemini, records all five values per question, and scores the five AIVS dimensions. Clients receive the question set so they can re-measure independently.
- AI Visibility Audit — full AIVS Score report plus a 30-minute call. THB 2,000, delivered in 24 hours, 7-day money back.
- Big Win (Audit + WebFix) — audit plus full implementation across crawlability, structure and schema. THB 20,000.
Get your free AI visibility scan, or contact sales@citelogics.com.
Frequently asked questions
Why publish the question set where competitors can see it?
The question set is not a competitive secret — anyone can write a comparable one in half a day. The value is in measuring consistently and interpreting results correctly. Publishing it lets clients audit our work, which matters more than withholding it.
How many questions are enough?
At least 25 per business for routine scoring. Below that you cannot separate system variance from a genuine trend. Our full set runs to 93 questions for cases needing a tier-by-tier picture.
How often should I measure?
Monthly is about right. More often and you mostly observe system variance rather than the effect of your work. Less often and you learn too late that a competitor has moved. Consistency of wording matters more than frequency.
Do Thai and English questions return different results?
Yes. In our testing, English and Thai questions covering the same topic returned different company lists. If your customers use both languages, you have to measure in both.
Results vary every run — does that mean the measurement is broken?
No, variance is normal for these systems. The way to handle it is to run each question several times and use the share of runs naming your brand as the headline figure, rather than reading a conclusion from a single attempt.

An engineer and AI enthusiast who reverse-engineers how AI models decide which brands to cite. He proved the method on his own websites first — then delivered the same results for brands like Coldtubb. He and the CiteLogics team have scanned 1,000+ websites for AI visibility.
Connect on LinkedIn →