Building a Query Taxonomy for an AI Brand Perception Audit
A repeatable framework for measuring how AI models describe and recommend your brand.

A query taxonomy is the classification system behind an AI brand perception audit: it decides which questions get asked, how often, and what the answers are supposed to prove. Skip it and an audit is just a pile of prompts somebody typed into ChatGPT one Tuesday afternoon. Build it properly and the same audit becomes repeatable, something you can run again next quarter, hold up against a competitor, and use to justify whatever fix comes next.
Buyer research has already moved past the point where this is optional. In the most recent purchase cycle software buyers went through, 54% used at least one generative AI tool, mostly during early research and shortlisting. That's the moment a taxonomy earns its keep. Buyers at that stage aren't typing keywords into a search box; they're asking a model a question and taking the answer at close to face value.
The traffic numbers back this up, though not quite the way people assume. Previsible's 2025 AI Traffic Report clocked AI-referred sessions growing 527% year over year in the first five months of 2025, which sounds enormous until you set it next to Google. All LLM referral traffic combined still sits at roughly 2 to 3% of what organic search alone sends a brand. I'd be careful reading too much into that growth curve on its own; the volume, right now, isn't the story. What matters is what the pattern already tells us about how buyers form opinions before a brand shows up anywhere in their search history.
The real shift shows up the moment an AI decides to name a brand at all. These systems run something like a confidence check before they'll surface a company by name, and they tend to hold back unless they can back the pick with more than one independent source. When ChatGPT or Google's AI Overview names a company, usually one to three per answer, the model is putting its own credibility on the line for that pick. That's closer to an endorsement than a listing. Getting left out is a trust problem as much as a ranking one, and the fix starts with finding out what the model currently believes about the brand. That's the whole point of building a taxonomy in the first place.
The gap between old SEO habits and new AI visibility shows up plainly in the data, if you're willing to look at it straight. The Fuel Online AI SEO report found 62% of enterprise brands invisible to generative AI models, despite 94% of them pouring money into traditional SEO. Fewer than 10% of the sources cited by ChatGPT, Gemini, and Copilot even rank in Google's organic top 10 for the matching query. Rank and citation aren't answering the same question anymore. They never really were. Now the gap is just measurable.
What an AI brand perception audit actually measures — and what it does not
An AI perception audit checks how platforms like ChatGPT, Perplexity, and Gemini describe a brand, how they stack it against competitors, and how confidently they'd recommend it. Counting mentions is a small part of the job, not the whole of it. The real work happens in four places: category clarity (does the model put you in the right competitive set), differentiation (can it say what actually makes you different), proof footprint (does it have sources to back up what it's claiming), and citation readiness (can it even find those sources). Each needs its own fix. An audit that just tallies "brand mentioned: yes/no" tells you almost nothing you can act on, and teams that anchor on it tend to find themselves without actionable direction when it matters most.
It's a narrower thing than a ranking audit, or social listening, or a snapshot you take once and file in a drawer. Ranking answers yesterday's question. The live risk is that an AI describes a brand wrong, unfavorably compares it to a rival, recommends someone else instead, or quietly skips the one thing that actually sets the brand apart. And because AI answers get recycled across adjacent queries, one bad description doesn't stay put. It spreads to every question that touches the same topic.
Cross-platform variance makes broad coverage a requirement, not a luxury. ChatGPT, Perplexity, and Gemini pull from different data and retrieve it differently, so the same brand gets described differently depending on who you ask. Wellows analyzed 11.1 million citations across 571,729 AI answers, spanning 363 brands and 35 regions, and found ChatGPT names brands explicitly about 65% more often than Perplexity does (2.4% of citations versus 1.5%). A taxonomy has to build that gap in from the start. Bolting a second platform on afterward isn't the same thing, and it never produces the same picture, no matter how thorough the retrofit.
None of this requires the brand to have done anything wrong, either. A new article, a competitor's product launch, a cluster of bad reviews: any of these can reshape how an AI describes a company with zero involvement from that company's own content. Which is the whole argument for treating this as something you run on a schedule, not a report you file once and forget about.
Why an informal prompt list fails as an audit methodology
Someone runs a dozen prompts, skims the answers, calls it an audit. It happens constantly, and it's a mistake almost every time. The prompts reflect whoever wrote them: their guesses about what buyers ask, which claims matter, which competitors even count. That's one person's intuition dressed up as data, and it can't be compared across time, across platforms, or against any benchmark worth the name.
Cadence makes it worse. Wellows's citation analysis found roughly two-thirds of cited sources churn between observations. A one-time audit goes stale almost the moment it's finished. Run it again next quarter without a fixed, classified prompt set, and you're asking different questions while calling the result "change." Nobody, at that point, can tell if the brand's standing actually moved or if someone just phrased things differently this time.
Selection bias hides in here too, and it stays invisible until something forces it into daylight. A team leaning on branded prompts misses the solution-level and comparison queries where the real buying decision gets shaped. A team fixated on one engine misses how much the picture shifts on a second one. Neither gap surfaces until a taxonomy makes you account for it, cell by cell.
What a taxonomy buys, ultimately, is comparability: same questions, same classification, same cadence, so the result reads as a trend line instead of a screenshot. It also forces coverage decisions up front, so gaps show up before the audit runs rather than after somebody asks why a category got skipped. And when a result looks wrong, the classification tells you where to look, whether it's a generic-description problem, a comparison-query problem, or a high-intent recommendation problem. A flat list of prompts can't tell you any of that.
The three classification axes every query taxonomy needs
Intent comes first: what is the person actually trying to do when they type this question. Awareness-stage questions ("what kind of software handles X") test whether the brand exists at all in the model's picture of the category. Evaluation-stage questions ("which tools do X best") test how accurate and complete the description is. Decision-stage questions ("what should I use for X, and why") test recommendation frequency and sentiment, the highest-stakes output of the three. A brand that scores fine at awareness but falls apart at decision stage has a specific, nameable problem, and the fix looks nothing like the fix for the reverse case.
Framing comes second: how the question actually gets asked. Branded prompts put the company name right in the question and test what the model already "knows" and how it feels about that. Solution or category prompts describe a need without naming anyone, testing whether the brand gets retrieved at all from a cold start. Comparison prompts name competitors directly and test positioning; this is probably the highest-leverage query type of the three, since it's where a lot of real decisions actually get made. A workable monitoring set runs around 10 to 15 "golden prompts," split roughly 20% branded, 60% solution, 20% comparison. That split isn't arbitrary. It mirrors how buyers actually search, and most discovery starts without a brand name in mind at all.
Evaluative dimension is the third axis: what the query is actually built to surface. Category placement. Attribute accuracy. Differentiation clarity. Proof and citation readiness. Sentiment and recommendation. Any prompt can hit any of these regardless of intent or framing, which is exactly why the three axes have to work together rather than stand alone. A branded, evaluation-stage prompt testing attribute accuracy tells you something completely different from a comparison, decision-stage prompt testing recommendation sentiment, even if both happen to mention the brand by name in the output. Treating those two results as the same kind of signal is where informal audits fall apart, every time.
How to populate each category with prompts that produce comparable signal
Start with how buyers actually talk, not how the brand talks about itself. The queries that belong in a taxonomy are the ones a real researcher would type into ChatGPT or Perplexity, not a rewritten homepage headline. "Prompt volume" is worth borrowing here, the AI-era stand-in for keyword search volume: how often people actually ask a model about a given topic. Sales call transcripts, support tickets, review text, and competitor comparison pages beat any internal messaging deck as raw material for this.
Each prompt should isolate one variable. Ask "what's the best tool for X" and you've mixed category placement, accuracy, and sentiment into a single blob, so a weak answer doesn't tell you which one actually failed. Write narrower instead. "What are the strongest arguments for using [Brand] over [Competitor] for [use case]" tests differentiation at the decision stage. "What do users commonly say are the limitations of [Brand]" tests sentiment at the evaluation stage. Narrow prompts produce narrow, usable answers, and that trade-off is worth making every time.
Wording has to stay fixed across runs. Paraphrase a prompt from one quarter to the next and you've introduced a variable that looks like a shift in brand perception but is really just a shift in phrasing. When a prompt genuinely needs updating, log the version and the date. Otherwise later comparisons stop being honest.
Multiple engines add structural coverage, not just extra data points. Because ChatGPT, Perplexity, and Gemini pull from different sources and behave differently as a result, the same classified prompt has to run on all of them. The variance between engines is itself useful information: a brand that shows up accurately on ChatGPT but disappears on Perplexity has a source-availability problem specific to how Perplexity retrieves and weighs citations.
Not every cell deserves equal attention. Decision-stage comparison queries carry more weight in an actual purchase than an awareness-stage category question does, so they earn denser coverage and a faster re-check cycle. Flag those cells as priority up front, and effort stays pointed at where the risk actually sits.
Scoring what the taxonomy surfaces — turning query outputs into structured measurement
An AI's answer is just text until somebody scores it. That's where the taxonomy earns its keep twice over, since its classification structure is what tells you what to score in the first place. Four dimensions map cleanly onto the axes above. Accuracy: are the facts right. Depth: how complete is the picture. Sentiment: favorable, neutral, or negative. Recommendation frequency: does the brand actually get surfaced when it's relevant.
A handful of newer, AI-specific metrics sharpen this further. Citation Sentiment Score looks at the tone of the sources the model draws from. Source Trust Differential checks whether those sources carry real authority or are thin filler. Narrative Consistency Index tracks whether the brand gets described the same way across engines and query types, or whether the story shifts depending on which model you ask. Entity Co-Occurrence Map tracks which competitors and category terms keep showing up next to the brand, often a faster way to spot a positioning problem than reading transcripts one by one. None of this is standardized yet, and every vendor scores it a little differently. Still, what they're measuring underneath is stable enough to build an internal rubric on, even without an industry standard to lean on.
Manually logging and scoring hundreds of prompt outputs across three or four engines every quarter is real work. It's exactly the kind of work purpose-built perception platforms, Scale Labs among them, exist to take off someone's plate, scoring across algorithmic, AI, and audience signals at once rather than one spreadsheet row at a time. The platform runs the taxonomy at scale; it doesn't replace the need to build one. The structure, coverage by intent, framing, and dimension, scored consistently on a fixed cadence, is what makes the output worth trusting.
That scoring turns a pile of transcripts into something a team can act on: a trend line showing whether a fix actually worked, a side-by-side against competitors on the same solution and comparison prompts, a plain answer to which cells are weak and high-stakes at the same time. No more guessing at the next fix.
What low scores in each taxonomy category actually point to — and where fixes begin
A weak score only matters if it points somewhere specific. That's the whole test of whether a taxonomy is doing its job or just generating noise.
A low score on a branded, evaluation-stage prompt, the kind testing attribute accuracy, usually means the public evidence about the brand is thin, stale, or contradicted somewhere the model can see. That's a source problem before it's anything else, and no amount of brand messaging fixes it if the underlying evidence isn't there.
A low score on a solution-stage prompt testing category placement means the model doesn't have enough consistent, corroborated signal to slot the brand into the right competitive set without being handed the name first. Call it an entity definition problem: the brand hasn't been described the same way, by enough independent sources, often enough for the model to trust the association on its own.
Decision-stage comparison prompts testing recommendation sentiment reveal a different animal when they score low. The model can describe the brand just fine, but it can't explain why that brand beats the alternative for a specific job. That's a differentiation gap, and no amount of source volume closes it if the underlying claim was never sharp to begin with.
One thread runs under all three. These systems favor brands that show up the same way everywhere, that get backed by more than one independent, credible source, and that hold real authority on the specific topic somebody's asking about. A brand described five different ways across five different sources doesn't get the benefit of the doubt from a model built to check for agreement before it commits to a recommendation. It never has, and there's no reason to expect that to change.


