Query Design for AI Perception Audits
How you write audit questions determines what an AI model reveals—or hides—about your brand.

Query design decides the outcome of an AI perception audit before anyone runs a single prompt. Whoever writes the questions decides which signals surface, which blind spots stay buried, and whether the findings can be acted on. Get the questions wrong and you can produce a report that reads beautifully and tells you almost nothing true.
Most people still run this like a brand survey. Open a chat window, type something loose like "tell me about Brand X," copy down whatever ChatGPT says, and call it a finding. I've watched teams do this and then present the output in a board deck. It fails for a structural reason, not a laziness reason: LLMs are probabilistic, not lookup tables. Ask the same question two slightly different ways and you can get two different answers, two different confidence levels, sometimes two entirely different sets of "facts." A badly built query doesn't come back looking incomplete, which would at least be a warning sign. Instead it comes back looking finished. A fluent, confident, well-organized answer to a bad question is more dangerous than a thin one, because nobody thinks to double-check it.
How AI systems have moved from presenting options to rendering verdicts
Search, the old kind, was built around choice. You type something into Google, you get ten blue links, and the engine steps out of the way. It doesn't tell you which one is right. It matches; it doesn't endorse.
Generative search doesn't work that way. Ask ChatGPT or Perplexity or Google's AI Overviews for a recommendation and you get one name, maybe two, rarely more. The model isn't handing you candidates anymore. It's handing you a verdict, and staking its own credibility on it. That changes what exclusion means. Missing from page one of ten results is a ranking problem. Missing from an answer that only names one company is closer to a rejection, and nobody explains the rejection or offers an appeal.
The scale of the shift isn't subtle, either. Gartner projects traditional search engine traffic falls 25% by 2027. Adobe's Digital Economy Index clocked AI-referred traffic up 1,200% between mid-2025 and early 2026. A Fuel Online review of roughly 1,000 enterprise brands found 62% effectively invisible to generative AI models, even though 94% of those same companies had spent real budget on traditional SEO. Optimizing for the channel that's shrinking while ignoring the one that's growing isn't a small miscalibration. It's planning for a landscape that's already gone.
That's exactly why the audit questions matter so much. Build them the way an SEO team thinks about search behavior, and you'll measure a kind of visibility that stopped predicting outcomes a while ago.
What LLMs actually evaluate when they assess a business entity
Here's where a lot of auditors trip: LLMs don't rank pages, they evaluate entities. A page can win over a human reader with clean copy, a strong headline, decent design. None of that buys trust from a model.
What actually earns weight is different, and less visual. Factual accuracy, corroborated across multiple independent sources, matters a great deal. So does entity consistency: your name, your leadership, your core messaging lining up the same way everywhere the model might have run into them. Expert authorship carries real weight too; sources with strong E-E-A-T signals (experience, expertise, authority, trust) account for 96% of AI Overview citations, which tells you plainly that the model cares who's talking. Citation patterns and referring-domain authority factor in. So does semantic completeness, a term for content that actually covers a topic rather than skimming it: in an analysis of nearly 16,000 AI Overview results, content scoring high on completeness was 4.2 times more likely to get cited.
Mention frequency gets less attention than it deserves. Research keeps landing on it as one of the strongest predictors of whether a brand shows up in AI-generated answers at all, right beside review sentiment and how fresh the underlying data is.
None of that is testable with one prompt. "Tell me about Brand X" flattens all these separate signals into a single paragraph you can't pull back apart afterward. If you want to know whether the model has your leadership right, whether it puts you in the correct category, whether other sources back up what you claim about yourself, you have to ask three separate questions. The entity is the unit of analysis, not the page. An audit checking only whether your URL comes up is answering a question nobody's asking anymore.
The hallucination problem and why it makes query sequencing critical
Hallucination here means something narrow and specific: an answer that sounds fluent and sure of itself but is wrong, invented, or misleading in ways you can't detect from tone.
This is the part that still catches people off guard. An MIT study from January 2025 found models are 34% more likely to reach for high-confidence language exactly when they're generating something false. The model doesn't hedge more as it gets less accurate; it gets more certain. Stanford's research adds a second wrinkle specific to naming: companies with common or ambiguous names face a 41% hallucination rate versus 23% for companies with distinctive names. Share your name with a common word, a person, or another company, and you're not just harder to find. You're measurably more likely to get misrepresented by the exact system replacing search for a growing share of your buyers.
Fabrication clusters hardest around long-tail knowledge, which is where most businesses actually live. The giants get corrected by sheer volume of training data. Everyone else is exposed.
That makes this a query design failure before it's a model failure. An open prompt, "what do you know about Brand X," hands the model room to fill gaps with inference dressed up as fact. A structured query, isolating one checkable claim, a founding year, a current CEO's name, forces commitment. Ask the same factual question three ways across three sessions and you find out whether confidence tracks accuracy or whether the model is just performing certainty regardless of whether it's earned. A single-prompt audit can't tell a model that genuinely knows your brand from one confidently making things up about it. Sequencing is the only thing that separates the two.
The four dimensions a well-designed query set must cover
A rigorous query set has to keep four things separate. Collapse them and you get an audit that's fun to read and useless to act on.
Entity recognition comes first, and it's the most basic question there is: does the model even know your brand is a distinct thing? Test it with unprompted recall, asking the model to name companies in your space without naming you, and with disambiguation checks if your name overlaps with something else in the world. What you're checking is whether the model holds a stable, accurate concept of you, or whether it's guessing and covering.
Factual accuracy comes next, and this is where hallucination testing actually lives. Isolated questions on founding date, leadership, product line, geography, run across multiple sessions and multiple temperature settings, give you a hallucination rate and, more usefully, tell you which specific claims are the fragile ones.
Evaluative positioning maps most directly onto the verdict-rendering shift I mentioned earlier. Decision-context prompts phrased the way a real buyer talks ("what's the best option for someone who needs X"), category prompts ("who leads in Y"), direct comparisons: all three reveal whether you land in the answer, what sentiment surrounds you when you do, and which attributes the model has quietly decided belong to you.
Narrative consistency closes it out. Does the model describe you the same way from a buyer's angle as a journalist's angle as an analyst's angle? The same way on ChatGPT as on Claude, Perplexity, Google's AI Overviews? High variance is itself the finding, not a footnote to it. It means your entity isn't well grounded in whatever corpus each model happened to draw from, and that instability is a risk independent of whether any one answer looks flattering.
How query framing shapes the evaluative signals a model surfaces
Wording doesn't just retrieve an answer. It decides which register the model answers in.
Frame a question as a first-time buyer and you pull sentiment. Frame the same underlying question as an industry analyst and you pull credibility signals instead. That's the persona effect, and it means an audit run entirely from one implied vantage point only ever sees half the picture. Decision-context framing works a bit differently: phrase the query around a real scenario, "I'm evaluating vendors for an enterprise deployment," and the model applies a selection filter, showing you which attributes it treats as decisive rather than incidental.
Comparison prompts force commitment in a way open questions don't. "How does Brand X compare to Brand Y" doesn't let the model hedge; it has to stake out relative positioning, and often that reveals a narrative gap a competitor has filled that you haven't touched. The negation probe runs against most brand instincts, but it might be the single most valuable question in the set: "what are the main criticisms of Brand X" surfaces negative associations a positively framed question will never turn up, because you never asked for them. The temporal probe, "what's changed about Brand X in the past year," tells you whether the model's knowledge is current or frozen on some older snapshot, which matters a lot if you've rebranded or changed leadership recently.
Watch for leading language. "Brand X is known for innovation, tell me more" doesn't evaluate anything, it just asks the model to agree with you. Every question in the set should read as though the person asking genuinely doesn't know the answer, because the model doesn't know what you were hoping to hear, and shouldn't be nudged toward guessing.
Controlling for model variability across sessions, temperatures, and platforms
Repetition isn't a nice extra step here. It's structural, because the same query run twice on a probabilistic system will not reliably give you the same answer twice.
Temperature makes this worse. Higher temperature settings produce more varied, occasionally more inventive, generally less dependable output, so audit queries need to run at documented, fixed settings, or comparisons across sessions stop meaning anything. Cross-platform testing adds a layer worth treating as a signal, not a formality. ChatGPT, Perplexity, Claude, and Google's AI Overviews draw on different training data and different retrieval setups. A brand that shows up consistently across all four has a kind of entity grounding a brand appearing on just one doesn't have, and the variance between platforms tells you something concrete: which data sources are actually feeding these systems, and which aren't.
The floor here is running each probe multiple times across multiple sessions before treating any single output as signal rather than noise. Below that, you're looking at a random draw and calling it a finding.
Documentation matters more than it sounds like it should. Log the exact query text, the model and version, the session date, and the full response, not a summary of it. Auditors who paraphrase instead of transcribe introduce their own drift into the dataset, which undoes the entire point of running a controlled comparison in the first place. Running the identical query set against direct competitors turns a brand perception exercise into competitive intelligence; the gap between how a model talks about you and how it talks about a rival is often the single most useful number in the whole audit.
Translating query outputs into a scored, actionable audit finding
An interesting answer and an actionable finding aren't the same thing, and plenty of audits stop at the first. An interesting answer tells you what the model said. An actionable finding tells you which specific signal is weak, why, and what kind of fix would move it.
The four dimensions above become four scoring categories. Entity recognition rate tracks how reliably the model identifies your brand across unprompted recall. Factual accuracy rate tracks the share of verifiable claims the model gets right versus invents, broken out by claim type, so you know whether the weak point is leadership, geography, or product detail. Evaluative positioning score combines sentiment, appearance rate in recommendation-style queries, and share of category-level answers where you're actually named. Narrative consistency index measures the variance in how you get described across framings and platforms.
Not every gap deserves the same urgency. A hallucinated leadership claim, your CEO's name coming back wrong across multiple sessions, sits in a different risk tier than a missing product attribute. The audit's output should rank findings by harm and by how tractable the fix actually is, not just by how surprising the finding sounded.
One fix worth flagging directly: benchmark research cited by Status Labs, drawing on Data World's model comparisons, found LLMs grounded in structured knowledge graphs achieve dramatically higher accuracy than models relying on unstructured data alone. Structured entity data is a direct lever against hallucination rate. It's not just good organizational hygiene.
Doing all of this by hand, one brand, one query set, scored manually, works fine at small scale. It stops working the moment you're tracking a portfolio of brands or trying to hold a baseline over months instead of capturing one snapshot. That's the gap platforms built for this specific job are meant to close. Evident, for instance, scores across hundreds of signals spanning algorithmic, AI, and human evaluation in one framework, turning a one-off audit into something closer to an ongoing measurement practice.
Query design produces a diagnosis. Without a repeatable scoring system behind it, that diagnosis is stale the day it's written. The quality of the whole audit was decided before anyone typed a single prompt, by whether the questions were built to surface what these models actually use to judge a business, or built around what the auditor already assumed the model would say.


