AI Brand Perception Audit for a SaaS Company With No Branded Press Coverage
LLMs rank brands by signals most marketing teams have never measured.

A SaaS company with no press coverage gets scored on signals most marketing teams have never audited, and the score usually comes back worse than they expect. This piece maps those signals and shows how to run the audit yourself.
Call it the press-dark condition: no TechCrunch writeup, no Gartner mention, no earned media of any kind. It's common among early-stage and mid-market SaaS companies that grew through product-led growth or direct outbound sales instead of PR. For years that gap didn't matter much. A company could rank on page one of Google through solid SEO, close deals through the sales motion, and never think twice about what a journalist made of them.
That gap now has a second dimension. Per 6sense's 2025 Buyer Experience Report, the vast majority of B2B buyers used generative AI somewhere in their purchase research. A VP of Marketing evaluating tools doesn't always start with a Google search anymore. She opens ChatGPT, asks for a recommendation, gets two or three confident answers, and never opens a second tab. If your company isn't one of those names, the deal is lost before she starts comparing anything at all.
The mechanism is omission. LLMs build brand familiarity from what showed up in training data and what gets pulled in at query time from sources the model trusts: news coverage, industry publications, review sites, forum threads. A company that never generated any of that has, from the model's vantage point, little evidence it exists as a credible option. It's simply absent from the conversation.
Search invisibility and AI invisibility are related but separate failures, and conflating the two is the first mistake most teams make. A company can rank first for its category on Google, run a mature content program, and still vanish from every major AI platform the moment a buyer asks for a recommendation, because search rankings and AI recommendations draw from overlapping but genuinely distinct signal sets. That's the reason this problem goes unnoticed for so long: the metrics that used to reassure you don't measure the thing that's now quietly failing.
How LLMs actually decide which brands to surface and recommend
Three things happen when a model decides whether to say your name out loud.
Training data prevalence comes first: how often the brand shows up in what the model learned from, in what context, alongside what sentiment. Retrieval (RAG) comes second; the model reaches out in real time to sources it trusts, Wikipedia, review aggregators, industry publications, to check or supplement what it already "knows." Third is confidence scoring, an internal probability estimate of whether naming your brand is safe and accurate. High confidence sounds like "X is excellent for this." Low confidence sounds like a hedge, a qualifier, or nothing at all.
What's changed is the shape of the output itself. A search engine used to hand back ten blue links and let you judge for yourself. An AI answer names one, maybe two or three brands, and puts its own credibility behind that shortlist. Missing the list now costs a lot more than ranking eighth on a results page ever did.
A June 2026 study on arxiv.org looked at 131,514 backbone citations across twelve languages and found Wikipedia was the most-cited domain in eleven of them. That tells you something concrete about what these systems trust: encyclopedia-grade, third-party-verified content, weighted well above anything self-published. No Wikipedia presence and no equivalent outside corroboration means you're already rowing against the current of how these models weight sources.
Semantic association makes the problem worse. Models connect brands to problems through patterns they've learned, not a lookup table. If your company has never been mentioned anywhere near "scaling customer support," the model has nothing to draw on when someone asks about scaling customer support, whether or not your product happens to be the right answer.
Platforms don't behave alike, either. Claude tends to be the most selective, naming a smaller share of tested companies than ChatGPT or Gemini. Perplexity, in that same citation study, pulled from the widest pool by far, citing 90,276 of the 131,514 total backbone citations across nearly 16,000 distinct domains. Test one platform and you've learned almost nothing about the rest. Any audit worth the name has to span several systems at once.
Once that logic clicks, the audit's job gets simple: measure the specific signals feeding training prevalence, retrieval, and confidence, one at a time, instead of squinting at "AI visibility" as one blurry thing.
The signal categories a press-dark SaaS company needs to audit
Four categories here, each worth auditing on its own terms.
Entity coherence first. Does your company's name, category label, founding date, product description, and use-case framing match across LinkedIn, Crunchbase, G2, Capterra, your own site, wherever else you turn up? LLMs synthesize across sources rather than trusting any single one, so contradictions between them read as ambiguity, and ambiguity drags confidence scores down. For a press-dark company, the entity definition is often self-reported and nowhere else, which is about as weak a signal as exists.
Third-party citation coverage is the second category, and probably the heaviest one. Otterly.ai's 2025 research found brand mentions correlate with AI visibility at 0.664, roughly three times the correlation of backlinks alone, which sat at 0.218. What counts here: independent mentions on review platforms, in community threads, in newsletters, comparison articles, podcast show notes. Your own backlinks and case studies carry much less weight unless someone outside the company picks them up and repeats them.
Semantic topical association is the third, and it's the sneaky one. Your product pages can be excellent, specific, well-written, and none of it matters if no third-party content ever puts your name next to the problem language your buyers actually type. Ask it plainly: when a competitor gets mentioned in a discussion of your category, does your name show up too, or only theirs?
Technical crawlability is the fourth, and it's the one people skip because it feels like an engineering problem rather than a marketing one. robots.txt restrictions, JavaScript-heavy rendering, certain CDN configurations, any of these can block AI crawlers outright. A company can rack up decent mentions elsewhere and still be functionally invisible because the crawler never got past the front door. Fix this before spending a dollar on anything else, because everything else depends on being reachable in the first place.
Paid placements, follower counts, and brand aesthetics carry little to no weight in an LLM's evaluation, worth saying plainly, so chasing them wastes time you don't have.
Running the AI presence baseline: what to query, where, and what to record
Don't run one prompt and call it done. Build an inventory, fifteen to twenty-five prompts that mirror how a real buyer actually asks about your category. Split them three ways: category-level ("what's the best tool for X"), problem-level ("how do teams solve Y"), and competitor-adjacent ("alternatives to [competitor]").
Run every prompt across at least four platforms, ChatGPT, Perplexity, Claude, and Google AI Overviews. The answers will diverge, sometimes sharply, and that divergence is data, not noise you average away.
For each response, log whether your company gets named at all, where it lands (first mention, buried mid-list, tucked into a parenthetical), the actual language used (confident, hedged, absent), and which sources the platform cites when it names a competitor instead.
There's a rough benchmark worth anchoring to. A citation rate of 20 to 30% across a well-built prompt set marks meaningful AI visibility. Below 10%, you're effectively invisible. Above 40% is rare and usually reflects sustained category leadership.
DerivateX's analysis of fifty B2B SaaS companies across 1,400 buyer-intent prompts found an average AI Presence Score of 56.9 out of 100, with 44% of companies scoring below 50. That's useful for calibrating expectations: even companies with active press and content programs often land in mediocre territory. A press-dark company should expect to start well below that average, not near it.
Track one more thing alongside the raw citation rate: how competitors get described versus how you get described, assuming you appear at all. That sentiment gap often tells you more than the absence does. Doing this by hand across dozens of prompts and four platforms gets tedious fast; tools built for scoring AI and algorithmic visibility signals, Scale Labs among them, exist specifically to make this phase rigorous instead of a handful of manual checks run once and forgotten.
The output here is a document, not an impression: citation rate by platform, sentiment register by prompt category, and a specific list of prompts where competitors show up and you don't.
Auditing entity coherence and third-party citation coverage
Start by searching your company name in quotes across Google, Bing, and each AI platform, and write down how you're described everywhere it surfaces. Then go source by source: LinkedIn company page, Crunchbase, G2, Capterra, Product Hunt, any niche directory relevant to your category. Check category label, founding date, employee count, product description, and use-case framing for consistency.
Small inconsistencies matter more than you'd guess. "Project management tool" on one page and "workflow automation platform" on another reads to a model as real ambiguity about what the entity even is, and ambiguity is exactly what lowers confidence scores. For a press-dark company this is often the whole problem in miniature: every piece of entity information came from the company itself, and self-reported information carries less weight with these systems than anything corroborated externally.
The citation side runs in parallel. List every indexed domain mentioning your company by name, then sort by type: review platforms, forums like Reddit or Hacker News, newsletters, comparison articles, podcasts with published show notes, industry directories. The arxiv.org study found Wikipedia-class sources, encyclopedic, cross-referenced, stable over time, dominate what these systems actually cite. A press-dark SaaS company almost certainly has zero presence in that tier, and that gap alone explains a large chunk of the visibility problem.
Run the identical audit against one or two direct competitors. Comparing source counts and source types side by side turns a vague feeling of "we're behind" into a concrete, ranked list of what to fix.
What "good" looks like to a model isn't a homepage or a LinkedIn post. It's several independent sources, in domains the model already trusts, discussing your company in the context of a specific problem it solves. The most common finding among press-dark SaaS companies: strong internal content, next to nothing external. That gap is usually the single highest-leverage place to start.
Auditing technical crawlability and structured data
If a crawler can't reach your content, none of your mention signals or entity data can get confirmed or updated at query time. Crawlability isn't a nice-to-have layered on top; it's the prerequisite everything else sits on.
Check robots.txt for explicit or accidental blocks on GPTBot, ClaudeBot, PerplexityBot, and Google-Extended. It's more common than it should be for a broad disallow rule, added during some unrelated security review, to end up excluding every major AI crawler without anyone noticing for months.
Rendering comes next to check. A heavy React or Vue frontend building content client-side can serve a crawler an empty shell while a human visitor sees a full page. Fetch the page the way a bot would and compare it against what loads in a browser. The gap, when there is one, is often stark.
Structured data comes after that. Is Schema.org markup present for Organization, Product, FAQPage, and HowTo where relevant? Structured markup hands machines an unambiguous, machine-readable description of what your company is and does, exactly the kind of disambiguation that raises confidence scores.
BrightEdge's September 2025 finding fits here: most AI Overview citations came from pages outside the traditional top-ten search results. A page doesn't need to dominate organic rankings to earn an AI citation. It needs to be crawlable, clearly structured, and ready to answer the question as asked.
The output of this stage is short and binary: crawlers allowed or blocked, rendering method, structured data present or absent. Answer those three and you know whether technical work has to happen before any content or citation effort will ever register.
Reading AI sentiment: the difference between being mentioned and being recommended
Being named and being recommended are separate events, and conflating them is the second common mistake. Traditional sentiment analysis reads explicit opinion in reviews and social posts; AI sentiment works differently, capturing the way a model constructs and frames your brand as it synthesizes everything it's absorbed, the tone, the confidence, the problem contexts it reaches for when your name comes up. It's platform-specific too. The same prompt can land warm on one system and cold on another.
There's a spectrum worth listening for. Strong endorsement sounds like "X is well-suited for" or "X is widely used by." Neutral mention sounds like "X is an option" or "X offers." Hedged language sounds like "X might be worth exploring" or "X has some reviews suggesting." And then there's plain absence: no mention at all, even on a prompt built directly around what you do.
Press-dark companies cluster in the hedged-to-neutral zone when they show up at all. The model picked up some signal, just not enough corroboration to commit to a real recommendation. That's a specific, diagnosable state, not a mystery to shrug at.
To audit it, classify every response in your baseline set against that spectrum, then look for patterns: which problem contexts get confident language, which get hedges. That split shows exactly where your third-party signal runs thick and where it's thin. Compare your language against a competitor's on identical prompts. If they get "excellent for X" and you get "might be worth considering" on the same query, that's a sentiment gap, and it exists even though both companies technically got mentioned.
This isn't cosmetic. Research from Seer Interactive found that being cited in AI responses, versus excluded, correlates with meaningfully stronger organic and paid click performance, and the trust gap between cited and uncited brands widens over time rather than holding steady. Sentiment sits upstream of traffic, and eventually, upstream of revenue.
Prioritizing fixes: which signal gaps to close first when starting from near zero
Not every gap deserves equal attention, and a press-dark company rarely has the budget to chase all four at once. The audit's real value is telling you what to spend on first.
Fix crawlability first, always. If AI crawlers are blocked, nothing else you fix, not a citation win, not entity cleanup, not a structured data project, will register, because the systems that would notice your improvements literally cannot reach the pages where those improvements live. This is cheap, fast, and sits almost entirely within an engineering team's control, which makes it the highest-leverage first move by a wide margin.
Entity coherence comes next, since it's also mostly within your control and doesn't depend on anyone outside the company agreeing to write about you. Fixing mismatched category labels and descriptions across LinkedIn, Crunchbase, G2, and your own site takes days, not months, and it clears out ambiguity that's been working against every other signal quietly the whole time.
Third-party citation coverage is the slow one. It depends on other people, reviewers, forum users, journalists, choosing to mention you, and that can't be forced onto a sprint schedule. Otterly.ai's correlation data points here hardest: mentions matter roughly three times as much as backlinks for AI visibility, so this is where the sustained effort belongs even though the results lag behind the work.
Semantic association tends to follow once citation coverage improves, since it's really just a byproduct of getting discussed in the right contexts often enough for a model to learn the pattern on its own.
Start with what you control, move to what you can influence, and treat what depends on others as the long game it actually is. A press-dark company that runs this audit honestly usually finds the same story: solid product, thin outside corroboration, a fixable technical gap sitting in front of all of it. That's a plain diagnosis, and precision is what makes the fix possible in the first place.


