GEO Guides AI Brand Perception

Why AI Visibility Scores Do Not Tell You What AI Believes About Your Brand

Learn why AI visibility scores are not enough to measure brand perception, and how an LLM Perception Score can evaluate favorability, narratives, message pull-through, citations, source authority, competitive position, and durability.

Man surrounded by a dense press crowd with cameras and microphones under overhead light

AI visibility scores answer a useful question:

Did the brand appear?

They do not answer the more important question:

What did the AI system believe about the brand when it appeared?

A company may receive a high AI visibility score because it is frequently named across monitored prompts.

But the brand may be described as:

  • A legacy incumbent

  • A lower-cost alternative

  • A company facing regulatory scrutiny

  • A secondary player

  • A product with limited functionality

  • A trusted market leader

  • An innovative category creator

  • A company defined by an outdated controversy

All of these outcomes can produce visibility.

They do not produce the same perception.

This is the central weakness of AI visibility scoring: it often treats brand presence as the outcome when presence is only the beginning of the analysis.

A meaningful measurement framework must determine:

  • Why the brand appeared

  • How prominently it appeared

  • Which narrative defined it

  • Whether the portrayal was favorable

  • Which messages pulled through

  • Whether the answer was accurate

  • How the brand compared with competitors

  • Which sources supported the interpretation

  • Whether the perception persisted across models and repeated runs

  • Whether web retrieval strengthened or weakened it

  • Whether the broader earned-media environment supported the same conclusion

Those dimensions can be combined into a score.

But that score should measure perception, not merely visibility.

What is an AI visibility score?

An AI visibility score generally measures how often a brand appears across a defined set of prompts or AI-generated answers.

Depending on the methodology, it may incorporate:

  • Brand mention frequency

  • Percentage of prompts containing the brand

  • Average answer position

  • Share of AI voice

  • Recommendation frequency

  • Citation frequency

  • Model coverage

  • Competitive inclusion

  • Search impressions

  • Page appearances in AI features

For example, a company may test 100 prompts across ChatGPT, Claude, Gemini, Perplexity, and Grok.

If the brand appears in 70 of those answers, it might receive a visibility score of 70%.

More sophisticated versions may assign additional weight when:

  • The brand appears first

  • The brand is recommended

  • The brand receives more answer space

  • Its website is cited

  • It appears across multiple models

  • It is named alongside fewer competitors

These additions improve the signal.

But visibility remains fundamentally a measure of presence and exposure.

It does not fully measure interpretation.

Visibility is not perception

Visibility tells you that the brand entered the answer.

Perception tells you what the answer caused a stakeholder to believe.

Consider a question such as:

Which companies lead enterprise cybersecurity?

A brand may appear in the first paragraph because it is described as:

One of the most established and widely deployed enterprise-security platforms, with strong threat intelligence and broad customer adoption.

The same brand may appear in another answer because it is described as:

A major cybersecurity provider that has faced criticism following a series of high-profile security incidents.

Both answers generate a brand mention.

Both may place the company prominently.

Both may cite authoritative sources.

One strengthens the brand.

The other creates reputational risk.

A visibility score may treat them as similarly successful.

A perception score should not.

High visibility can reflect a negative narrative

Brand prominence is not inherently positive.

Companies often become highly visible in AI-generated answers because of:

  • Controversies

  • Lawsuits

  • Product failures

  • Regulatory actions

  • Security incidents

  • Executive misconduct

  • Customer complaints

  • Financial distress

  • Political disputes

  • Workforce reductions

  • Failed launches

  • Public criticism

During a crisis, a brand may achieve its highest-ever visibility.

That does not mean its AI performance improved.

This is why visibility cannot be interpreted without brand-centric favorability.

The analysis must ask:

  • Does the answer increase or reduce confidence in the brand?

  • Is the company portrayed as responsible for the problem?

  • Is it shown responding effectively?

  • Is the issue described as resolved or ongoing?

  • Does the answer distinguish allegation from confirmed fact?

  • Does the system repeat outdated or inaccurate claims?

  • Is the company framed more negatively than competitors?

Without this context, a visibility score can reward reputational damage.

High visibility can reinforce the wrong category

A brand can appear frequently while being associated with a perception it is actively trying to change.

For example, a company may want to be understood as a broad enterprise platform.

AI systems may mention it in most relevant answers but continue describing it as:

  • A point solution

  • A single-product company

  • A small-business tool

  • A legacy provider

  • A consumer brand

  • A lower-cost alternative

  • A niche specialist

The brand is visible.

The repositioning has failed.

This matters when companies are trying to:

  • Move beyond a legacy product

  • Enter a new market

  • Establish category leadership

  • Expand into enterprise accounts

  • Build an AI-forward identity

  • Shift from services to software

  • Position a portfolio rather than one product

  • Recover from an outdated reputation

  • Compete on value rather than price

Visibility alone cannot reveal whether the desired category association is present.

That requires narrative and message analysis.

A citation does not make the perception favorable

Some AI visibility tools treat citations as an inherently positive outcome.

Citations are valuable, but they are not automatically beneficial.

A brand may be cited through:

  • A product page

  • An investigative article

  • A regulator

  • A lawsuit

  • A customer-review site

  • A security notice

  • A negative analyst report

  • An outdated press release

  • A competitor comparison

  • A favorable customer story

Each citation contributes different evidence.

The important questions are:

  • Which source was cited?

  • What claim did it support?

  • Was the source authoritative for that claim?

  • Was the information current?

  • Was the brand central to the source?

  • Did the source reinforce the desired or undesired narrative?

  • Did the citation recur across models and runs?

  • Did the answer accurately represent the source?

  • Was the brand cited as evidence of leadership, weakness, risk, or controversy?

Citation frequency measures observed source selection.

Citation interpretation reveals reputational meaning.

Search impressions do not measure AI belief

Search platforms increasingly provide reporting on how websites appear in generative AI experiences.

These reports are useful, but they primarily measure content visibility.

Google’s generative AI reporting in Search Console includes information such as:

  • Impressions

  • Pages

  • Countries

  • Devices

  • Performance over time

Bing Webmaster Tools’ AI Performance reporting includes:

  • Citation activity

  • Cited URLs

  • Citation trends

  • Grounding queries

  • Page-level activity

These tools help website owners understand whether their pages are being surfaced.

They do not directly determine:

  • Whether the brand was portrayed favorably

  • Which corporate narrative dominated

  • Whether the answer was accurate

  • Whether priority messages appeared

  • How the brand compared with competitors

  • Whether the source supported a negative conclusion

  • Whether the company was recommended

  • Whether the perception persisted without web retrieval

  • Whether the answer reflected the wider earned-media environment

Website visibility and brand perception are related.

They are not the same metric.

Share of AI voice has the same limitation

Share of AI voice generally measures the brand’s portion of appearances relative to competitors.

A simple version may use:

Brand appearances ÷ Total appearances across measured brands

This can help reveal competitive presence.

But it may treat all appearances as equal.

Consider two brands:

Brand A

  • Appears in 70% of answers

  • Is usually named in a long list

  • Receives little substantive explanation

  • Is rarely recommended

  • Has weak message pull-through

  • Is described neutrally

Brand B

  • Appears in 50% of answers

  • Is frequently named first

  • Is described as the category leader

  • Receives detailed favorable analysis

  • Is recommended for the highest-value use cases

  • Is supported by authoritative citations

Brand A has greater raw share of AI voice.

Brand B has stronger perception.

A useful competitive framework should therefore account for:

  • Answer prominence

  • Narrative relevance

  • Favorability

  • Recommendation

  • Category ownership

  • Message pull-through

  • Citation support

  • Source authority

  • Substantive discussion

Share of AI voice should remain one component, not the final measure.

Recommendation rate is also incomplete

Recommendation is often more meaningful than presence.

But recommendation alone still requires context.

An AI system may recommend a company:

  • As the best enterprise option

  • As the cheapest option

  • For small teams only

  • For users willing to accept limited functionality

  • As one of several acceptable choices

  • For a narrow use case

  • With significant qualifications

  • Despite concerns about service or reliability

These are materially different outcomes.

A recommendation score should record:

  • Whether the brand was recommended

  • For which customer or use case

  • Whether the recommendation was qualified

  • Which strengths drove it

  • Which weaknesses limited it

  • Whether a competitor was preferred overall

  • Which evidence supported the recommendation

  • Whether the recommendation persisted across repeated runs

A recommendation without interpretation can create the same false confidence as a visibility score.

General sentiment does not solve the problem

Some AI visibility systems add sentiment analysis to improve the metric.

That is better than counting mentions alone, but traditional sentiment has a major limitation: it often measures the emotional tone of the answer rather than the brand’s position within it.

For example:

Small businesses are facing severe financial pressure, but the company has helped thousands of employers reduce administrative costs and preserve jobs.

The broader subject is negative.

The brand’s position is positive.

A general sentiment system may misclassify the answer because it detects words such as:

  • Severe

  • Pressure

  • Costs

  • Job losses

  • Financial difficulty

Brand-centric favorability asks a more precise question:

Does this answer strengthen or weaken confidence in the brand?

That is the relevant measure for AI perception.

The prompt set can distort the score

Every AI visibility score is partly a product of the prompts selected.

A company may appear highly visible because the monitoring program includes questions that naturally favor it.

For example:

  • Prompts may use the company’s preferred category language.

  • Competitors may be omitted.

  • Questions may be narrowly tailored to the brand’s strongest product.

  • Adverse questions may not be included.

  • Open-ended prompts may be underrepresented.

  • The prompts may reflect internal messaging rather than real stakeholder language.

  • Similar prompts may be counted as separate evidence.

  • A brand name may be included in the prompt itself.

A score of 85 does not mean much without knowing what was tested.

A credible methodology should include prompt families across:

  • Factual questions

  • Category questions

  • Comparative questions

  • Evaluative questions

  • Adverse questions

  • Open-ended questions

  • Narrative-specific questions

  • Recommendation questions

The prompt set should be representative, not promotional.

Exact prompts do not represent every stakeholder question

People can ask about the same brand narrative in thousands of ways.

They may vary:

  • Terminology

  • Context

  • Use case

  • Geography

  • Industry

  • Company size

  • Risk tolerance

  • Time horizon

  • Competitors

  • Decision criteria

  • Follow-up questions

A communications team may monitor:

Is Apple a leader in artificial intelligence?

A customer may ask:

Which smartphone company has the most useful consumer AI features?

An investor may ask:

Is Apple’s AI strategy strengthening or weakening its competitive position?

A journalist may ask:

What evidence supports Apple’s claims about privacy-focused AI?

A developer may ask:

How open is Apple’s AI ecosystem compared with Google or Microsoft?

These prompts belong to the same broad narrative.

They require different evidence.

The narrative should therefore be the unit of analysis.

Prompts are tools used to test it.

One answer is not a score

AI-generated answers can vary between runs.

The same prompt may produce different:

  • Brand lists

  • Rankings

  • Recommendations

  • Claims

  • Citations

  • Competitive comparisons

  • Levels of detail

  • Favorability

  • Narrative framing

A score based on one response per prompt may create false precision.

Repeated testing is necessary to determine:

  • Whether the brand appears consistently

  • Whether the same narrative recurs

  • Whether favorability is stable

  • Whether citations repeat

  • Whether recommendations change

  • Whether one answer was an outlier

  • Whether the result is model-specific

  • Whether prompt wording changes the conclusion

A meaningful score should include a confidence rating based partly on cross-run stability.

Web-on and web-off answers measure different conditions

AI perception may change substantially when web retrieval is enabled.

With web retrieval

The answer may reflect:

  • Current news

  • Current owned content

  • Recent regulatory developments

  • Newly published research

  • Updated product information

  • Visible citations

  • Current competitor evidence

Without visible web retrieval

The answer may express:

  • More established associations

  • Older narratives

  • Persistent category perceptions

  • Outdated facts

  • Different competitive framing

  • A different level of confidence

These conditions should be measured separately.

A brand can have:

  • Strong web-on visibility but weak web-off perception

  • Strong web-off perception but unfavorable current retrieval

  • High visibility in both conditions but different narratives

  • Accurate web-on answers but persistent web-off inaccuracies

  • High citation activity without durable perception

A single visibility score may hide these differences.

The gap between web-on and web-off performance matters

The difference between retrieval conditions can reveal whether a narrative is:

  • Emerging

  • Durable

  • Outdated

  • Fragile

  • Dependent on current news

  • Vulnerable to negative retrieval

  • Not yet established across models

For example:

High web-on, low web-off performance

The current information environment supports the desired narrative, but the perception may not yet be durable.

Low web-on, high web-off performance

The brand has a favorable established reputation, but current evidence is weakening it.

High performance in both

The perception may be strong and resilient.

Low performance in both

The desired narrative lacks support across both current retrieval and persistent model outputs.

This can be measured as retrieval resilience.

Visibility alone does not capture it.

Visibility scores often ignore the evidence environment

An AI answer is connected to a broader body of evidence.

That environment may include:

  • Earned media

  • Owned content

  • Product documentation

  • Regulatory records

  • Customer evidence

  • Research

  • Expert commentary

  • Social and community sources

  • Business directories

  • Public filings

  • Competitor content

A visibility score may record that the brand appeared.

It may not explain why.

For example, a competitor may outperform the brand because it has:

  • Clearer category-definition pages

  • Stronger product documentation

  • More authoritative earned media

  • Better customer evidence

  • Greater brand prominence

  • More current information

  • More consistent executive messaging

  • Stronger independent corroboration

  • Higher-quality comparison coverage

  • Better crawl accessibility

Without evidence analysis, the score identifies the symptom but not the cause.

Earned media should be part of perception measurement

Earned media helps establish how independent sources understand the brand.

It can:

  • Validate claims

  • Challenge positioning

  • Define categories

  • Compare competitors

  • Document outcomes

  • Build executive authority

  • Establish market significance

  • Surface reputational risk

  • Create durable associations

A perception framework should examine:

  • Which narratives dominate coverage

  • Whether the brand is central or incidental

  • How the brand is positioned

  • Which sources have authority

  • Which claims recur

  • Whether the coverage is independently reported

  • Which messages pull through

  • Whether sources converge on the same interpretation

  • Which competitor narratives are stronger

  • Whether the coverage supports or conflicts with AI answers

Raw AI visibility ignores much of this evidence.

Visible citations are only one evidence layer

Visible citations show which sources were presented with a particular answer.

They are valuable and directly observable.

But they should not automatically be treated as a complete explanation of every factor involved in producing the response.

AI companies do not publish a full formula for every answer.

A complete source analysis should distinguish among:

Observed citation behavior

Which URLs and domains visibly appear across tested answers?

Likely citation influence

Which sources are best positioned to shape future retrieval and citations based on:

  • Authority

  • Relevance

  • Factual specificity

  • Brand prominence

  • Freshness

  • Independence

  • Accessibility

  • Repetition

  • Narrative alignment

The broader evidence environment

Which earned, owned, expert, regulatory, customer, social, and research sources collectively support the narrative?

A visibility score may capture a portion of the first layer.

A perception framework should examine all three.

What AI visibility scores are good for

AI visibility scores still have value.

They can help answer:

  • Is the brand appearing at all?

  • How often does it appear?

  • Which models include it?

  • Which competitors appear more often?

  • Is the brand being recommended?

  • Is its website cited?

  • Are certain pages surfacing?

  • Is visibility changing over time?

  • Is the brand absent from important category questions?

  • Has a campaign increased exposure?

These are legitimate measurement questions.

Visibility can be particularly useful for:

  • Establishing a baseline

  • Tracking broad category presence

  • Identifying competitive gaps

  • Finding prompts that require deeper analysis

  • Monitoring changes after a launch

  • Measuring website citation activity

  • Detecting emerging inclusion

  • Prioritizing narratives for investigation

The problem is not visibility measurement.

The problem is treating visibility as the complete definition of AI performance.

The better alternative: an LLM Perception Score

A defensible LLM Perception Score should measure the quality, strength, accuracy, and durability of the brand’s interpretation across AI systems.

It can combine:

  • Narrative presence

  • Brand prominence

  • Brand-centric favorability

  • Message pull-through

  • Factual accuracy

  • Competitive position

  • Recommendation quality

  • Citation consistency

  • Citation quality

  • Source authority

  • Cross-model consistency

  • Cross-run stability

  • Retrieval resilience

  • Earned-media evidence strength

  • Narrative trajectory

This score answers a different question.

An AI visibility score asks:

How often did the brand appear?

An LLM Perception Score asks:

How strong, favorable, accurate, consistent, competitive, and well-supported is the brand’s perception?

That is a more valuable executive measure.

A perception score should include visibility

Visibility should not be discarded.

It should be incorporated as one component of the broader score.

For example:

Component What it measures
Narrative presence Whether the brand appears within the priority narrative
Prominence How central the brand is to the answer
Favorability Whether the answer strengthens or weakens confidence
Message pull-through Whether priority messages survive synthesis
Accuracy Whether material claims are correct and current
Competitive position How the brand compares with alternatives
Recommendation quality Whether and why the brand is recommended
Citation performance Which sources recur and what they support
Evidence strength Authority, independence, specificity, and corroboration
Model consistency Stability across AI systems
Run stability Stability across repeated responses
Retrieval resilience Performance with web retrieval on and off
Earned-media environment Strength of the wider independent evidence
Narrative trajectory Whether perception is strengthening or weakening

Visibility matters.

It simply should not dominate everything else.

How to score narrative presence

Narrative presence should distinguish between levels of inclusion.

A useful rubric may classify a result as:

Primary

The brand is presented as a leading example or central subject.

Substantive

The brand is meaningfully connected to the narrative.

Incidental

The brand appears but contributes little to the answer.

Absent

The brand does not appear.

This prevents a passing mention from receiving the same credit as category leadership.

How to score prominence

Prominence can incorporate:

  • Appearance in the opening

  • Position in a list

  • Amount of substantive discussion

  • Inclusion in headings

  • Presence in the conclusion

  • Use as the primary example

  • Share of answer space

  • Whether the brand defines the category

A brand should receive more credit when it frames the answer than when it appears as an afterthought.

How to score brand-centric favorability

Favorability should assess the brand’s position rather than general emotional tone.

A practical rubric may classify answers as:

  • Strongly positive

  • Positive

  • Neutral

  • Mixed

  • Negative

  • Strongly negative

The classification should consider whether the answer:

  • Builds confidence

  • Validates leadership

  • Reinforces differentiation

  • Raises concerns

  • Describes failures

  • Introduces qualifications

  • Associates the company with risk

  • Shows effective response to a difficult issue

The reason behind the classification should remain visible.

How to score message pull-through

For each priority message, classify whether it is:

  • Explicit

  • Implied

  • Absent

  • Contradicted

A brand may appear frequently while its most important message never appears.

For example, a company may want to be understood as more than its traditional product.

If AI systems repeatedly name the company but continue describing it only through the legacy category, visibility is high and message pull-through is low.

That distinction is central to perception measurement.

How to score factual accuracy

Accuracy should be evaluated at the claim level.

Classifications may include:

  • Accurate

  • Incomplete

  • Outdated

  • Unsupported

  • Incorrect

A high-visibility answer containing material inaccuracies should not receive a strong perception score.

Accuracy is especially important for:

  • Executive leadership

  • Product availability

  • Pricing

  • Regulatory status

  • Company ownership

  • Transactions

  • Market data

  • Corporate history

  • Security events

  • Litigation

  • Product capabilities

A score that ignores accuracy can reward confident misinformation.

How to score competitive position

Competitive analysis should examine:

  • Which brands appear

  • Which brand appears first

  • Which company is framed as the leader

  • Which differentiators are assigned

  • Which weaknesses are attached

  • Which use cases each brand owns

  • Whether the brand is recommended

  • Whether competitors have stronger evidence

  • Whether the category definition favors one company

A brand may have strong standalone visibility but weak comparative perception.

That difference should affect the score.

How to score citation performance

Citation performance should go beyond counting URLs.

It may incorporate:

  • Citation presence

  • URL recurrence

  • Domain recurrence

  • Source authority

  • Source freshness

  • Brand prominence in the source

  • Claim relevance

  • Independent reporting

  • Citation-to-claim support

  • Cross-model recurrence

  • Cross-run recurrence

  • Whether the source supports the desired or undesired narrative

A cited corporate page may be highly useful for product specifications.

A cited regulator may be highly authoritative for an enforcement action.

A cited review site may be relevant to usability.

The source must be evaluated in relation to the claim.

How to score source authority

Authority is contextual.

Relevant factors include:

  • Institutional responsibility

  • Subject-matter expertise

  • Editorial standards

  • Original reporting

  • Firsthand evidence

  • Independent validation

  • Research quality

  • Transparency

  • Factual specificity

  • Current relevance

The largest publication is not always the strongest source.

A specialist journal may carry greater authority for a technical claim.

A regulator may carry greater authority for an approval.

A company page may carry greater authority for its current pricing.

The scoring rubric should recognize this.

How to score cross-model consistency

ChatGPT, Claude, Gemini, Perplexity, and Grok may produce different interpretations.

Cross-model consistency should measure whether the underlying conclusion remains stable.

Possible classifications include:

  • Highly consistent

  • Generally consistent

  • Mixed

  • Contradictory

  • Insufficient evidence

The goal is not identical wording.

It is a stable underlying perception.

A high average score combined with severe model disagreement should carry less confidence than a similar score supported by broad agreement.

How to score cross-run stability

Repeated runs reveal whether the result persists.

A stable perception should maintain similar:

  • Narrative presence

  • Favorability

  • Message pull-through

  • Competitive position

  • Recommendation

  • Factual claims

  • Citation patterns

One favorable answer should not receive the same weight as a narrative that recurs consistently.

Run stability is therefore both a performance input and a confidence input.

How to score retrieval resilience

Retrieval resilience measures whether perception holds when web retrieval is turned on and off.

A strong result may show:

  • Accurate perception in both conditions

  • Consistent favorability

  • Stable message pull-through

  • Limited narrative fragmentation

  • Current retrieval reinforcing the established perception

A weak result may show:

  • Favorable web-off answers but negative current retrieval

  • Accurate web-on answers but persistent outdated web-off beliefs

  • Large changes in competitive position

  • Current citations contradicting the desired narrative

  • Extreme sensitivity to retrieval conditions

This distinction is invisible in a single visibility score.

How to score earned-media evidence strength

Earned-media evidence strength should measure how well independent coverage supports the narrative.

Relevant dimensions include:

  • Source authority

  • Brand prominence

  • Brand-centric favorability

  • Narrative relevance

  • Factual specificity

  • Independent corroboration

  • Original reporting

  • Freshness

  • Narrative momentum

  • Competitive framing

  • Consistency across sources

  • Conflicting evidence

Coverage volume may contribute context.

It should not automatically increase the score.

One authoritative, substantive story may provide stronger evidence than dozens of passing mentions or syndicated articles.

How to score narrative trajectory

A perception score should not be static.

Narrative trajectory evaluates whether the brand’s perception is:

  • Strengthening

  • Weakening

  • Correcting

  • Fragmenting

  • Hardening

  • Fading

  • Remaining stable

This helps communications teams understand whether recent activity is changing the information environment.

A score of 70 that is strengthening may represent a different strategic situation from a score of 70 that is deteriorating.

Build the score from an explicit rubric

A composite score is valuable only when the methodology is understandable.

The rubric should define:

  • Each component

  • The scoring scale

  • The evidence required

  • The weighting

  • How missing data is handled

  • How web-on and web-off results are combined

  • How repeated runs affect confidence

  • How models are weighted

  • How earned-media evidence contributes

  • How adverse narratives are handled

  • How accuracy affects the result

  • How score changes are calculated

An unexplained number creates false authority.

A transparent score creates an executive summary that can be traced back to evidence.

Add a confidence rating

Performance and confidence should be reported separately.

Two brands can receive the same LLM Perception Score while the evidence behind the scores differs significantly.

A confidence rating may account for:

  • Number of models tested

  • Number of prompt families

  • Number of repeated runs

  • Cross-model agreement

  • Cross-run stability

  • Citation recurrence

  • Source quality

  • Evidence freshness

  • Earned-media corpus size

  • Retrieval-condition coverage

  • Longitudinal consistency

For example:

Score: 82

Confidence: High

This may indicate favorable perception supported across five models, repeated runs, strong citations, and authoritative earned media.

Score: 82

Confidence: Low

This may indicate a favorable result based on limited prompts, one model, or unstable outputs.

The score communicates performance.

Confidence communicates how strongly the evidence supports it.

A sample LLM Perception Score architecture

A company could organize the overall score into component categories such as:

Score component Example measures
Perception quality Favorability, accuracy, overall framing
Narrative strength Presence, prominence, message pull-through
Competitive position Leadership, differentiation, recommendation
Evidence strength Authority, specificity, independence, corroboration
Citation intelligence Citation quality, recurrence, claim support
Model consistency Cross-model and cross-run stability
Retrieval resilience Web-on and web-off alignment
Earned-media environment Independent narrative support
Narrative trajectory Direction and durability over time

The exact weighting should reflect the purpose of the analysis.

For example:

Corporate reputation

May place greater weight on:

  • Favorability

  • Accuracy

  • Source authority

  • Cross-model consistency

  • Earned-media evidence

  • Narrative durability

Product launch

May place greater weight on:

  • Narrative presence

  • Message pull-through

  • Product accuracy

  • Competitive position

  • Retrieval freshness

  • Citation activity

  • Earned-media momentum

Crisis analysis

May place greater weight on:

  • Factual accuracy

  • Negative narrative prevalence

  • Source authority

  • Current retrieval

  • Narrative trajectory

  • Correction persistence

  • Cross-model consistency

The score should remain comparable over time while allowing the framework to reflect the business question.

Visibility should remain visible within the score

A broader perception score should not hide the underlying visibility data.

A useful report may include:

  • Raw brand presence

  • Primary prominence

  • Share of AI voice

  • Recommendation frequency

  • Citation presence

  • Overall LLM Perception Score

  • Confidence rating

  • Narrative-level scores

  • Web-on score

  • Web-off score

  • Earned-evidence score

  • Competitive-position score

  • Message-pull-through score

This allows teams to see both exposure and interpretation.

For example:

Metric Result
Raw visibility 84%
Primary prominence 43%
Positive favorability 51%
Message pull-through 38%
Factual accuracy 93%
Recommendation rate 47%
High-authority citation share 42%
LLM Perception Score 61
Confidence High

This scorecard reveals that the brand appears often but is not converting that exposure into strong perception.

A high visibility score can coexist with a low perception score

Consider a hypothetical enterprise-software company.

The company appears in 88% of monitored answers.

A deeper analysis finds:

  • It is usually named after two competitors.

  • It is described as a legacy platform.

  • Its AI capabilities are framed as add-ons.

  • Priority messages appear in only 22% of answers.

  • Current citations include several older product articles.

  • Competitors receive stronger enterprise recommendations.

  • Favorability is neutral to mixed.

  • Web retrieval makes the comparison less favorable.

  • Earned media focuses heavily on implementation complexity.

  • The result is stable across models.

The company may have:

  • AI Visibility Score: 88

  • LLM Perception Score: 54

  • Confidence: High

The visibility score says the company is widely present.

The perception score says the current narrative is strategically weak.

Both numbers are useful.

Only one captures the reputational outcome.

A lower visibility score can coexist with a strong perception score

Now consider a new category challenger.

The company appears in only 42% of relevant answers.

When it appears:

  • It is described as innovative.

  • It receives strong message pull-through.

  • It is recommended for sophisticated use cases.

  • Its claims are supported by authoritative research.

  • It has favorable independent coverage.

  • Its citations are current and relevant.

  • Cross-run stability is strong.

  • Web-on perception is highly favorable.

  • Web-off perception is beginning to emerge.

The company may have:

  • AI Visibility Score: 42

  • LLM Perception Score: 79

  • Confidence: Moderate

The strategic issue is not perception quality.

It is narrative reach and durability.

The company should amplify strong evidence rather than rebuild its positioning from scratch.

Visibility alone would miss this distinction.

The score should diagnose the evidence gap

The purpose of measurement is not simply to produce a number.

It should reveal why the number exists.

Common evidence gaps include:

Presence gap

The brand is missing from relevant answers.

Prominence gap

The brand appears but is not central.

Favorability gap

The brand is visible but portrayed negatively or with significant qualifications.

Message gap

Priority claims do not survive synthesis.

Accuracy gap

Answers contain outdated or incorrect information.

Competitive gap

Competitors own the category or preferred use case.

Citation gap

The brand lacks recurring authoritative source support.

Authority gap

Most supporting evidence is company-authored or low quality.

Retrieval gap

Strong information exists but is difficult to access.

Durability gap

The narrative appears only under narrow prompts, one model, or web-on conditions.

Earned-evidence gap

The desired perception lacks independent corroboration.

Different gaps require different actions.

Connect the score to communications action

A perception score should help determine what the company should do next.

Amplify

Use when the desired perception is favorable but insufficiently prominent.

Clarify

Use when the right narrative appears but important distinctions are misunderstood.

Counter

Use when inaccurate or incomplete narratives are gaining strength.

Canonicalize

Use when important facts lack a clear permanent owned source.

Create

Use when the evidence required to support the desired narrative does not yet exist.

Validate

Use when a company claim lacks credible independent support.

Correct

Use when outdated or inaccurate information remains accessible.

Consolidate

Use when fragmented owned content creates conflicting interpretations.

Monitor

Use when the narrative is emerging but not yet stable.

The score should point to the evidence gap.

The evidence gap should determine the action.

Measure change over time

A one-time score is a baseline.

The strategic value comes from longitudinal measurement.

Track:

  • Overall LLM Perception Score

  • Raw visibility

  • Narrative-level scores

  • Favorability

  • Message pull-through

  • Competitive position

  • Citation quality

  • Earned-evidence strength

  • Web-on and web-off differences

  • Cross-model consistency

  • Cross-run stability

  • Narrative trajectory

  • Confidence

Then connect changes to:

  • Product launches

  • Earned-media campaigns

  • Research releases

  • Executive announcements

  • Issues and crises

  • Website updates

  • Category campaigns

  • Competitor activity

  • Regulatory developments

  • Customer evidence

This allows communications teams to show whether the information environment changed.

Report the conclusion before the metrics

Executives should not have to interpret hundreds of prompt results.

A strong report should begin with:

What AI systems currently believe

State the dominant perception directly.

How strong the perception is

Provide the LLM Perception Score and confidence rating.

What is driving the score

Identify the strongest positive and negative components.

How the brand compares with competitors

Explain leadership, differentiation, recommendation, and category position.

What changed

Describe narrative strengthening, weakening, correction, or fragmentation.

Why it changed

Connect the movement to sources, claims, earned media, owned content, and current events.

What should happen next

Provide evidence-based communications actions.

Visibility metrics should support this interpretation.

They should not replace it.

Questions to ask an AI visibility vendor

Before relying on an AI visibility score, ask:

  1. What exactly does the score measure?

  2. Does it measure presence or perception?

  3. Does it evaluate brand-centric favorability?

  4. Does it measure message pull-through?

  5. Does it assess factual accuracy?

  6. Does it distinguish primary prominence from passing mentions?

  7. Does it measure competitive positioning?

  8. Does it classify recommendation quality?

  9. Does it use repeated runs?

  10. Does it test multiple models?

  11. Does it separate web-on and web-off conditions?

  12. Does it map citations to claims?

  13. Does it evaluate source authority?

  14. Does it connect answers to the broader earned-media environment?

  15. Does it measure narrative durability?

  16. Does it include a confidence rating?

  17. Can the score be traced to the underlying evidence?

  18. Can it explain why the result exists?

  19. Can it identify what should change?

  20. Can it demonstrate movement over time?

A score that cannot answer these questions may still be useful.

It should be described accurately as a visibility metric.

Common mistakes with AI visibility scores

Treating all appearances as positive

A brand can be visible because of reputational risk.

Treating all mentions equally

A passing mention is not equivalent to category leadership.

Ignoring brand-centric favorability

General sentiment may misread how the brand itself is positioned.

Ignoring accuracy

A highly visible incorrect answer is not success.

Treating citations as inherently favorable

A citation may support criticism, controversy, or outdated information.

Using one response per prompt

Single-run results may be unstable.

Testing only favorable prompts

The score may reflect the assumptions of the testing program.

Ignoring web retrieval conditions

Web-on and web-off perception may differ materially.

Ignoring model differences

A combined score can conceal major disagreement.

Ignoring earned media

Visible AI outputs are connected to a larger evidence environment.

Hiding the rubric

An unexplained score creates false precision.

Reporting performance without confidence

A high score based on weak evidence should not be treated as durable.

Optimizing the score rather than the narrative

The goal is not to manipulate tested prompts.

It is to strengthen the underlying perception.

What good AI perception measurement should reveal

A complete measurement program should answer:

  1. Is the brand visible?

  2. How prominent is it?

  3. Which narrative defines it?

  4. Is the portrayal favorable?

  5. Which messages pull through?

  6. Are the claims accurate?

  7. How does the brand compare with competitors?

  8. Is it recommended?

  9. Why is it recommended or rejected?

  10. Which sources are cited?

  11. What do those sources support?

  12. How authoritative are they?

  13. Does the perception persist across models?

  14. Does it persist across repeated runs?

  15. Does web retrieval improve or weaken it?

  16. Does earned media support the same conclusion?

  17. Is the narrative strengthening or fading?

  18. How confident should leadership be in the result?

  19. What evidence gap is driving weakness?

  20. What communications action should follow?

A visibility score answers the first question.

An LLM Perception Score should answer the complete set.

The central lesson

AI visibility matters.

Brands need to know whether they appear in the answers their stakeholders receive.

But visibility is not the final outcome.

A company can be highly visible and poorly perceived.

It can be cited often and associated with risk.

It can lead share of AI voice while losing the category narrative.

It can appear across every model while its priority messages remain absent.

It can receive strong current visibility while an outdated perception remains durable without web retrieval.

The better measure is not whether the brand appeared.

It is whether AI systems consistently reproduce an accurate, favorable, differentiated, and well-supported interpretation of the brand.

That can be measured.

The underlying attributes can be defined in a rubric.

The rubric can produce a score.

The score can be compared across narratives, models, competitors, retrieval conditions, and time.

An AI visibility score measures exposure.

An LLM Perception Score measures what the exposure means.

For communications leaders, that is the outcome that matters.

Frequently asked questions

An AI visibility score measures whether or how frequently a brand appears across a defined set of AI-generated answers. It may also incorporate answer position, share of AI voice, recommendation frequency, citations, or model coverage.

Visibility measures presence. AI brand perception measures how the brand is interpreted, including its narrative, prominence, favorability, accuracy, message pull-through, competitive position, recommendations, citations, and supporting evidence.

Yes. A company may receive high visibility because of controversy, regulatory action, product failure, customer criticism, or another unfavorable narrative. Visibility must be interpreted alongside brand-centric favorability.

They are a useful observable signal, but not automatically a positive one. A citation may support a favorable claim, a negative claim, an outdated fact, a regulatory action, or a customer complaint. Citation quality and narrative contribution should be evaluated.

Yes, as one component of competitive measurement. It becomes misleading when every appearance is treated equally and prominence, favorability, recommendation, narrative relevance, and citation support are ignored.

Recommendation can be more meaningful, but it still requires interpretation. A brand may be recommended only for a narrow segment, as a budget option, or with substantial qualifications.

Brand-centric favorability measures whether the answer strengthens or weakens confidence in the brand itself. It differs from general sentiment, which may focus on the emotional tone of the broader subject.

An LLM Perception Score is a composite measure of the quality, strength, accuracy, consistency, and durability of a brand’s perception across AI systems. It can combine visibility, prominence, favorability, message pull-through, accuracy, competitive position, recommendation, citation quality, source authority, cross-model consistency, cross-run stability, retrieval resilience, earned-media evidence, and narrative trajectory.

Yes. Visibility and narrative presence are important components. They should not be treated as the complete score.

It can be synthesized into an overall score when the rubric is transparent and the underlying components remain visible. The score should act as an executive summary, not a substitute for evidence.

The same score may be supported by very different amounts and qualities of evidence. Confidence reflects sample depth, repeated-run stability, cross-model agreement, citation recurrence, source quality, retrieval coverage, and earned-media evidence.

Web-on testing shows how current accessible evidence changes the answer and which sources are surfaced. Web-off testing shows the perception expressed without visible current retrieval under the tested conditions. The difference reveals retrieval resilience and narrative durability.

Earned media provides independent evidence about the brand. It can validate or challenge company claims, establish significance, compare competitors, document outcomes, and create recurring narratives that shape both human and AI perception.

No. Search impressions measure whether pages appeared within relevant search or generative AI features. They do not directly measure how the brand was characterized or whether the perception was favorable, accurate, or strategically useful.

Not by themselves. Citation counts show observed activity. They do not automatically establish source authority, answer placement, narrative value, or influence across all contexts.

Report: • Overall LLM Perception Score • Confidence rating • Raw visibility • Narrative presence • Brand prominence • Brand-centric favorability • Message pull-through • Accuracy • Competitive position • Recommendation • Citation intelligence • Web-on and web-off results • Earned-media evidence strength • Narrative trajectory

Yes. A brand can improve the score by strengthening the information environment through clearer owned content, stronger factual evidence, authoritative earned media, independent corroboration, current sources, better message pull-through, corrected inaccuracies, and more consistent narrative support.

Yes. A newer or emerging brand may appear less frequently but be described very favorably and supported by strong evidence when it does appear. The strategic priority may be amplification rather than repositioning.

Yes. A well-known company may appear frequently while being characterized as outdated, risky, undifferentiated, or weaker than competitors.

They can confuse being mentioned with being understood favorably. The brand’s interpretation, not its mere appearance, is the real reputational outcome.