AI visibility scores answer a useful question:
Did the brand appear?
They do not answer the more important question:
What did the AI system believe about the brand when it appeared?
A company may receive a high AI visibility score because it is frequently named across monitored prompts.
But the brand may be described as:
A legacy incumbent
A lower-cost alternative
A company facing regulatory scrutiny
A secondary player
A product with limited functionality
A trusted market leader
An innovative category creator
A company defined by an outdated controversy
All of these outcomes can produce visibility.
They do not produce the same perception.
This is the central weakness of AI visibility scoring: it often treats brand presence as the outcome when presence is only the beginning of the analysis.
A meaningful measurement framework must determine:
Why the brand appeared
How prominently it appeared
Which narrative defined it
Whether the portrayal was favorable
Which messages pulled through
Whether the answer was accurate
How the brand compared with competitors
Which sources supported the interpretation
Whether the perception persisted across models and repeated runs
Whether web retrieval strengthened or weakened it
Whether the broader earned-media environment supported the same conclusion
Those dimensions can be combined into a score.
But that score should measure perception, not merely visibility.
What is an AI visibility score?
An AI visibility score generally measures how often a brand appears across a defined set of prompts or AI-generated answers.
Depending on the methodology, it may incorporate:
Brand mention frequency
Percentage of prompts containing the brand
Average answer position
Share of AI voice
Recommendation frequency
Citation frequency
Model coverage
Competitive inclusion
Search impressions
Page appearances in AI features
For example, a company may test 100 prompts across ChatGPT, Claude, Gemini, Perplexity, and Grok.
If the brand appears in 70 of those answers, it might receive a visibility score of 70%.
More sophisticated versions may assign additional weight when:
The brand appears first
The brand is recommended
The brand receives more answer space
Its website is cited
It appears across multiple models
It is named alongside fewer competitors
These additions improve the signal.
But visibility remains fundamentally a measure of presence and exposure.
It does not fully measure interpretation.
Visibility is not perception
Visibility tells you that the brand entered the answer.
Perception tells you what the answer caused a stakeholder to believe.
Consider a question such as:
Which companies lead enterprise cybersecurity?
A brand may appear in the first paragraph because it is described as:
One of the most established and widely deployed enterprise-security platforms, with strong threat intelligence and broad customer adoption.
The same brand may appear in another answer because it is described as:
A major cybersecurity provider that has faced criticism following a series of high-profile security incidents.
Both answers generate a brand mention.
Both may place the company prominently.
Both may cite authoritative sources.
One strengthens the brand.
The other creates reputational risk.
A visibility score may treat them as similarly successful.
A perception score should not.
High visibility can reflect a negative narrative
Brand prominence is not inherently positive.
Companies often become highly visible in AI-generated answers because of:
Controversies
Lawsuits
Product failures
Regulatory actions
Security incidents
Executive misconduct
Customer complaints
Financial distress
Political disputes
Workforce reductions
Failed launches
Public criticism
During a crisis, a brand may achieve its highest-ever visibility.
That does not mean its AI performance improved.
This is why visibility cannot be interpreted without brand-centric favorability.
The analysis must ask:
Does the answer increase or reduce confidence in the brand?
Is the company portrayed as responsible for the problem?
Is it shown responding effectively?
Is the issue described as resolved or ongoing?
Does the answer distinguish allegation from confirmed fact?
Does the system repeat outdated or inaccurate claims?
Is the company framed more negatively than competitors?
Without this context, a visibility score can reward reputational damage.
High visibility can reinforce the wrong category
A brand can appear frequently while being associated with a perception it is actively trying to change.
For example, a company may want to be understood as a broad enterprise platform.
AI systems may mention it in most relevant answers but continue describing it as:
A point solution
A single-product company
A small-business tool
A legacy provider
A consumer brand
A lower-cost alternative
A niche specialist
The brand is visible.
The repositioning has failed.
This matters when companies are trying to:
Move beyond a legacy product
Enter a new market
Establish category leadership
Expand into enterprise accounts
Build an AI-forward identity
Shift from services to software
Position a portfolio rather than one product
Recover from an outdated reputation
Compete on value rather than price
Visibility alone cannot reveal whether the desired category association is present.
That requires narrative and message analysis.
A citation does not make the perception favorable
Some AI visibility tools treat citations as an inherently positive outcome.
Citations are valuable, but they are not automatically beneficial.
A brand may be cited through:
A product page
An investigative article
A regulator
A lawsuit
A customer-review site
A security notice
A negative analyst report
An outdated press release
A competitor comparison
A favorable customer story
Each citation contributes different evidence.
The important questions are:
Which source was cited?
What claim did it support?
Was the source authoritative for that claim?
Was the information current?
Was the brand central to the source?
Did the source reinforce the desired or undesired narrative?
Did the citation recur across models and runs?
Did the answer accurately represent the source?
Was the brand cited as evidence of leadership, weakness, risk, or controversy?
Citation frequency measures observed source selection.
Citation interpretation reveals reputational meaning.
Search impressions do not measure AI belief
Search platforms increasingly provide reporting on how websites appear in generative AI experiences.
These reports are useful, but they primarily measure content visibility.
Google’s generative AI reporting in Search Console includes information such as:
Impressions
Pages
Countries
Devices
Performance over time
Bing Webmaster Tools’ AI Performance reporting includes:
Citation activity
Cited URLs
Citation trends
Grounding queries
Page-level activity
These tools help website owners understand whether their pages are being surfaced.
They do not directly determine:
Whether the brand was portrayed favorably
Which corporate narrative dominated
Whether the answer was accurate
Whether priority messages appeared
How the brand compared with competitors
Whether the source supported a negative conclusion
Whether the company was recommended
Whether the perception persisted without web retrieval
Whether the answer reflected the wider earned-media environment
Website visibility and brand perception are related.
They are not the same metric.
Share of AI voice has the same limitation
Share of AI voice generally measures the brand’s portion of appearances relative to competitors.
A simple version may use:
Brand appearances ÷ Total appearances across measured brands
This can help reveal competitive presence.
But it may treat all appearances as equal.
Consider two brands:
Brand A
Appears in 70% of answers
Is usually named in a long list
Receives little substantive explanation
Is rarely recommended
Has weak message pull-through
Is described neutrally
Brand B
Appears in 50% of answers
Is frequently named first
Is described as the category leader
Receives detailed favorable analysis
Is recommended for the highest-value use cases
Is supported by authoritative citations
Brand A has greater raw share of AI voice.
Brand B has stronger perception.
A useful competitive framework should therefore account for:
Answer prominence
Narrative relevance
Favorability
Recommendation
Category ownership
Message pull-through
Citation support
Source authority
Substantive discussion
Share of AI voice should remain one component, not the final measure.
Recommendation rate is also incomplete
Recommendation is often more meaningful than presence.
But recommendation alone still requires context.
An AI system may recommend a company:
As the best enterprise option
As the cheapest option
For small teams only
For users willing to accept limited functionality
As one of several acceptable choices
For a narrow use case
With significant qualifications
Despite concerns about service or reliability
These are materially different outcomes.
A recommendation score should record:
Whether the brand was recommended
For which customer or use case
Whether the recommendation was qualified
Which strengths drove it
Which weaknesses limited it
Whether a competitor was preferred overall
Which evidence supported the recommendation
Whether the recommendation persisted across repeated runs
A recommendation without interpretation can create the same false confidence as a visibility score.
General sentiment does not solve the problem
Some AI visibility systems add sentiment analysis to improve the metric.
That is better than counting mentions alone, but traditional sentiment has a major limitation: it often measures the emotional tone of the answer rather than the brand’s position within it.
For example:
Small businesses are facing severe financial pressure, but the company has helped thousands of employers reduce administrative costs and preserve jobs.
The broader subject is negative.
The brand’s position is positive.
A general sentiment system may misclassify the answer because it detects words such as:
Severe
Pressure
Costs
Job losses
Financial difficulty
Brand-centric favorability asks a more precise question:
Does this answer strengthen or weaken confidence in the brand?
That is the relevant measure for AI perception.
The prompt set can distort the score
Every AI visibility score is partly a product of the prompts selected.
A company may appear highly visible because the monitoring program includes questions that naturally favor it.
For example:
Prompts may use the company’s preferred category language.
Competitors may be omitted.
Questions may be narrowly tailored to the brand’s strongest product.
Adverse questions may not be included.
Open-ended prompts may be underrepresented.
The prompts may reflect internal messaging rather than real stakeholder language.
Similar prompts may be counted as separate evidence.
A brand name may be included in the prompt itself.
A score of 85 does not mean much without knowing what was tested.
A credible methodology should include prompt families across:
Factual questions
Category questions
Comparative questions
Evaluative questions
Adverse questions
Open-ended questions
Narrative-specific questions
Recommendation questions
The prompt set should be representative, not promotional.
Exact prompts do not represent every stakeholder question
People can ask about the same brand narrative in thousands of ways.
They may vary:
Terminology
Context
Use case
Geography
Industry
Company size
Risk tolerance
Time horizon
Competitors
Decision criteria
Follow-up questions
A communications team may monitor:
Is Apple a leader in artificial intelligence?
A customer may ask:
Which smartphone company has the most useful consumer AI features?
An investor may ask:
Is Apple’s AI strategy strengthening or weakening its competitive position?
A journalist may ask:
What evidence supports Apple’s claims about privacy-focused AI?
A developer may ask:
How open is Apple’s AI ecosystem compared with Google or Microsoft?
These prompts belong to the same broad narrative.
They require different evidence.
The narrative should therefore be the unit of analysis.
Prompts are tools used to test it.
One answer is not a score
AI-generated answers can vary between runs.
The same prompt may produce different:
Brand lists
Rankings
Recommendations
Claims
Citations
Competitive comparisons
Levels of detail
Favorability
Narrative framing
A score based on one response per prompt may create false precision.
Repeated testing is necessary to determine:
Whether the brand appears consistently
Whether the same narrative recurs
Whether favorability is stable
Whether citations repeat
Whether recommendations change
Whether one answer was an outlier
Whether the result is model-specific
Whether prompt wording changes the conclusion
A meaningful score should include a confidence rating based partly on cross-run stability.
Web-on and web-off answers measure different conditions
AI perception may change substantially when web retrieval is enabled.
With web retrieval
The answer may reflect:
Current news
Current owned content
Recent regulatory developments
Newly published research
Updated product information
Visible citations
Current competitor evidence
Without visible web retrieval
The answer may express:
More established associations
Older narratives
Persistent category perceptions
Outdated facts
Different competitive framing
A different level of confidence
These conditions should be measured separately.
A brand can have:
Strong web-on visibility but weak web-off perception
Strong web-off perception but unfavorable current retrieval
High visibility in both conditions but different narratives
Accurate web-on answers but persistent web-off inaccuracies
High citation activity without durable perception
A single visibility score may hide these differences.
The gap between web-on and web-off performance matters
The difference between retrieval conditions can reveal whether a narrative is:
Emerging
Durable
Outdated
Fragile
Dependent on current news
Vulnerable to negative retrieval
Not yet established across models
For example:
High web-on, low web-off performance
The current information environment supports the desired narrative, but the perception may not yet be durable.
Low web-on, high web-off performance
The brand has a favorable established reputation, but current evidence is weakening it.
High performance in both
The perception may be strong and resilient.
Low performance in both
The desired narrative lacks support across both current retrieval and persistent model outputs.
This can be measured as retrieval resilience.
Visibility alone does not capture it.
Visibility scores often ignore the evidence environment
An AI answer is connected to a broader body of evidence.
That environment may include:
Earned media
Owned content
Product documentation
Regulatory records
Customer evidence
Research
Expert commentary
Social and community sources
Business directories
Public filings
Competitor content
A visibility score may record that the brand appeared.
It may not explain why.
For example, a competitor may outperform the brand because it has:
Clearer category-definition pages
Stronger product documentation
More authoritative earned media
Better customer evidence
Greater brand prominence
More current information
More consistent executive messaging
Stronger independent corroboration
Higher-quality comparison coverage
Better crawl accessibility
Without evidence analysis, the score identifies the symptom but not the cause.
Earned media should be part of perception measurement
Earned media helps establish how independent sources understand the brand.
It can:
Validate claims
Challenge positioning
Define categories
Compare competitors
Document outcomes
Build executive authority
Establish market significance
Surface reputational risk
Create durable associations
A perception framework should examine:
Which narratives dominate coverage
Whether the brand is central or incidental
How the brand is positioned
Which sources have authority
Which claims recur
Whether the coverage is independently reported
Which messages pull through
Whether sources converge on the same interpretation
Which competitor narratives are stronger
Whether the coverage supports or conflicts with AI answers
Raw AI visibility ignores much of this evidence.
Visible citations are only one evidence layer
Visible citations show which sources were presented with a particular answer.
They are valuable and directly observable.
But they should not automatically be treated as a complete explanation of every factor involved in producing the response.
AI companies do not publish a full formula for every answer.
A complete source analysis should distinguish among:
Observed citation behavior
Which URLs and domains visibly appear across tested answers?
Likely citation influence
Which sources are best positioned to shape future retrieval and citations based on:
Authority
Relevance
Factual specificity
Brand prominence
Freshness
Independence
Accessibility
Repetition
Narrative alignment
The broader evidence environment
Which earned, owned, expert, regulatory, customer, social, and research sources collectively support the narrative?
A visibility score may capture a portion of the first layer.
A perception framework should examine all three.
What AI visibility scores are good for
AI visibility scores still have value.
They can help answer:
Is the brand appearing at all?
How often does it appear?
Which models include it?
Which competitors appear more often?
Is the brand being recommended?
Is its website cited?
Are certain pages surfacing?
Is visibility changing over time?
Is the brand absent from important category questions?
Has a campaign increased exposure?
These are legitimate measurement questions.
Visibility can be particularly useful for:
Establishing a baseline
Tracking broad category presence
Identifying competitive gaps
Finding prompts that require deeper analysis
Monitoring changes after a launch
Measuring website citation activity
Detecting emerging inclusion
Prioritizing narratives for investigation
The problem is not visibility measurement.
The problem is treating visibility as the complete definition of AI performance.
The better alternative: an LLM Perception Score
A defensible LLM Perception Score should measure the quality, strength, accuracy, and durability of the brand’s interpretation across AI systems.
It can combine:
Narrative presence
Brand prominence
Brand-centric favorability
Message pull-through
Factual accuracy
Competitive position
Recommendation quality
Citation consistency
Citation quality
Source authority
Cross-model consistency
Cross-run stability
Retrieval resilience
Earned-media evidence strength
Narrative trajectory
This score answers a different question.
An AI visibility score asks:
How often did the brand appear?
An LLM Perception Score asks:
How strong, favorable, accurate, consistent, competitive, and well-supported is the brand’s perception?
That is a more valuable executive measure.
A perception score should include visibility
Visibility should not be discarded.
It should be incorporated as one component of the broader score.
For example:
| Component | What it measures |
|---|---|
| Narrative presence | Whether the brand appears within the priority narrative |
| Prominence | How central the brand is to the answer |
| Favorability | Whether the answer strengthens or weakens confidence |
| Message pull-through | Whether priority messages survive synthesis |
| Accuracy | Whether material claims are correct and current |
| Competitive position | How the brand compares with alternatives |
| Recommendation quality | Whether and why the brand is recommended |
| Citation performance | Which sources recur and what they support |
| Evidence strength | Authority, independence, specificity, and corroboration |
| Model consistency | Stability across AI systems |
| Run stability | Stability across repeated responses |
| Retrieval resilience | Performance with web retrieval on and off |
| Earned-media environment | Strength of the wider independent evidence |
| Narrative trajectory | Whether perception is strengthening or weakening |
Visibility matters.
It simply should not dominate everything else.
How to score narrative presence
Narrative presence should distinguish between levels of inclusion.
A useful rubric may classify a result as:
Primary
The brand is presented as a leading example or central subject.
Substantive
The brand is meaningfully connected to the narrative.
Incidental
The brand appears but contributes little to the answer.
Absent
The brand does not appear.
This prevents a passing mention from receiving the same credit as category leadership.
How to score prominence
Prominence can incorporate:
Appearance in the opening
Position in a list
Amount of substantive discussion
Inclusion in headings
Presence in the conclusion
Use as the primary example
Share of answer space
Whether the brand defines the category
A brand should receive more credit when it frames the answer than when it appears as an afterthought.
How to score brand-centric favorability
Favorability should assess the brand’s position rather than general emotional tone.
A practical rubric may classify answers as:
Strongly positive
Positive
Neutral
Mixed
Negative
Strongly negative
The classification should consider whether the answer:
Builds confidence
Validates leadership
Reinforces differentiation
Raises concerns
Describes failures
Introduces qualifications
Associates the company with risk
Shows effective response to a difficult issue
The reason behind the classification should remain visible.
How to score message pull-through
For each priority message, classify whether it is:
Explicit
Implied
Absent
Contradicted
A brand may appear frequently while its most important message never appears.
For example, a company may want to be understood as more than its traditional product.
If AI systems repeatedly name the company but continue describing it only through the legacy category, visibility is high and message pull-through is low.
That distinction is central to perception measurement.
How to score factual accuracy
Accuracy should be evaluated at the claim level.
Classifications may include:
Accurate
Incomplete
Outdated
Unsupported
Incorrect
A high-visibility answer containing material inaccuracies should not receive a strong perception score.
Accuracy is especially important for:
Executive leadership
Product availability
Pricing
Regulatory status
Company ownership
Transactions
Market data
Corporate history
Security events
Litigation
Product capabilities
A score that ignores accuracy can reward confident misinformation.
How to score competitive position
Competitive analysis should examine:
Which brands appear
Which brand appears first
Which company is framed as the leader
Which differentiators are assigned
Which weaknesses are attached
Which use cases each brand owns
Whether the brand is recommended
Whether competitors have stronger evidence
Whether the category definition favors one company
A brand may have strong standalone visibility but weak comparative perception.
That difference should affect the score.
How to score citation performance
Citation performance should go beyond counting URLs.
It may incorporate:
Citation presence
URL recurrence
Domain recurrence
Source authority
Source freshness
Brand prominence in the source
Claim relevance
Independent reporting
Citation-to-claim support
Cross-model recurrence
Cross-run recurrence
Whether the source supports the desired or undesired narrative
A cited corporate page may be highly useful for product specifications.
A cited regulator may be highly authoritative for an enforcement action.
A cited review site may be relevant to usability.
The source must be evaluated in relation to the claim.
How to score source authority
Authority is contextual.
Relevant factors include:
Institutional responsibility
Subject-matter expertise
Editorial standards
Original reporting
Firsthand evidence
Independent validation
Research quality
Transparency
Factual specificity
Current relevance
The largest publication is not always the strongest source.
A specialist journal may carry greater authority for a technical claim.
A regulator may carry greater authority for an approval.
A company page may carry greater authority for its current pricing.
The scoring rubric should recognize this.
How to score cross-model consistency
ChatGPT, Claude, Gemini, Perplexity, and Grok may produce different interpretations.
Cross-model consistency should measure whether the underlying conclusion remains stable.
Possible classifications include:
Highly consistent
Generally consistent
Mixed
Contradictory
Insufficient evidence
The goal is not identical wording.
It is a stable underlying perception.
A high average score combined with severe model disagreement should carry less confidence than a similar score supported by broad agreement.
How to score cross-run stability
Repeated runs reveal whether the result persists.
A stable perception should maintain similar:
Narrative presence
Favorability
Message pull-through
Competitive position
Recommendation
Factual claims
Citation patterns
One favorable answer should not receive the same weight as a narrative that recurs consistently.
Run stability is therefore both a performance input and a confidence input.
How to score retrieval resilience
Retrieval resilience measures whether perception holds when web retrieval is turned on and off.
A strong result may show:
Accurate perception in both conditions
Consistent favorability
Stable message pull-through
Limited narrative fragmentation
Current retrieval reinforcing the established perception
A weak result may show:
Favorable web-off answers but negative current retrieval
Accurate web-on answers but persistent outdated web-off beliefs
Large changes in competitive position
Current citations contradicting the desired narrative
Extreme sensitivity to retrieval conditions
This distinction is invisible in a single visibility score.
How to score earned-media evidence strength
Earned-media evidence strength should measure how well independent coverage supports the narrative.
Relevant dimensions include:
Source authority
Brand prominence
Brand-centric favorability
Narrative relevance
Factual specificity
Independent corroboration
Original reporting
Freshness
Narrative momentum
Competitive framing
Consistency across sources
Conflicting evidence
Coverage volume may contribute context.
It should not automatically increase the score.
One authoritative, substantive story may provide stronger evidence than dozens of passing mentions or syndicated articles.
How to score narrative trajectory
A perception score should not be static.
Narrative trajectory evaluates whether the brand’s perception is:
Strengthening
Weakening
Correcting
Fragmenting
Hardening
Fading
Remaining stable
This helps communications teams understand whether recent activity is changing the information environment.
A score of 70 that is strengthening may represent a different strategic situation from a score of 70 that is deteriorating.
Build the score from an explicit rubric
A composite score is valuable only when the methodology is understandable.
The rubric should define:
Each component
The scoring scale
The evidence required
The weighting
How missing data is handled
How web-on and web-off results are combined
How repeated runs affect confidence
How models are weighted
How earned-media evidence contributes
How adverse narratives are handled
How accuracy affects the result
How score changes are calculated
An unexplained number creates false authority.
A transparent score creates an executive summary that can be traced back to evidence.
Add a confidence rating
Performance and confidence should be reported separately.
Two brands can receive the same LLM Perception Score while the evidence behind the scores differs significantly.
A confidence rating may account for:
Number of models tested
Number of prompt families
Number of repeated runs
Cross-model agreement
Cross-run stability
Citation recurrence
Source quality
Evidence freshness
Earned-media corpus size
Retrieval-condition coverage
Longitudinal consistency
For example:
Score: 82
Confidence: High
This may indicate favorable perception supported across five models, repeated runs, strong citations, and authoritative earned media.
Score: 82
Confidence: Low
This may indicate a favorable result based on limited prompts, one model, or unstable outputs.
The score communicates performance.
Confidence communicates how strongly the evidence supports it.
A sample LLM Perception Score architecture
A company could organize the overall score into component categories such as:
| Score component | Example measures |
|---|---|
| Perception quality | Favorability, accuracy, overall framing |
| Narrative strength | Presence, prominence, message pull-through |
| Competitive position | Leadership, differentiation, recommendation |
| Evidence strength | Authority, specificity, independence, corroboration |
| Citation intelligence | Citation quality, recurrence, claim support |
| Model consistency | Cross-model and cross-run stability |
| Retrieval resilience | Web-on and web-off alignment |
| Earned-media environment | Independent narrative support |
| Narrative trajectory | Direction and durability over time |
The exact weighting should reflect the purpose of the analysis.
For example:
Corporate reputation
May place greater weight on:
Favorability
Accuracy
Source authority
Cross-model consistency
Earned-media evidence
Narrative durability
Product launch
May place greater weight on:
Narrative presence
Message pull-through
Product accuracy
Competitive position
Retrieval freshness
Citation activity
Earned-media momentum
Crisis analysis
May place greater weight on:
Factual accuracy
Negative narrative prevalence
Source authority
Current retrieval
Narrative trajectory
Correction persistence
Cross-model consistency
The score should remain comparable over time while allowing the framework to reflect the business question.
Visibility should remain visible within the score
A broader perception score should not hide the underlying visibility data.
A useful report may include:
Raw brand presence
Primary prominence
Share of AI voice
Recommendation frequency
Citation presence
Overall LLM Perception Score
Confidence rating
Narrative-level scores
Web-on score
Web-off score
Earned-evidence score
Competitive-position score
Message-pull-through score
This allows teams to see both exposure and interpretation.
For example:
| Metric | Result |
|---|---|
| Raw visibility | 84% |
| Primary prominence | 43% |
| Positive favorability | 51% |
| Message pull-through | 38% |
| Factual accuracy | 93% |
| Recommendation rate | 47% |
| High-authority citation share | 42% |
| LLM Perception Score | 61 |
| Confidence | High |
This scorecard reveals that the brand appears often but is not converting that exposure into strong perception.
A high visibility score can coexist with a low perception score
Consider a hypothetical enterprise-software company.
The company appears in 88% of monitored answers.
A deeper analysis finds:
It is usually named after two competitors.
It is described as a legacy platform.
Its AI capabilities are framed as add-ons.
Priority messages appear in only 22% of answers.
Current citations include several older product articles.
Competitors receive stronger enterprise recommendations.
Favorability is neutral to mixed.
Web retrieval makes the comparison less favorable.
Earned media focuses heavily on implementation complexity.
The result is stable across models.
The company may have:
AI Visibility Score: 88
LLM Perception Score: 54
Confidence: High
The visibility score says the company is widely present.
The perception score says the current narrative is strategically weak.
Both numbers are useful.
Only one captures the reputational outcome.
A lower visibility score can coexist with a strong perception score
Now consider a new category challenger.
The company appears in only 42% of relevant answers.
When it appears:
It is described as innovative.
It receives strong message pull-through.
It is recommended for sophisticated use cases.
Its claims are supported by authoritative research.
It has favorable independent coverage.
Its citations are current and relevant.
Cross-run stability is strong.
Web-on perception is highly favorable.
Web-off perception is beginning to emerge.
The company may have:
AI Visibility Score: 42
LLM Perception Score: 79
Confidence: Moderate
The strategic issue is not perception quality.
It is narrative reach and durability.
The company should amplify strong evidence rather than rebuild its positioning from scratch.
Visibility alone would miss this distinction.
The score should diagnose the evidence gap
The purpose of measurement is not simply to produce a number.
It should reveal why the number exists.
Common evidence gaps include:
Presence gap
The brand is missing from relevant answers.
Prominence gap
The brand appears but is not central.
Favorability gap
The brand is visible but portrayed negatively or with significant qualifications.
Message gap
Priority claims do not survive synthesis.
Accuracy gap
Answers contain outdated or incorrect information.
Competitive gap
Competitors own the category or preferred use case.
Citation gap
The brand lacks recurring authoritative source support.
Authority gap
Most supporting evidence is company-authored or low quality.
Retrieval gap
Strong information exists but is difficult to access.
Durability gap
The narrative appears only under narrow prompts, one model, or web-on conditions.
Earned-evidence gap
The desired perception lacks independent corroboration.
Different gaps require different actions.
Connect the score to communications action
A perception score should help determine what the company should do next.
Amplify
Use when the desired perception is favorable but insufficiently prominent.
Clarify
Use when the right narrative appears but important distinctions are misunderstood.
Counter
Use when inaccurate or incomplete narratives are gaining strength.
Canonicalize
Use when important facts lack a clear permanent owned source.
Create
Use when the evidence required to support the desired narrative does not yet exist.
Validate
Use when a company claim lacks credible independent support.
Correct
Use when outdated or inaccurate information remains accessible.
Consolidate
Use when fragmented owned content creates conflicting interpretations.
Monitor
Use when the narrative is emerging but not yet stable.
The score should point to the evidence gap.
The evidence gap should determine the action.
Measure change over time
A one-time score is a baseline.
The strategic value comes from longitudinal measurement.
Track:
Overall LLM Perception Score
Raw visibility
Narrative-level scores
Favorability
Message pull-through
Competitive position
Citation quality
Earned-evidence strength
Web-on and web-off differences
Cross-model consistency
Cross-run stability
Narrative trajectory
Confidence
Then connect changes to:
Product launches
Earned-media campaigns
Research releases
Executive announcements
Issues and crises
Website updates
Category campaigns
Competitor activity
Regulatory developments
Customer evidence
This allows communications teams to show whether the information environment changed.
Report the conclusion before the metrics
Executives should not have to interpret hundreds of prompt results.
A strong report should begin with:
What AI systems currently believe
State the dominant perception directly.
How strong the perception is
Provide the LLM Perception Score and confidence rating.
What is driving the score
Identify the strongest positive and negative components.
How the brand compares with competitors
Explain leadership, differentiation, recommendation, and category position.
What changed
Describe narrative strengthening, weakening, correction, or fragmentation.
Why it changed
Connect the movement to sources, claims, earned media, owned content, and current events.
What should happen next
Provide evidence-based communications actions.
Visibility metrics should support this interpretation.
They should not replace it.
Questions to ask an AI visibility vendor
Before relying on an AI visibility score, ask:
What exactly does the score measure?
Does it measure presence or perception?
Does it evaluate brand-centric favorability?
Does it measure message pull-through?
Does it assess factual accuracy?
Does it distinguish primary prominence from passing mentions?
Does it measure competitive positioning?
Does it classify recommendation quality?
Does it use repeated runs?
Does it test multiple models?
Does it separate web-on and web-off conditions?
Does it map citations to claims?
Does it evaluate source authority?
Does it connect answers to the broader earned-media environment?
Does it measure narrative durability?
Does it include a confidence rating?
Can the score be traced to the underlying evidence?
Can it explain why the result exists?
Can it identify what should change?
Can it demonstrate movement over time?
A score that cannot answer these questions may still be useful.
It should be described accurately as a visibility metric.
Common mistakes with AI visibility scores
Treating all appearances as positive
A brand can be visible because of reputational risk.
Treating all mentions equally
A passing mention is not equivalent to category leadership.
Ignoring brand-centric favorability
General sentiment may misread how the brand itself is positioned.
Ignoring accuracy
A highly visible incorrect answer is not success.
Treating citations as inherently favorable
A citation may support criticism, controversy, or outdated information.
Using one response per prompt
Single-run results may be unstable.
Testing only favorable prompts
The score may reflect the assumptions of the testing program.
Ignoring web retrieval conditions
Web-on and web-off perception may differ materially.
Ignoring model differences
A combined score can conceal major disagreement.
Ignoring earned media
Visible AI outputs are connected to a larger evidence environment.
Hiding the rubric
An unexplained score creates false precision.
Reporting performance without confidence
A high score based on weak evidence should not be treated as durable.
Optimizing the score rather than the narrative
The goal is not to manipulate tested prompts.
It is to strengthen the underlying perception.
What good AI perception measurement should reveal
A complete measurement program should answer:
Is the brand visible?
How prominent is it?
Which narrative defines it?
Is the portrayal favorable?
Which messages pull through?
Are the claims accurate?
How does the brand compare with competitors?
Is it recommended?
Why is it recommended or rejected?
Which sources are cited?
What do those sources support?
How authoritative are they?
Does the perception persist across models?
Does it persist across repeated runs?
Does web retrieval improve or weaken it?
Does earned media support the same conclusion?
Is the narrative strengthening or fading?
How confident should leadership be in the result?
What evidence gap is driving weakness?
What communications action should follow?
A visibility score answers the first question.
An LLM Perception Score should answer the complete set.
The central lesson
AI visibility matters.
Brands need to know whether they appear in the answers their stakeholders receive.
But visibility is not the final outcome.
A company can be highly visible and poorly perceived.
It can be cited often and associated with risk.
It can lead share of AI voice while losing the category narrative.
It can appear across every model while its priority messages remain absent.
It can receive strong current visibility while an outdated perception remains durable without web retrieval.
The better measure is not whether the brand appeared.
It is whether AI systems consistently reproduce an accurate, favorable, differentiated, and well-supported interpretation of the brand.
That can be measured.
The underlying attributes can be defined in a rubric.
The rubric can produce a score.
The score can be compared across narratives, models, competitors, retrieval conditions, and time.
An AI visibility score measures exposure.
An LLM Perception Score measures what the exposure means.
For communications leaders, that is the outcome that matters.
Frequently asked questions
An AI visibility score measures whether or how frequently a brand appears across a defined set of AI-generated answers. It may also incorporate answer position, share of AI voice, recommendation frequency, citations, or model coverage.
Visibility measures presence. AI brand perception measures how the brand is interpreted, including its narrative, prominence, favorability, accuracy, message pull-through, competitive position, recommendations, citations, and supporting evidence.
Yes. A company may receive high visibility because of controversy, regulatory action, product failure, customer criticism, or another unfavorable narrative. Visibility must be interpreted alongside brand-centric favorability.
They are a useful observable signal, but not automatically a positive one. A citation may support a favorable claim, a negative claim, an outdated fact, a regulatory action, or a customer complaint. Citation quality and narrative contribution should be evaluated.
Yes, as one component of competitive measurement. It becomes misleading when every appearance is treated equally and prominence, favorability, recommendation, narrative relevance, and citation support are ignored.
Recommendation can be more meaningful, but it still requires interpretation. A brand may be recommended only for a narrow segment, as a budget option, or with substantial qualifications.
Brand-centric favorability measures whether the answer strengthens or weakens confidence in the brand itself. It differs from general sentiment, which may focus on the emotional tone of the broader subject.
An LLM Perception Score is a composite measure of the quality, strength, accuracy, consistency, and durability of a brand’s perception across AI systems. It can combine visibility, prominence, favorability, message pull-through, accuracy, competitive position, recommendation, citation quality, source authority, cross-model consistency, cross-run stability, retrieval resilience, earned-media evidence, and narrative trajectory.
Yes. Visibility and narrative presence are important components. They should not be treated as the complete score.
It can be synthesized into an overall score when the rubric is transparent and the underlying components remain visible. The score should act as an executive summary, not a substitute for evidence.
The same score may be supported by very different amounts and qualities of evidence. Confidence reflects sample depth, repeated-run stability, cross-model agreement, citation recurrence, source quality, retrieval coverage, and earned-media evidence.
Web-on testing shows how current accessible evidence changes the answer and which sources are surfaced. Web-off testing shows the perception expressed without visible current retrieval under the tested conditions. The difference reveals retrieval resilience and narrative durability.
Earned media provides independent evidence about the brand. It can validate or challenge company claims, establish significance, compare competitors, document outcomes, and create recurring narratives that shape both human and AI perception.
No. Search impressions measure whether pages appeared within relevant search or generative AI features. They do not directly measure how the brand was characterized or whether the perception was favorable, accurate, or strategically useful.
Not by themselves. Citation counts show observed activity. They do not automatically establish source authority, answer placement, narrative value, or influence across all contexts.
Report: • Overall LLM Perception Score • Confidence rating • Raw visibility • Narrative presence • Brand prominence • Brand-centric favorability • Message pull-through • Accuracy • Competitive position • Recommendation • Citation intelligence • Web-on and web-off results • Earned-media evidence strength • Narrative trajectory
Yes. A brand can improve the score by strengthening the information environment through clearer owned content, stronger factual evidence, authoritative earned media, independent corroboration, current sources, better message pull-through, corrected inaccuracies, and more consistent narrative support.
Yes. A newer or emerging brand may appear less frequently but be described very favorably and supported by strong evidence when it does appear. The strategic priority may be amplification rather than repositioning.
Yes. A well-known company may appear frequently while being characterized as outdated, risky, undifferentiated, or weaker than competitors.
They can confuse being mentioned with being understood favorably. The brand’s interpretation, not its mere appearance, is the real reputational outcome.