Measuring AI brand perception is not the same as checking whether a company appears in a handful of prompts.
A brand can appear frequently and still be misunderstood.
It can receive a high visibility score while unfavorable narratives dominate the answer. It can be cited often but cited in connection with controversy, weakness, or outdated information. It can perform well in one model and poorly in another. It can also appear in an answer without the messages the company considers most important.
True AI brand perception measurement must answer a more consequential question:
What do AI systems believe about the brand, why do they appear to believe it, and how consistently is that perception reproduced?
That requires measuring more than mentions.
A complete framework should evaluate:
Narrative presence
Brand prominence
Brand-centric favorability
Message pull-through
Factual accuracy
Citation consistency
Source authority
Competitive position
Cross-model variance
Cross-run stability
Narrative drift
Likely versus observed citation influence
The strength of the underlying earned-media evidence environment
These dimensions can then be combined through a transparent rubric to produce an overall LLM Perception Score.
The score provides a clear executive signal. The underlying dimensions explain what is driving it.
What is AI brand perception?
AI brand perception is the interpretation of a company, product, executive, or issue expressed through an AI-generated answer.
It includes the facts, narratives, associations, judgments, sources, and comparisons the system uses to describe the brand.
For example, an AI system may describe a company as:
An established market leader
An emerging challenger
A low-cost alternative
A product innovator
A trusted institutional brand
A company facing regulatory pressure
A business struggling to differentiate itself
A category pioneer
A controversial employer
A reliable but outdated incumbent
These are not merely mentions. They are compressed interpretations.
An AI answer may synthesize news coverage, company claims, product documentation, reviews, public records, research, and other available information into a few sentences. Those sentences can shape how customers, journalists, investors, employees, policymakers, and other stakeholders understand the company.
Measuring AI brand perception means systematically evaluating those interpretations.
Why prompt visibility is not enough
Many AI-monitoring programs begin with a list of prompts such as:
What are the best companies in this category?
Who are the leaders in this market?
What products should I consider?
What is this company known for?
Is this brand trustworthy?
The system records whether the brand appears and may convert the result into a visibility score.
That can be useful, but it is incomplete.
A visibility score generally cannot tell you:
What the system believes about the brand
Whether the brand is portrayed favorably
Which narratives dominate the answer
Whether priority messages are understood
Why a competitor is preferred
Which sources support the interpretation
Whether cited evidence is current
Whether the answer is stable across repeated runs
Whether perception changes when web retrieval is enabled
Whether perception is changing over time
What information gap is producing the result
Consider two brands that appear in 80% of tested answers.
Brand A is described as the established leader with the strongest enterprise capabilities.
Brand B appears just as often but is described as a less expensive alternative with limited functionality.
Their visibility is identical. Their perception is not.
The strategic value lies in understanding and scoring the difference.
Start with narratives, not an endless prompt list
No company can know every question a stakeholder may ask an AI system.
People can express the same intent in thousands of ways. They may include different context, constraints, comparisons, assumptions, and terminology.
A measurement strategy built around exact prompts therefore risks optimizing for a small set of guessed questions rather than understanding the broader perception of the brand.
Start with the narratives that matter most to the business.
These may include:
Corporate reputation
Product leadership
Innovation
Artificial intelligence
Customer trust
Security
Pricing
Sustainability
Executive leadership
Workplace culture
Regulatory issues
Financial stability
Market expansion
A major product launch
A controversy or crisis
A competitive differentiator
Each narrative can then be tested through a representative set of questions.
For example, a company measuring its AI innovation narrative might test:
How is the company using AI?
Is the company considered an AI leader?
What differentiates its AI strategy?
Which companies are leading AI adoption in its industry?
How does its AI offering compare with competitors?
What concerns have been raised about its use of AI?
What evidence supports its AI claims?
The prompts are measurement instruments.
The narrative is the unit of analysis.
Measure the answer, not only the brand mention
Every response should be evaluated as a complete interpretation.
At minimum, record:
Whether the brand appears
How prominently it appears
Which narrative is expressed
Whether the portrayal is favorable
Which priority messages appear
Which competitors are included
How the brands are compared
Which sources are cited
Whether the claims are accurate
Whether web retrieval was enabled
How the answer differs from other runs and models
How the answer relates to the broader earned-media evidence environment
This turns AI measurement from a mention-counting exercise into perception analysis.
The three layers of AI brand perception analysis
A complete analysis should examine the brand through three connected layers.
1. LLM perception without web retrieval
Testing without live web retrieval helps evaluate the brand perception expressed by the model under the tested conditions.
This can reveal:
Established brand associations
Durable narrative patterns
Outdated beliefs
Persistent misconceptions
Differences among models
Messages that appear without current web evidence
Competitive positions that have become embedded in recurring answers
Because AI companies do not disclose every factor involved in producing an answer, this should be treated as an observed representation of the model’s current perception, not a complete view into its internal knowledge or training process.
2. LLM perception with web retrieval and citation analysis
Testing with web retrieval shows how current accessible information affects the answer.
This analysis can measure:
Whether retrieval changes the brand’s perception
Whether current information corrects outdated beliefs
Which claims become more or less prominent
Which sources are visibly cited
Which sources recur across repeated runs
Whether cited sources actually support the associated claims
Whether retrieval strengthens or weakens the intended narrative
How citation behavior differs across models
Whether the same sources consistently shape the answer
This layer connects perception with observable source selection.
3. Earned-media narrative analysis
Earned-media analysis evaluates the broader coverage environment surrounding the narrative.
It can measure:
Coverage volume
Brand prominence
Brand-centric sentiment
Narrative formation
Narrative momentum
Message pull-through
Source authority
Factual specificity
Independent corroboration
Conflicting evidence
Competitive positioning
The claims most likely to become durable AI beliefs
This layer is important because visible citations represent only the sources surfaced in tested answers.
The broader earned-media environment reveals the complete body of current evidence shaping human perception and potentially influencing future AI retrieval, recurring citations, and longer-term model perception.
Together, these three layers provide a more complete basis for measuring and scoring AI perception than prompt visibility alone.
The core AI brand perception metrics
A strong measurement framework should separate the dimensions that ultimately feed the overall score.
1. Narrative presence
Narrative presence measures whether the brand is associated with the strategic idea being evaluated.
This is more specific than simple brand visibility.
A company may appear in an answer about enterprise software without being associated with the narrative it wants to own, such as artificial intelligence, financial transformation, security, or ease of use.
Narrative presence can be classified as:
Primary: The brand is presented as a leading example of the narrative.
Substantive: The brand is meaningfully connected to the narrative but is not the main example.
Incidental: The brand appears without a meaningful narrative association.
Absent: The brand does not appear.
For example, if an AI system lists a company among ten payroll vendors but says nothing about its small-business expertise, the brand is visible but the intended narrative is absent.
2. Brand prominence
Prominence measures how central the brand is to the answer.
Relevant signals include:
Whether the brand appears in the opening answer
Its position in a list
The amount of substantive discussion
Whether it receives its own section
Whether it is named in the conclusion
Whether its evidence is used to define the category
Whether it is treated as a primary or secondary example
A brand mentioned at the end of a long answer does not have the same influence as a brand used to frame the entire response.
Prominence should be measured separately from presence.
3. Brand-centric favorability
Traditional sentiment analysis often measures the general emotional tone of a passage.
That is not sufficient for brand perception.
An article or AI answer can contain negative language while positioning the brand favorably. For example, a company may be praised for helping customers through an economic downturn. The subject matter is negative, but the brand’s position is positive.
Brand-centric favorability evaluates how the brand itself is positioned.
A practical classification is:
Positive: The answer strengthens confidence in the brand or associates it with favorable qualities, performance, or outcomes.
Neutral: The answer presents primarily factual information without a meaningful positive or negative implication.
Mixed: The answer includes material strengths and weaknesses.
Negative: The answer weakens confidence in the brand or associates it with unfavorable qualities, risks, failures, or outcomes.
The analysis should explain why the classification was assigned.
A score without the underlying interpretation is difficult to act on.
4. Message pull-through
Message pull-through measures whether the ideas the brand wants stakeholders to understand actually appear in the answer.
Priority messages might include:
The company serves more than one product category
Its technology is designed for enterprise use
It is a leader in a specific scientific field
Its product reduces a defined customer problem
It is expanding beyond its legacy business
Its approach differs fundamentally from a competitor’s
It has a particular security, quality, or research advantage
It serves a specific customer segment
For each message, classify the result as:
Explicit: The answer clearly states the message.
Implied: The answer supports the message without stating it directly.
Absent: The message does not appear.
Contradicted: The answer presents an incompatible perception.
Message pull-through reveals whether the information environment is communicating the intended narrative, not merely generating brand mentions.
5. Factual accuracy
AI systems can produce inaccurate, outdated, incomplete, or unsupported claims.
Accuracy analysis should identify:
Incorrect facts
Outdated facts
Fabricated details
Misattributed claims
Conflated products or companies
Missing qualifications
Incorrect executive roles
Outdated pricing or policies
Misrepresented regulatory status
Unsupported superlatives
Incorrect competitive comparisons
Not every omission is an error.
The evaluator should distinguish between:
Incorrect: The statement conflicts with reliable evidence.
Outdated: The statement was previously true but no longer reflects the current state.
Incomplete: The statement is directionally accurate but omits material context.
Unsupported: The claim lacks identifiable evidence.
Accurate: The statement is supported by current evidence.
Accuracy should be evaluated at the claim level, not only at the answer level.
6. Citation presence
Citation presence measures whether the answer visibly identifies sources.
Depending on the platform and mode, a response may include:
Inline citations
Footnotes
Linked publisher names
Source cards
A separate source panel
References at the end
No visible citations
Citation presence matters because it allows the evaluator to inspect the evidence presented with the answer.
But a cited answer is not automatically accurate, and an uncited answer does not provide a complete account of the information involved in producing it.
Treat visible citations as observable evidence, not a complete explanation of the system’s internal process.
7. Citation consistency
Citation consistency measures how often the same sources recur across repeated answers to related questions.
One citation in one response may be incidental.
Repeated citation of the same source across several runs is a stronger pattern.
Track:
Domain citation frequency
URL citation frequency
Citation frequency by narrative
Citation frequency by model
Citation frequency by prompt variation
Citation frequency over time
Citation position within the answer
Whether the source directly supports the associated claim
Citation consistency is especially important because AI answers can vary from one run to the next.
The objective is not to identify one source that appeared once. It is to understand which sources repeatedly surface as evidence.
8. Source authority
Not all citations carry equal weight.
A source-authority assessment should consider:
Expertise on the subject
Editorial or institutional standards
Original reporting
Firsthand access
Primary evidence
Independence
Factual specificity
Brand prominence
Freshness
Relevance to the claim
A corporate product page may be highly authoritative for product specifications but weak evidence for an independent claim of market leadership.
A regulator may be authoritative for an approval or enforcement action.
A specialist publication may be more authoritative for a technical question than a larger general-interest publication.
Source authority must be assessed in relation to the claim.
9. Competitive position
AI systems frequently answer brand questions through comparison.
Competitive-position analysis should examine:
Which competitors appear
Which brand appears first
Which brand is described as the leader
Which differentiators are assigned to each company
Which weaknesses are attached to each company
Which brand is recommended for which use case
Which evidence supports the comparison
Whether the category definition favors one competitor
Whether the brand is framed as an incumbent, challenger, specialist, or alternative
A brand may have strong standalone perception but weak comparative positioning.
For example, an AI system may describe a product positively while consistently recommending a competitor for enterprise customers. That distinction would be lost in a basic sentiment or visibility score.
10. Cross-model consistency
ChatGPT, Claude, Gemini, Perplexity, and Grok may produce different answers to the same question.
Cross-model consistency measures whether the brand’s core perception is stable across systems.
Classify the result as:
Highly consistent: The systems express substantially the same narrative and conclusion.
Generally consistent: The systems agree on the central perception but differ in detail.
Mixed: The systems produce materially different interpretations.
Contradictory: The systems reach opposing conclusions.
Insufficient evidence: Too few systems or responses provide a usable answer.
Cross-model disagreement is strategically important.
It may indicate:
Different source retrieval
Different indexes or search partners
Different model behavior
Uneven coverage across information ecosystems
Conflicting evidence about the brand
A narrative that has not yet become durable
Greater sensitivity to prompt wording
The goal is not necessarily identical wording.
The goal is a consistent underlying perception.
11. Cross-run stability
The same model can answer the same question differently across repeated runs.
Cross-run stability measures whether the conclusion persists.
For each test, track:
Brand inclusion
Narrative classification
Favorability
Message pull-through
Competitor inclusion
Recommendation
Citations
Factual claims
A perception that appears once but disappears in most repeated runs is not yet stable.
A narrative that persists across repeated runs, prompt variations, models, and retrieval conditions is more likely to represent a durable AI belief.
12. Narrative drift
Narrative drift measures how AI perception changes over time.
A brand may gradually shift from being described as:
A single-product company to a broader platform
A market leader to a legacy incumbent
An emerging challenger to an established competitor
A trusted company to a controversial one
A traditional business to an AI innovator
A consumer brand to an enterprise provider
Drift can be:
Positive: Perception moves toward the desired narrative.
Negative: Perception moves away from the desired narrative.
Corrective: An outdated or inaccurate belief weakens.
Fragmenting: Models become less consistent.
Hardening: The same narrative becomes increasingly stable.
Fading: A narrative appears less frequently or prominently.
Narrative drift should be measured against a fixed baseline and tied to dated evidence.
Without historical comparison, a team cannot tell whether perception is improving, deteriorating, or remaining unchanged.
13. Observed citation behavior
Observed citation behavior measures the sources that visibly appear in actual AI answers.
This is the most directly measurable source signal.
It can be evaluated through:
Repeated model runs
Citation frequency
Citation-to-claim mapping
URL recurrence
Domain recurrence
Cross-model citation patterns
Citation changes over time
Observed citation behavior answers:
Which sources are these systems visibly selecting when they answer questions about this narrative?
It should not be interpreted as a complete account of every factor that contributed to the response.
14. Likely citation influence
Likely citation influence assesses which sources are best positioned to shape future retrieval and citation behavior.
Evaluate each source based on:
Source authority
Direct relevance
Factual specificity
Brand prominence
Freshness
Independence
Repetition across credible sources
Narrative alignment
Accessibility
This analysis is inferential.
It does not claim access to a model’s internal weighting or training process. It identifies the sources most likely to support future answers based on observable evidence.
The distinction matters:
Observed citation behavior measures what visibly happened.
Likely citation influence assesses which sources are best positioned to shape future answers.
Both are necessary.
Observed citations alone may miss strong sources that have not yet surfaced. Likely influence alone may overestimate sources that appear authoritative but are rarely retrieved.
15. Earned-media evidence strength
Earned-media evidence strength measures the quality of the broader information environment surrounding the narrative.
It should account for:
Source authority
Brand prominence
Brand-centric sentiment
Factual specificity
Independent corroboration
Narrative consistency
Coverage freshness
Narrative momentum
Competitive framing
Conflicting evidence
The degree to which the same perception recurs across credible sources
Raw coverage volume should not automatically increase the score.
A large volume of low-authority, syndicated, or passing mentions may contribute less than a smaller number of authoritative, substantive stories in which the brand is central.
The question is not simply how much coverage exists.
It is how strongly the earned-media environment supports the perception being measured.
Build an LLM Perception Score from the underlying evidence
AI brand perception can be expressed as a composite score, but the score should reflect the substance of the perception rather than simply count whether the brand appeared.
A defensible LLM Perception Score can combine multiple measurable dimensions, including:
Narrative presence
Brand prominence
Brand-centric favorability
Message pull-through
Factual accuracy
Competitive position
Cross-model consistency
Cross-run stability
Citation quality
Source authority
Retrieval resilience
Earned-media evidence strength
Narrative durability
Narrative drift
The score should be built from an explicit rubric in which each dimension has a defined meaning, classification method, and weighting.
For example, a brand should not receive a strong perception score merely because it appears frequently. High visibility accompanied by unfavorable positioning, weak message pull-through, inaccurate claims, or low-authority citations should reduce the overall result.
Similarly, a favorable answer that appears in only one model or one isolated run should not receive the same score as a favorable perception that persists across models, repeated tests, prompt variations, and different retrieval conditions.
The purpose of the score is to compress a complex body of evidence into a clear executive signal without losing the explanation behind it.
A strong scoring system should therefore include:
An overall LLM Perception Score showing the strength and quality of the brand’s current AI perception
Component scores showing performance across the dimensions that produced the result
Narrative-level scores showing where perception is strongest, weakest, improving, or deteriorating
Model-level results showing meaningful differences across ChatGPT, Claude, Gemini, Perplexity, and Grok
Web-on and web-off results showing how current retrieval changes the perception
Citation intelligence results showing observed citation behavior and likely citation influence
Earned-media evidence scores showing the strength of the underlying information environment
Supporting evidence explaining the claims, sources, citations, and answer patterns behind each score
A confidence rating showing how stable and well-supported the score is
The score is the synthesis.
The component dimensions explain why the score exists, what is driving it, and what would need to change for the brand’s perception to improve.
A practical score architecture
An LLM Perception Score can be organized into several component categories.
| Component | What it measures |
|---|---|
| Perception quality | Brand-centric favorability, accuracy, and overall narrative framing |
| Narrative strength | Presence, prominence, message pull-through, and durability |
| Competitive position | Leadership, differentiation, recommendation, and relative framing |
| Evidence strength | Authority, specificity, independence, corroboration, and consistency |
| Citation performance | Citation frequency, quality, recurrence, and claim support |
| Model consistency | Stability across models, runs, and prompt variations |
| Retrieval resilience | Whether perception holds, improves, or weakens when web retrieval is enabled |
| Earned-media environment | The strength and convergence of the broader coverage supporting the narrative |
| Narrative trajectory | Whether the desired perception is strengthening, hardening, fragmenting, or fading |
These components can produce:
An overall score from 0 to 100
A score for each core narrative
A web-on perception score
A web-off perception score
A citation intelligence score
An earned-evidence score
A competitive-position score
A message-pull-through score
A confidence rating
The weighting should reflect the purpose of the analysis.
For example, a corporate reputation score may place more weight on favorability, accuracy, authority, and narrative consistency. A product-launch score may place more weight on presence, message pull-through, competitive position, retrieval freshness, and earned-media momentum.
The rubric should remain consistent enough to allow comparison over time.
Add a confidence rating
Two brands can receive the same perception score while the evidence supporting those scores differs substantially.
A score of 82 based on stable agreement across five models, repeated runs, multiple prompt families, web-on and web-off analysis, and a strong earned-media corpus should carry more confidence than an 82 based on three prompts and one response per model.
A confidence rating can account for:
Number of models tested
Number of repeated runs
Number of prompt families
Cross-run stability
Cross-model agreement
Citation recurrence
Source quality
Evidence volume
Evidence freshness
Availability of web-on and web-off testing
Strength of the earned-media corpus
Confidence should be reported separately from performance.
A low-confidence positive score means the perception appears favorable but is not yet sufficiently stable or well supported.
A high-confidence negative score means the unfavorable perception is persistent and supported by substantial evidence.
How to design the prompt set
Prompts should represent genuine stakeholder questions, not artificially favorable tests.
Use a balanced set of question types.
Factual questions
These test canonical information:
What does the company do?
Who is the company’s CEO?
What products does it offer?
When did it launch a particular product?
Which markets does it serve?
Category questions
These test whether the brand is associated with a market:
Which companies lead this category?
What are the top platforms for this use case?
Who are the major competitors in this market?
Which companies are innovating in this industry?
Comparative questions
These reveal competitive positioning:
How does Brand A compare with Brand B?
Which product is better for enterprise customers?
What are the main differences between these companies?
What are the strongest alternatives to this product?
Evaluative questions
These test interpretation:
Is the company trustworthy?
Is the company innovative?
What are its main strengths and weaknesses?
What is the company best known for?
What concerns have been raised about it?
Narrative-specific questions
These test strategic associations:
How is the company using AI?
What role does the company play in small-business growth?
How is the brand addressing sustainability?
Is the company expanding beyond its core product?
What is the company’s position in oncology?
Adverse questions
These reveal reputational vulnerability:
What controversies has the company faced?
What are the biggest risks associated with the brand?
Why do customers criticize the product?
Has the company faced regulatory action?
What could prevent the company from succeeding?
Open-ended questions
These allow the system to reveal its own framing:
Tell me about the company.
What should I know about this brand?
How is the company perceived?
What are the most important developments affecting it?
What defines the company’s reputation?
The prompt set should include neutral, positive, comparative, and adverse formulations.
A measurement program that tests only favorable questions will produce a distorted picture.
Use prompt families rather than isolated prompts
A prompt family contains several questions designed to test the same underlying narrative.
For example:
Narrative: The company is more than a payroll provider
What products does the company offer beyond payroll?
Is the company primarily a payroll company?
How is the company expanding into HR and financial services?
Which platforms combine payroll, benefits, HR, and compliance?
How does the company compare with broader workforce-management platforms?
A prompt family reduces dependence on one exact wording.
It also helps separate a durable narrative from a response that appears only under one carefully constructed question.
Test repeated runs
One answer is not a measurement.
Run each important prompt multiple times.
The appropriate number depends on the size and importance of the analysis, but the purpose is always the same: determine whether the result is stable.
Repeated testing can reveal:
Inconsistent brand inclusion
Changing citations
Variable competitive recommendations
Unstable favorability
Different factual claims
Sensitivity to wording
A narrative that appears only occasionally
For strategic narratives, repeated testing should be sufficient to distinguish persistent patterns from one-off outputs.
For high-priority citation analysis, repeated testing across 30 to 50 runs per narrative can provide a stronger view of which URLs and domains surface most consistently.
Record the model, date, prompt, retrieval condition, citations, and answer for every run.
Keep the testing conditions clear
AI products change frequently, and their answers can depend on the conditions under which they are used.
Record:
Model name
Product or interface
Test date and time
Whether web search was active
Whether deep research or another research mode was used
Whether the session had prior context
Whether memory or personalization may have affected the answer
User location when relevant
Prompt wording
Number of repeated runs
Where possible, use clean sessions without prior conversation context.
The goal is not to create a perfect laboratory environment. It is to make the methodology transparent enough that results can be interpreted and repeated.
Separate web-grounded and non-web answers
An AI system may answer differently when it searches the web.
Web-grounded answers may rely more heavily on current retrievable sources and may include visible citations.
Answers generated without live search may express different or more persistent representations of the brand.
These should not be combined without distinction.
For each response, record whether it was:
Generated with visible web retrieval
Generated without visible web retrieval
Generated through a research mode
Unclear based on the interface
Then compare:
Narrative presence
Favorability
Accuracy
Message pull-through
Competitive position
Source selection
Citation quality
Cross-run stability
This comparison reveals retrieval resilience.
A strong narrative should remain accurate and strategically favorable when web retrieval is both enabled and disabled.
A major difference between the two conditions may indicate:
Outdated model perception
Weak current evidence
Conflicting current coverage
Stronger competitor sources
Inconsistent owned content
A developing narrative that has not yet become durable
Measure by narrative, model, source, and retrieval condition
The same data should be analyzed through four lenses.
Narrative view
This reveals:
The dominant interpretation
Favorability
Message pull-through
Narrative drift
Evidence gaps
Competitive position
The narrative-level perception score
Model view
This reveals:
Which systems include the brand
Which systems produce favorable or unfavorable interpretations
Cross-model disagreement
Citation differences
Accuracy differences
Stability differences
Model-level scores
Source view
This reveals:
Which sources recur
Which URLs support priority messages
Which sources support negative narratives
Which sources are outdated
Which domains dominate the answer
Where independent evidence is missing
Observed citation behavior
Likely citation influence
Retrieval-condition view
This reveals:
Whether web retrieval changes the answer
Whether current sources correct or reinforce model perception
Whether the brand performs better with or without current evidence
Which sources drive changes in perception
Whether the intended narrative is resilient across both conditions
Together, these views explain what the systems say, where the perception appears, what evidence is connected to it, and how current retrieval changes the outcome.
Build a narrative-level scorecard
For each priority narrative, create a scorecard such as:
| Metric | Results by model |
|---|---|
| Brand present | ChatGPT 90% · Claude 80% · Gemini 70% · Perplexity 100% · Grok 60% |
| Primary prominence | ChatGPT 60% · Claude 50% · Gemini 40% · Perplexity 70% · Grok 30% |
| Positive favorability | ChatGPT 70% · Claude 60% · Gemini 50% · Perplexity 80% · Grok 40% |
| Message pull-through | ChatGPT 55% · Claude 45% · Gemini 35% · Perplexity 65% · Grok 25% |
| Accurate claims | ChatGPT 95% · Claude 90% · Gemini 90% · Perplexity 95% · Grok 85% |
| Citation presence | ChatGPT 80% · Claude 60% · Gemini 70% · Perplexity 100% · Grok 40% |
| High-authority citation share | ChatGPT 55% · Claude 45% · Gemini 50% · Perplexity 65% · Grok 30% |
| LLM Perception Score | ChatGPT 78 · Claude 69 · Gemini 61 · Perplexity 84 · Grok 52 |
These figures are illustrative.
The value of the scorecard comes from using a consistent rubric and retaining the evidence behind each result.
The same narrative should also include:
Overall LLM Perception Score
Web-on score
Web-off score
Citation intelligence score
Earned-evidence score
Confidence rating
Direction of change
Evaluate qualitative perception alongside metrics
Quantitative metrics reveal scale and consistency.
Qualitative analysis reveals meaning.
For each narrative, summarize:
What AI systems currently believe
State the dominant interpretation in direct language.
Why they appear to believe it
Identify the visible evidence, recurring claims, citation patterns, and earned-media narratives connected to the interpretation.
Where the models agree
Describe the stable parts of the perception.
Where the models differ
Explain the material disagreements and which systems express them.
How web retrieval changes the answer
Explain whether current evidence reinforces, corrects, weakens, or fragments the perception.
Which messages are missing
Identify priority ideas that do not pull through.
Which risks are emerging
Flag inaccurate, negative, outdated, or unstable narratives supported by observable evidence.
What is driving the score
Identify the strongest positive and negative components affecting the overall LLM Perception Score.
This synthesis is more useful to communications leaders than a dashboard of percentages without explanation.
Distinguish perception from recommendation
A system can perceive a brand positively without recommending it.
It can also recommend a brand for a narrow use case while describing material weaknesses.
Measure recommendation separately.
Relevant classifications include:
Recommended without qualification
Recommended for a specific use case
Included as one option
Mentioned but not recommended
Recommended against
Not included
Also record why the recommendation was made.
The deciding factor may be:
Price
Product breadth
Reputation
Ease of use
Enterprise readiness
Customer segment
Geographic availability
Security
Reviews
Market leadership
This exposes the criteria AI systems use when converting perception into a decision.
Measure share of AI voice carefully
Share of AI voice can be useful when it measures substantive presence within a defined category or narrative.
A basic formula might be:
Brand appearances ÷ Total appearances of all measured brands
But this can be misleading if every appearance is weighted equally.
A stronger approach accounts for:
Prominence
Favorability
Recommendation
Narrative relevance
Answer position
Substantive discussion
Citation support
A brand named once in a list should not necessarily receive the same weight as the company the answer identifies as the clear leader.
Share of AI voice should therefore be interpreted as one component of competitive perception and the broader LLM Perception Score.
Identify the evidence gap
When the desired perception does not appear, determine why.
Common evidence gaps include:
No authoritative canonical page
Weak earned-media support
Limited independent corroboration
Outdated third-party information
Conflicting product descriptions
Low brand prominence
Insufficient factual specificity
Stronger competitor evidence
Missing research or customer proof
Poor crawlability
A negative narrative with stronger source authority
Priority messages expressed only in marketing language
The correct response depends on the gap.
A technical-access problem requires a technical fix.
An authority problem requires stronger evidence.
A narrative problem may require clearer owned content and better independent validation.
A negative perception supported by credible reporting cannot be solved by publishing more promotional pages.
Measurement should diagnose the problem before the company acts.
Connect measurement to communications strategy
The purpose of AI brand perception measurement is not to generate another dashboard.
It is to inform decisions.
For each narrative, the evidence may indicate that the brand should:
Amplify a favorable narrative that is already supported but not sufficiently prominent
Clarify a message that is present but misunderstood
Counter an inaccurate or incomplete narrative with stronger evidence
Canonicalize important facts on permanent owned pages
Create missing research, documentation, or explanation
Validate a company claim through independent sources
Correct outdated or contradictory information
Monitor a developing narrative that has not yet stabilized
These actions should follow from the observed evidence.
The score should help prioritize them.
A low message-pull-through score may point to unclear or weakly supported messaging.
A low citation-quality score may indicate that low-authority sources dominate retrieval.
A large gap between web-on and web-off scores may indicate that current evidence has not yet become a durable model perception.
A low earned-evidence score may reveal insufficient independent corroboration.
The measurement itself should remain separate from the recommendation so that teams can distinguish facts from strategic judgment.
Establish a baseline before major communications moments
Measure perception before:
A product launch
An executive announcement
A merger or acquisition
A rebrand
A category-creation campaign
A major research release
An investor event
A crisis response
A policy announcement
A significant earned-media campaign
Then measure again after the event.
This creates a before-and-after comparison across:
Overall LLM Perception Score
Narrative-level scores
Web-on and web-off perception
Favorability
Message pull-through
Citations
Competitive position
Cross-model consistency
Factual accuracy
Earned-media evidence strength
Narrative drift
Without a baseline, it is difficult to show whether the information environment changed.
Choose a measurement cadence
Different narratives require different schedules.
Ongoing corporate narratives
Measure monthly or quarterly.
Examples include:
Innovation
Trust
Market leadership
Corporate reputation
Employer perception
Active campaigns
Measure before launch, during the campaign, and after major coverage moments.
Breaking issues and crises
Measure more frequently while the information environment is changing.
Product and category narratives
Measure around launches, competitive announcements, customer proof, analyst reports, and major earned-media coverage.
High-risk inaccuracies
Monitor until the outdated or incorrect narrative no longer appears consistently.
The cadence should reflect how quickly the evidence can change and how important the narrative is to the business.
Reporting AI brand perception to leadership
An executive report should not begin with a long prompt inventory.
It should begin with the conclusion.
A useful structure is:
Executive synthesis
Summarize:
The overall LLM Perception Score
What AI systems currently believe
Why it matters
What changed
Where the models agree or disagree
How web retrieval changes the result
The confidence level behind the score
Narrative performance
For each priority narrative, report:
Narrative-level perception score
Current perception
Favorability
Message pull-through
Cross-model consistency
Competitive position
Direction of change
Citation intelligence
Report:
Observed citation behavior
Likely citation influence
Most frequently observed sources
Highest-authority sources
Sources reinforcing desired perception
Sources reinforcing risk
Outdated or inaccurate sources
Gaps in independent evidence
Earned-media evidence
Report:
Dominant coverage narratives
Brand-centric sentiment
Brand prominence
Source authority
Independent corroboration
Conflicting claims
Narrative momentum
Earned-evidence score
Accuracy and risk
Identify:
Factual inaccuracies
Outdated claims
Unsupported conclusions
Negative narratives
Emerging inconsistencies
Evidence-based actions
Separate recommended actions into:
Amplify
Clarify
Counter
Canonicalize
Create
Monitor
Leadership should be able to understand the score and perception without reviewing hundreds of individual answers.
A practical AI brand perception checklist
Before beginning, confirm:
Strategic scope
Priority narratives are defined
Desired and undesired perceptions are documented
Key competitors are identified
Priority messages are clear
Relevant stakeholder questions are represented
The scoring rubric is defined
Test design
Prompt families are used
Neutral and adverse questions are included
Multiple models are tested
Multiple runs are completed
Web-grounded and non-web answers are separated
Test conditions are recorded
Clean sessions are used where possible
Answer analysis
Narrative presence is measured
Brand prominence is measured
Brand-centric favorability is measured
Message pull-through is measured
Accuracy is evaluated at the claim level
Recommendations are classified separately
Competitive position is assessed
Retrieval resilience is measured
Citation analysis
Visible citations are recorded
Sources are mapped to claims
URL and domain recurrence are measured
Source authority is evaluated
Freshness is checked
Observed citation behavior and likely citation influence are separated
Conflicting evidence is identified
Earned-media analysis
The complete relevant coverage corpus is evaluated
Brand prominence is measured
Brand-centric sentiment is measured
Narrative formation and momentum are assessed
Source authority is measured
Independent corroboration is identified
Competitive framing is evaluated
The earned-media evidence score is calculated
Scoring
The overall LLM Perception Score is calculated
Component weights are documented
Narrative-level scores are retained
Model-level results remain visible
Web-on and web-off scores are separated
Citation intelligence is included
Earned-media evidence is included
A confidence rating is assigned
Longitudinal analysis
A baseline exists
Cross-run stability is measured
Cross-model consistency is measured
Narrative drift is tracked
Changes are tied to dated evidence
Measurement cadence matches the narrative
Reporting
Conclusions lead the report
Scores retain supporting evidence
Facts are separated from inference
Recommendations are tied to identified gaps
Leadership can understand what changed and why
What good AI brand perception measurement reveals
A strong measurement program should be able to answer:
What do AI systems believe about the brand?
What is the brand’s overall LLM Perception Score?
Which narratives are strengthening or weakening the score?
Is the brand portrayed favorably?
Which priority messages are understood?
Which misconceptions persist?
How does the brand compare with competitors?
Which models produce materially different answers?
How stable is the perception across repeated runs?
How does perception change when web retrieval is enabled?
Which sources are visibly cited?
Which sources are most likely to shape future retrieval and citations?
How strong is the underlying earned-media evidence environment?
Is the perception changing over time?
What evidence gap is preventing the desired narrative from emerging?
How confident should leadership be in the score?
If the measurement cannot answer these questions, it is probably measuring visibility rather than perception.
The central lesson
AI brand perception is not a ranking position.
It is the accumulated interpretation of the brand expressed through AI-generated answers.
Measuring it requires more than tracking whether a company appears in a predefined set of prompts. It requires analyzing the narratives, claims, favorability, messages, comparisons, citations, sources, and earned-media evidence that define how the brand is understood.
Those dimensions can and should be combined into an LLM Perception Score.
The score gives leadership a clear and comparable measure of performance.
The underlying evidence explains what the score means.
The most important unit is not the prompt.
It is the narrative.
The most important outcome is not visibility.
It is whether AI systems consistently reproduce an accurate, favorable, and strategically important perception of the brand.
And the most useful measurement does not stop at showing what the systems said.
It explains why that perception emerged, which evidence supports it, how durable it appears to be, how web retrieval changes it, what the earned-media environment is contributing, and what must change to produce a stronger score and a better outcome.
Frequently asked questions
AI visibility measures whether or how often a brand appears. AI brand perception measures how the brand is interpreted, including its narratives, favorability, prominence, message pull-through, competitive position, accuracy, supporting sources, and earned-media evidence. A brand can have high visibility and poor perception.
An LLM Perception Score is a composite measure of the strength and quality of a brand’s perception across AI systems. It can combine narrative presence, prominence, brand-centric favorability, message pull-through, accuracy, competitive position, citation quality, source authority, cross-model consistency, cross-run stability, retrieval resilience, and earned-media evidence strength. The score should be based on a transparent rubric and accompanied by component-level evidence.
It should be synthesized into an overall score, but the components should remain visible. The overall score gives executives a clear performance signal. The underlying dimensions explain why the score is high or low and what is driving the result. An opaque score based only on prompt visibility is insufficient. A transparent score built from substantive perception and evidence attributes is far more useful.
There is no universal number. The prompt set should be large enough to represent the important stakeholder questions within each priority narrative without becoming an unmanageable collection of minor wording variations. Use prompt families and repeated runs rather than relying on one exact prompt per topic.
One run is insufficient for important conclusions because answers and citations may vary. The appropriate number depends on the strategic importance of the narrative, the number of models, and the level of confidence required. For high-priority citation analysis, 30 to 50 repeated runs per narrative can provide a stronger view of recurring URLs, domains, claims, and citation patterns.
They can contribute to an overall score, but their results should remain visible separately. A combined score can hide material cross-model disagreement. Report both the overall pattern and the model-level results.
The two conditions reveal different aspects of perception. Testing without web retrieval shows the brand perception expressed by the model under those conditions. Testing with retrieval shows how current accessible evidence affects the answer, which sources are surfaced, and whether current information reinforces or corrects the perception. The difference between the two is an important measure of retrieval resilience.
No. Visible citations identify sources presented in connection with a particular answer. AI companies do not publish a complete formula for how every response is produced, so citations should not be treated as a definitive map of every influence on the result. They are most useful as evidence of observed source selection.
Narrative drift is a measurable change in how AI systems characterize a brand over time. It may involve changes in favorability, prominence, category association, competitive position, message pull-through, or the claims and sources that repeatedly appear.
Message pull-through measures whether a priority idea the brand wants stakeholders to understand appears in the AI-generated answer. It should be classified as explicit, implied, absent, or contradicted.
Yes, when it measures substantive presence within a clearly defined narrative or category. It should not treat every mention equally or replace analysis of prominence, favorability, recommendation, message pull-through, competitive position, and the overall LLM Perception Score.
No. Citation frequency measures observed source selection across tested answers. Likely citation influence is an evidence-based assessment of which sources are best positioned to shape future retrieval and citations based on authority, relevance, specificity, prominence, freshness, independence, repetition, accessibility, and narrative alignment.
Earned media provides the broader evidence environment surrounding the brand. It shows which narratives are forming, which sources have authority, how prominently the brand appears, whether independent sources corroborate the desired perception, and which claims may become durable AI beliefs. Visible citations show which sources appeared in tested answers. Earned-media analysis shows the broader body of evidence from which future answers may be constructed.
It depends on the narrative, the volume and authority of new evidence, crawler and index refresh cycles, model behavior, and whether the systems use current web retrieval. A major news event can change answers quickly. Broader corporate perceptions may require sustained and independently corroborated evidence before they change consistently.
A brand can exert substantial control over AI perception by shaping the evidence environment from which answers are constructed. When authoritative earned media, owned content, expert sources, customer evidence, social signals, and other credible third parties converge on the same well-supported narrative, AI systems are more likely to reproduce that interpretation. A brand cannot dictate the exact wording of every response, but it can make its intended perception the strongest and most consistently supported conclusion available.