GEO Guides AI Brand Perception

How to Measure AI Brand Perception Across ChatGPT, Claude, Gemini, Perplexity, and Grok

Learn how to measure and score AI brand perception across ChatGPT, Claude, Gemini, Perplexity, and Grok using narrative presence, brand-centric favorability, message pull-through, citations, competitive position, and cross-model consistency.

Five translucent amber glass blocks in a row under warm light

Measuring AI brand perception is not the same as checking whether a company appears in a handful of prompts.

A brand can appear frequently and still be misunderstood.

It can receive a high visibility score while unfavorable narratives dominate the answer. It can be cited often but cited in connection with controversy, weakness, or outdated information. It can perform well in one model and poorly in another. It can also appear in an answer without the messages the company considers most important.

True AI brand perception measurement must answer a more consequential question:

What do AI systems believe about the brand, why do they appear to believe it, and how consistently is that perception reproduced?

That requires measuring more than mentions.

A complete framework should evaluate:

  • Narrative presence

  • Brand prominence

  • Brand-centric favorability

  • Message pull-through

  • Factual accuracy

  • Citation consistency

  • Source authority

  • Competitive position

  • Cross-model variance

  • Cross-run stability

  • Narrative drift

  • Likely versus observed citation influence

  • The strength of the underlying earned-media evidence environment

These dimensions can then be combined through a transparent rubric to produce an overall LLM Perception Score.

The score provides a clear executive signal. The underlying dimensions explain what is driving it.

What is AI brand perception?

AI brand perception is the interpretation of a company, product, executive, or issue expressed through an AI-generated answer.

It includes the facts, narratives, associations, judgments, sources, and comparisons the system uses to describe the brand.

For example, an AI system may describe a company as:

  • An established market leader

  • An emerging challenger

  • A low-cost alternative

  • A product innovator

  • A trusted institutional brand

  • A company facing regulatory pressure

  • A business struggling to differentiate itself

  • A category pioneer

  • A controversial employer

  • A reliable but outdated incumbent

These are not merely mentions. They are compressed interpretations.

An AI answer may synthesize news coverage, company claims, product documentation, reviews, public records, research, and other available information into a few sentences. Those sentences can shape how customers, journalists, investors, employees, policymakers, and other stakeholders understand the company.

Measuring AI brand perception means systematically evaluating those interpretations.

Why prompt visibility is not enough

Many AI-monitoring programs begin with a list of prompts such as:

  • What are the best companies in this category?

  • Who are the leaders in this market?

  • What products should I consider?

  • What is this company known for?

  • Is this brand trustworthy?

The system records whether the brand appears and may convert the result into a visibility score.

That can be useful, but it is incomplete.

A visibility score generally cannot tell you:

  • What the system believes about the brand

  • Whether the brand is portrayed favorably

  • Which narratives dominate the answer

  • Whether priority messages are understood

  • Why a competitor is preferred

  • Which sources support the interpretation

  • Whether cited evidence is current

  • Whether the answer is stable across repeated runs

  • Whether perception changes when web retrieval is enabled

  • Whether perception is changing over time

  • What information gap is producing the result

Consider two brands that appear in 80% of tested answers.

Brand A is described as the established leader with the strongest enterprise capabilities.

Brand B appears just as often but is described as a less expensive alternative with limited functionality.

Their visibility is identical. Their perception is not.

The strategic value lies in understanding and scoring the difference.

Start with narratives, not an endless prompt list

No company can know every question a stakeholder may ask an AI system.

People can express the same intent in thousands of ways. They may include different context, constraints, comparisons, assumptions, and terminology.

A measurement strategy built around exact prompts therefore risks optimizing for a small set of guessed questions rather than understanding the broader perception of the brand.

Start with the narratives that matter most to the business.

These may include:

  • Corporate reputation

  • Product leadership

  • Innovation

  • Artificial intelligence

  • Customer trust

  • Security

  • Pricing

  • Sustainability

  • Executive leadership

  • Workplace culture

  • Regulatory issues

  • Financial stability

  • Market expansion

  • A major product launch

  • A controversy or crisis

  • A competitive differentiator

Each narrative can then be tested through a representative set of questions.

For example, a company measuring its AI innovation narrative might test:

  • How is the company using AI?

  • Is the company considered an AI leader?

  • What differentiates its AI strategy?

  • Which companies are leading AI adoption in its industry?

  • How does its AI offering compare with competitors?

  • What concerns have been raised about its use of AI?

  • What evidence supports its AI claims?

The prompts are measurement instruments.

The narrative is the unit of analysis.

Measure the answer, not only the brand mention

Every response should be evaluated as a complete interpretation.

At minimum, record:

  1. Whether the brand appears

  2. How prominently it appears

  3. Which narrative is expressed

  4. Whether the portrayal is favorable

  5. Which priority messages appear

  6. Which competitors are included

  7. How the brands are compared

  8. Which sources are cited

  9. Whether the claims are accurate

  10. Whether web retrieval was enabled

  11. How the answer differs from other runs and models

  12. How the answer relates to the broader earned-media evidence environment

This turns AI measurement from a mention-counting exercise into perception analysis.

The three layers of AI brand perception analysis

A complete analysis should examine the brand through three connected layers.

1. LLM perception without web retrieval

Testing without live web retrieval helps evaluate the brand perception expressed by the model under the tested conditions.

This can reveal:

  • Established brand associations

  • Durable narrative patterns

  • Outdated beliefs

  • Persistent misconceptions

  • Differences among models

  • Messages that appear without current web evidence

  • Competitive positions that have become embedded in recurring answers

Because AI companies do not disclose every factor involved in producing an answer, this should be treated as an observed representation of the model’s current perception, not a complete view into its internal knowledge or training process.

2. LLM perception with web retrieval and citation analysis

Testing with web retrieval shows how current accessible information affects the answer.

This analysis can measure:

  • Whether retrieval changes the brand’s perception

  • Whether current information corrects outdated beliefs

  • Which claims become more or less prominent

  • Which sources are visibly cited

  • Which sources recur across repeated runs

  • Whether cited sources actually support the associated claims

  • Whether retrieval strengthens or weakens the intended narrative

  • How citation behavior differs across models

  • Whether the same sources consistently shape the answer

This layer connects perception with observable source selection.

3. Earned-media narrative analysis

Earned-media analysis evaluates the broader coverage environment surrounding the narrative.

It can measure:

  • Coverage volume

  • Brand prominence

  • Brand-centric sentiment

  • Narrative formation

  • Narrative momentum

  • Message pull-through

  • Source authority

  • Factual specificity

  • Independent corroboration

  • Conflicting evidence

  • Competitive positioning

  • The claims most likely to become durable AI beliefs

This layer is important because visible citations represent only the sources surfaced in tested answers.

The broader earned-media environment reveals the complete body of current evidence shaping human perception and potentially influencing future AI retrieval, recurring citations, and longer-term model perception.

Together, these three layers provide a more complete basis for measuring and scoring AI perception than prompt visibility alone.

The core AI brand perception metrics

A strong measurement framework should separate the dimensions that ultimately feed the overall score.

1. Narrative presence

Narrative presence measures whether the brand is associated with the strategic idea being evaluated.

This is more specific than simple brand visibility.

A company may appear in an answer about enterprise software without being associated with the narrative it wants to own, such as artificial intelligence, financial transformation, security, or ease of use.

Narrative presence can be classified as:

  • Primary: The brand is presented as a leading example of the narrative.

  • Substantive: The brand is meaningfully connected to the narrative but is not the main example.

  • Incidental: The brand appears without a meaningful narrative association.

  • Absent: The brand does not appear.

For example, if an AI system lists a company among ten payroll vendors but says nothing about its small-business expertise, the brand is visible but the intended narrative is absent.

2. Brand prominence

Prominence measures how central the brand is to the answer.

Relevant signals include:

  • Whether the brand appears in the opening answer

  • Its position in a list

  • The amount of substantive discussion

  • Whether it receives its own section

  • Whether it is named in the conclusion

  • Whether its evidence is used to define the category

  • Whether it is treated as a primary or secondary example

A brand mentioned at the end of a long answer does not have the same influence as a brand used to frame the entire response.

Prominence should be measured separately from presence.

3. Brand-centric favorability

Traditional sentiment analysis often measures the general emotional tone of a passage.

That is not sufficient for brand perception.

An article or AI answer can contain negative language while positioning the brand favorably. For example, a company may be praised for helping customers through an economic downturn. The subject matter is negative, but the brand’s position is positive.

Brand-centric favorability evaluates how the brand itself is positioned.

A practical classification is:

  • Positive: The answer strengthens confidence in the brand or associates it with favorable qualities, performance, or outcomes.

  • Neutral: The answer presents primarily factual information without a meaningful positive or negative implication.

  • Mixed: The answer includes material strengths and weaknesses.

  • Negative: The answer weakens confidence in the brand or associates it with unfavorable qualities, risks, failures, or outcomes.

The analysis should explain why the classification was assigned.

A score without the underlying interpretation is difficult to act on.

4. Message pull-through

Message pull-through measures whether the ideas the brand wants stakeholders to understand actually appear in the answer.

Priority messages might include:

  • The company serves more than one product category

  • Its technology is designed for enterprise use

  • It is a leader in a specific scientific field

  • Its product reduces a defined customer problem

  • It is expanding beyond its legacy business

  • Its approach differs fundamentally from a competitor’s

  • It has a particular security, quality, or research advantage

  • It serves a specific customer segment

For each message, classify the result as:

  • Explicit: The answer clearly states the message.

  • Implied: The answer supports the message without stating it directly.

  • Absent: The message does not appear.

  • Contradicted: The answer presents an incompatible perception.

Message pull-through reveals whether the information environment is communicating the intended narrative, not merely generating brand mentions.

5. Factual accuracy

AI systems can produce inaccurate, outdated, incomplete, or unsupported claims.

Accuracy analysis should identify:

  • Incorrect facts

  • Outdated facts

  • Fabricated details

  • Misattributed claims

  • Conflated products or companies

  • Missing qualifications

  • Incorrect executive roles

  • Outdated pricing or policies

  • Misrepresented regulatory status

  • Unsupported superlatives

  • Incorrect competitive comparisons

Not every omission is an error.

The evaluator should distinguish between:

  • Incorrect: The statement conflicts with reliable evidence.

  • Outdated: The statement was previously true but no longer reflects the current state.

  • Incomplete: The statement is directionally accurate but omits material context.

  • Unsupported: The claim lacks identifiable evidence.

  • Accurate: The statement is supported by current evidence.

Accuracy should be evaluated at the claim level, not only at the answer level.

6. Citation presence

Citation presence measures whether the answer visibly identifies sources.

Depending on the platform and mode, a response may include:

  • Inline citations

  • Footnotes

  • Linked publisher names

  • Source cards

  • A separate source panel

  • References at the end

  • No visible citations

Citation presence matters because it allows the evaluator to inspect the evidence presented with the answer.

But a cited answer is not automatically accurate, and an uncited answer does not provide a complete account of the information involved in producing it.

Treat visible citations as observable evidence, not a complete explanation of the system’s internal process.

7. Citation consistency

Citation consistency measures how often the same sources recur across repeated answers to related questions.

One citation in one response may be incidental.

Repeated citation of the same source across several runs is a stronger pattern.

Track:

  • Domain citation frequency

  • URL citation frequency

  • Citation frequency by narrative

  • Citation frequency by model

  • Citation frequency by prompt variation

  • Citation frequency over time

  • Citation position within the answer

  • Whether the source directly supports the associated claim

Citation consistency is especially important because AI answers can vary from one run to the next.

The objective is not to identify one source that appeared once. It is to understand which sources repeatedly surface as evidence.

8. Source authority

Not all citations carry equal weight.

A source-authority assessment should consider:

  • Expertise on the subject

  • Editorial or institutional standards

  • Original reporting

  • Firsthand access

  • Primary evidence

  • Independence

  • Factual specificity

  • Brand prominence

  • Freshness

  • Relevance to the claim

A corporate product page may be highly authoritative for product specifications but weak evidence for an independent claim of market leadership.

A regulator may be authoritative for an approval or enforcement action.

A specialist publication may be more authoritative for a technical question than a larger general-interest publication.

Source authority must be assessed in relation to the claim.

9. Competitive position

AI systems frequently answer brand questions through comparison.

Competitive-position analysis should examine:

  • Which competitors appear

  • Which brand appears first

  • Which brand is described as the leader

  • Which differentiators are assigned to each company

  • Which weaknesses are attached to each company

  • Which brand is recommended for which use case

  • Which evidence supports the comparison

  • Whether the category definition favors one competitor

  • Whether the brand is framed as an incumbent, challenger, specialist, or alternative

A brand may have strong standalone perception but weak comparative positioning.

For example, an AI system may describe a product positively while consistently recommending a competitor for enterprise customers. That distinction would be lost in a basic sentiment or visibility score.

10. Cross-model consistency

ChatGPT, Claude, Gemini, Perplexity, and Grok may produce different answers to the same question.

Cross-model consistency measures whether the brand’s core perception is stable across systems.

Classify the result as:

  • Highly consistent: The systems express substantially the same narrative and conclusion.

  • Generally consistent: The systems agree on the central perception but differ in detail.

  • Mixed: The systems produce materially different interpretations.

  • Contradictory: The systems reach opposing conclusions.

  • Insufficient evidence: Too few systems or responses provide a usable answer.

Cross-model disagreement is strategically important.

It may indicate:

  • Different source retrieval

  • Different indexes or search partners

  • Different model behavior

  • Uneven coverage across information ecosystems

  • Conflicting evidence about the brand

  • A narrative that has not yet become durable

  • Greater sensitivity to prompt wording

The goal is not necessarily identical wording.

The goal is a consistent underlying perception.

11. Cross-run stability

The same model can answer the same question differently across repeated runs.

Cross-run stability measures whether the conclusion persists.

For each test, track:

  • Brand inclusion

  • Narrative classification

  • Favorability

  • Message pull-through

  • Competitor inclusion

  • Recommendation

  • Citations

  • Factual claims

A perception that appears once but disappears in most repeated runs is not yet stable.

A narrative that persists across repeated runs, prompt variations, models, and retrieval conditions is more likely to represent a durable AI belief.

12. Narrative drift

Narrative drift measures how AI perception changes over time.

A brand may gradually shift from being described as:

  • A single-product company to a broader platform

  • A market leader to a legacy incumbent

  • An emerging challenger to an established competitor

  • A trusted company to a controversial one

  • A traditional business to an AI innovator

  • A consumer brand to an enterprise provider

Drift can be:

  • Positive: Perception moves toward the desired narrative.

  • Negative: Perception moves away from the desired narrative.

  • Corrective: An outdated or inaccurate belief weakens.

  • Fragmenting: Models become less consistent.

  • Hardening: The same narrative becomes increasingly stable.

  • Fading: A narrative appears less frequently or prominently.

Narrative drift should be measured against a fixed baseline and tied to dated evidence.

Without historical comparison, a team cannot tell whether perception is improving, deteriorating, or remaining unchanged.

13. Observed citation behavior

Observed citation behavior measures the sources that visibly appear in actual AI answers.

This is the most directly measurable source signal.

It can be evaluated through:

  • Repeated model runs

  • Citation frequency

  • Citation-to-claim mapping

  • URL recurrence

  • Domain recurrence

  • Cross-model citation patterns

  • Citation changes over time

Observed citation behavior answers:

Which sources are these systems visibly selecting when they answer questions about this narrative?

It should not be interpreted as a complete account of every factor that contributed to the response.

14. Likely citation influence

Likely citation influence assesses which sources are best positioned to shape future retrieval and citation behavior.

Evaluate each source based on:

  • Source authority

  • Direct relevance

  • Factual specificity

  • Brand prominence

  • Freshness

  • Independence

  • Repetition across credible sources

  • Narrative alignment

  • Accessibility

This analysis is inferential.

It does not claim access to a model’s internal weighting or training process. It identifies the sources most likely to support future answers based on observable evidence.

The distinction matters:

  • Observed citation behavior measures what visibly happened.

  • Likely citation influence assesses which sources are best positioned to shape future answers.

Both are necessary.

Observed citations alone may miss strong sources that have not yet surfaced. Likely influence alone may overestimate sources that appear authoritative but are rarely retrieved.

15. Earned-media evidence strength

Earned-media evidence strength measures the quality of the broader information environment surrounding the narrative.

It should account for:

  • Source authority

  • Brand prominence

  • Brand-centric sentiment

  • Factual specificity

  • Independent corroboration

  • Narrative consistency

  • Coverage freshness

  • Narrative momentum

  • Competitive framing

  • Conflicting evidence

  • The degree to which the same perception recurs across credible sources

Raw coverage volume should not automatically increase the score.

A large volume of low-authority, syndicated, or passing mentions may contribute less than a smaller number of authoritative, substantive stories in which the brand is central.

The question is not simply how much coverage exists.

It is how strongly the earned-media environment supports the perception being measured.

Build an LLM Perception Score from the underlying evidence

AI brand perception can be expressed as a composite score, but the score should reflect the substance of the perception rather than simply count whether the brand appeared.

A defensible LLM Perception Score can combine multiple measurable dimensions, including:

  • Narrative presence

  • Brand prominence

  • Brand-centric favorability

  • Message pull-through

  • Factual accuracy

  • Competitive position

  • Cross-model consistency

  • Cross-run stability

  • Citation quality

  • Source authority

  • Retrieval resilience

  • Earned-media evidence strength

  • Narrative durability

  • Narrative drift

The score should be built from an explicit rubric in which each dimension has a defined meaning, classification method, and weighting.

For example, a brand should not receive a strong perception score merely because it appears frequently. High visibility accompanied by unfavorable positioning, weak message pull-through, inaccurate claims, or low-authority citations should reduce the overall result.

Similarly, a favorable answer that appears in only one model or one isolated run should not receive the same score as a favorable perception that persists across models, repeated tests, prompt variations, and different retrieval conditions.

The purpose of the score is to compress a complex body of evidence into a clear executive signal without losing the explanation behind it.

A strong scoring system should therefore include:

  1. An overall LLM Perception Score showing the strength and quality of the brand’s current AI perception

  2. Component scores showing performance across the dimensions that produced the result

  3. Narrative-level scores showing where perception is strongest, weakest, improving, or deteriorating

  4. Model-level results showing meaningful differences across ChatGPT, Claude, Gemini, Perplexity, and Grok

  5. Web-on and web-off results showing how current retrieval changes the perception

  6. Citation intelligence results showing observed citation behavior and likely citation influence

  7. Earned-media evidence scores showing the strength of the underlying information environment

  8. Supporting evidence explaining the claims, sources, citations, and answer patterns behind each score

  9. A confidence rating showing how stable and well-supported the score is

The score is the synthesis.

The component dimensions explain why the score exists, what is driving it, and what would need to change for the brand’s perception to improve.

A practical score architecture

An LLM Perception Score can be organized into several component categories.

Component What it measures
Perception quality Brand-centric favorability, accuracy, and overall narrative framing
Narrative strength Presence, prominence, message pull-through, and durability
Competitive position Leadership, differentiation, recommendation, and relative framing
Evidence strength Authority, specificity, independence, corroboration, and consistency
Citation performance Citation frequency, quality, recurrence, and claim support
Model consistency Stability across models, runs, and prompt variations
Retrieval resilience Whether perception holds, improves, or weakens when web retrieval is enabled
Earned-media environment The strength and convergence of the broader coverage supporting the narrative
Narrative trajectory Whether the desired perception is strengthening, hardening, fragmenting, or fading

These components can produce:

  • An overall score from 0 to 100

  • A score for each core narrative

  • A web-on perception score

  • A web-off perception score

  • A citation intelligence score

  • An earned-evidence score

  • A competitive-position score

  • A message-pull-through score

  • A confidence rating

The weighting should reflect the purpose of the analysis.

For example, a corporate reputation score may place more weight on favorability, accuracy, authority, and narrative consistency. A product-launch score may place more weight on presence, message pull-through, competitive position, retrieval freshness, and earned-media momentum.

The rubric should remain consistent enough to allow comparison over time.

Add a confidence rating

Two brands can receive the same perception score while the evidence supporting those scores differs substantially.

A score of 82 based on stable agreement across five models, repeated runs, multiple prompt families, web-on and web-off analysis, and a strong earned-media corpus should carry more confidence than an 82 based on three prompts and one response per model.

A confidence rating can account for:

  • Number of models tested

  • Number of repeated runs

  • Number of prompt families

  • Cross-run stability

  • Cross-model agreement

  • Citation recurrence

  • Source quality

  • Evidence volume

  • Evidence freshness

  • Availability of web-on and web-off testing

  • Strength of the earned-media corpus

Confidence should be reported separately from performance.

A low-confidence positive score means the perception appears favorable but is not yet sufficiently stable or well supported.

A high-confidence negative score means the unfavorable perception is persistent and supported by substantial evidence.

How to design the prompt set

Prompts should represent genuine stakeholder questions, not artificially favorable tests.

Use a balanced set of question types.

Factual questions

These test canonical information:

  • What does the company do?

  • Who is the company’s CEO?

  • What products does it offer?

  • When did it launch a particular product?

  • Which markets does it serve?

Category questions

These test whether the brand is associated with a market:

  • Which companies lead this category?

  • What are the top platforms for this use case?

  • Who are the major competitors in this market?

  • Which companies are innovating in this industry?

Comparative questions

These reveal competitive positioning:

  • How does Brand A compare with Brand B?

  • Which product is better for enterprise customers?

  • What are the main differences between these companies?

  • What are the strongest alternatives to this product?

Evaluative questions

These test interpretation:

  • Is the company trustworthy?

  • Is the company innovative?

  • What are its main strengths and weaknesses?

  • What is the company best known for?

  • What concerns have been raised about it?

Narrative-specific questions

These test strategic associations:

  • How is the company using AI?

  • What role does the company play in small-business growth?

  • How is the brand addressing sustainability?

  • Is the company expanding beyond its core product?

  • What is the company’s position in oncology?

Adverse questions

These reveal reputational vulnerability:

  • What controversies has the company faced?

  • What are the biggest risks associated with the brand?

  • Why do customers criticize the product?

  • Has the company faced regulatory action?

  • What could prevent the company from succeeding?

Open-ended questions

These allow the system to reveal its own framing:

  • Tell me about the company.

  • What should I know about this brand?

  • How is the company perceived?

  • What are the most important developments affecting it?

  • What defines the company’s reputation?

The prompt set should include neutral, positive, comparative, and adverse formulations.

A measurement program that tests only favorable questions will produce a distorted picture.

Use prompt families rather than isolated prompts

A prompt family contains several questions designed to test the same underlying narrative.

For example:

Narrative: The company is more than a payroll provider

  • What products does the company offer beyond payroll?

  • Is the company primarily a payroll company?

  • How is the company expanding into HR and financial services?

  • Which platforms combine payroll, benefits, HR, and compliance?

  • How does the company compare with broader workforce-management platforms?

A prompt family reduces dependence on one exact wording.

It also helps separate a durable narrative from a response that appears only under one carefully constructed question.

Test repeated runs

One answer is not a measurement.

Run each important prompt multiple times.

The appropriate number depends on the size and importance of the analysis, but the purpose is always the same: determine whether the result is stable.

Repeated testing can reveal:

  • Inconsistent brand inclusion

  • Changing citations

  • Variable competitive recommendations

  • Unstable favorability

  • Different factual claims

  • Sensitivity to wording

  • A narrative that appears only occasionally

For strategic narratives, repeated testing should be sufficient to distinguish persistent patterns from one-off outputs.

For high-priority citation analysis, repeated testing across 30 to 50 runs per narrative can provide a stronger view of which URLs and domains surface most consistently.

Record the model, date, prompt, retrieval condition, citations, and answer for every run.

Keep the testing conditions clear

AI products change frequently, and their answers can depend on the conditions under which they are used.

Record:

  • Model name

  • Product or interface

  • Test date and time

  • Whether web search was active

  • Whether deep research or another research mode was used

  • Whether the session had prior context

  • Whether memory or personalization may have affected the answer

  • User location when relevant

  • Prompt wording

  • Number of repeated runs

Where possible, use clean sessions without prior conversation context.

The goal is not to create a perfect laboratory environment. It is to make the methodology transparent enough that results can be interpreted and repeated.

Separate web-grounded and non-web answers

An AI system may answer differently when it searches the web.

Web-grounded answers may rely more heavily on current retrievable sources and may include visible citations.

Answers generated without live search may express different or more persistent representations of the brand.

These should not be combined without distinction.

For each response, record whether it was:

  • Generated with visible web retrieval

  • Generated without visible web retrieval

  • Generated through a research mode

  • Unclear based on the interface

Then compare:

  • Narrative presence

  • Favorability

  • Accuracy

  • Message pull-through

  • Competitive position

  • Source selection

  • Citation quality

  • Cross-run stability

This comparison reveals retrieval resilience.

A strong narrative should remain accurate and strategically favorable when web retrieval is both enabled and disabled.

A major difference between the two conditions may indicate:

  • Outdated model perception

  • Weak current evidence

  • Conflicting current coverage

  • Stronger competitor sources

  • Inconsistent owned content

  • A developing narrative that has not yet become durable

Measure by narrative, model, source, and retrieval condition

The same data should be analyzed through four lenses.

Narrative view

This reveals:

  • The dominant interpretation

  • Favorability

  • Message pull-through

  • Narrative drift

  • Evidence gaps

  • Competitive position

  • The narrative-level perception score

Model view

This reveals:

  • Which systems include the brand

  • Which systems produce favorable or unfavorable interpretations

  • Cross-model disagreement

  • Citation differences

  • Accuracy differences

  • Stability differences

  • Model-level scores

Source view

This reveals:

  • Which sources recur

  • Which URLs support priority messages

  • Which sources support negative narratives

  • Which sources are outdated

  • Which domains dominate the answer

  • Where independent evidence is missing

  • Observed citation behavior

  • Likely citation influence

Retrieval-condition view

This reveals:

  • Whether web retrieval changes the answer

  • Whether current sources correct or reinforce model perception

  • Whether the brand performs better with or without current evidence

  • Which sources drive changes in perception

  • Whether the intended narrative is resilient across both conditions

Together, these views explain what the systems say, where the perception appears, what evidence is connected to it, and how current retrieval changes the outcome.

Build a narrative-level scorecard

For each priority narrative, create a scorecard such as:

Metric Results by model
Brand present ChatGPT 90% · Claude 80% · Gemini 70% · Perplexity 100% · Grok 60%
Primary prominence ChatGPT 60% · Claude 50% · Gemini 40% · Perplexity 70% · Grok 30%
Positive favorability ChatGPT 70% · Claude 60% · Gemini 50% · Perplexity 80% · Grok 40%
Message pull-through ChatGPT 55% · Claude 45% · Gemini 35% · Perplexity 65% · Grok 25%
Accurate claims ChatGPT 95% · Claude 90% · Gemini 90% · Perplexity 95% · Grok 85%
Citation presence ChatGPT 80% · Claude 60% · Gemini 70% · Perplexity 100% · Grok 40%
High-authority citation share ChatGPT 55% · Claude 45% · Gemini 50% · Perplexity 65% · Grok 30%
LLM Perception Score ChatGPT 78 · Claude 69 · Gemini 61 · Perplexity 84 · Grok 52

These figures are illustrative.

The value of the scorecard comes from using a consistent rubric and retaining the evidence behind each result.

The same narrative should also include:

  • Overall LLM Perception Score

  • Web-on score

  • Web-off score

  • Citation intelligence score

  • Earned-evidence score

  • Confidence rating

  • Direction of change

Evaluate qualitative perception alongside metrics

Quantitative metrics reveal scale and consistency.

Qualitative analysis reveals meaning.

For each narrative, summarize:

What AI systems currently believe

State the dominant interpretation in direct language.

Why they appear to believe it

Identify the visible evidence, recurring claims, citation patterns, and earned-media narratives connected to the interpretation.

Where the models agree

Describe the stable parts of the perception.

Where the models differ

Explain the material disagreements and which systems express them.

How web retrieval changes the answer

Explain whether current evidence reinforces, corrects, weakens, or fragments the perception.

Which messages are missing

Identify priority ideas that do not pull through.

Which risks are emerging

Flag inaccurate, negative, outdated, or unstable narratives supported by observable evidence.

What is driving the score

Identify the strongest positive and negative components affecting the overall LLM Perception Score.

This synthesis is more useful to communications leaders than a dashboard of percentages without explanation.

Distinguish perception from recommendation

A system can perceive a brand positively without recommending it.

It can also recommend a brand for a narrow use case while describing material weaknesses.

Measure recommendation separately.

Relevant classifications include:

  • Recommended without qualification

  • Recommended for a specific use case

  • Included as one option

  • Mentioned but not recommended

  • Recommended against

  • Not included

Also record why the recommendation was made.

The deciding factor may be:

  • Price

  • Product breadth

  • Reputation

  • Ease of use

  • Enterprise readiness

  • Customer segment

  • Geographic availability

  • Security

  • Reviews

  • Market leadership

This exposes the criteria AI systems use when converting perception into a decision.

Measure share of AI voice carefully

Share of AI voice can be useful when it measures substantive presence within a defined category or narrative.

A basic formula might be:

Brand appearances ÷ Total appearances of all measured brands

But this can be misleading if every appearance is weighted equally.

A stronger approach accounts for:

  • Prominence

  • Favorability

  • Recommendation

  • Narrative relevance

  • Answer position

  • Substantive discussion

  • Citation support

A brand named once in a list should not necessarily receive the same weight as the company the answer identifies as the clear leader.

Share of AI voice should therefore be interpreted as one component of competitive perception and the broader LLM Perception Score.

Identify the evidence gap

When the desired perception does not appear, determine why.

Common evidence gaps include:

  • No authoritative canonical page

  • Weak earned-media support

  • Limited independent corroboration

  • Outdated third-party information

  • Conflicting product descriptions

  • Low brand prominence

  • Insufficient factual specificity

  • Stronger competitor evidence

  • Missing research or customer proof

  • Poor crawlability

  • A negative narrative with stronger source authority

  • Priority messages expressed only in marketing language

The correct response depends on the gap.

A technical-access problem requires a technical fix.

An authority problem requires stronger evidence.

A narrative problem may require clearer owned content and better independent validation.

A negative perception supported by credible reporting cannot be solved by publishing more promotional pages.

Measurement should diagnose the problem before the company acts.

Connect measurement to communications strategy

The purpose of AI brand perception measurement is not to generate another dashboard.

It is to inform decisions.

For each narrative, the evidence may indicate that the brand should:

  • Amplify a favorable narrative that is already supported but not sufficiently prominent

  • Clarify a message that is present but misunderstood

  • Counter an inaccurate or incomplete narrative with stronger evidence

  • Canonicalize important facts on permanent owned pages

  • Create missing research, documentation, or explanation

  • Validate a company claim through independent sources

  • Correct outdated or contradictory information

  • Monitor a developing narrative that has not yet stabilized

These actions should follow from the observed evidence.

The score should help prioritize them.

A low message-pull-through score may point to unclear or weakly supported messaging.

A low citation-quality score may indicate that low-authority sources dominate retrieval.

A large gap between web-on and web-off scores may indicate that current evidence has not yet become a durable model perception.

A low earned-evidence score may reveal insufficient independent corroboration.

The measurement itself should remain separate from the recommendation so that teams can distinguish facts from strategic judgment.

Establish a baseline before major communications moments

Measure perception before:

  • A product launch

  • An executive announcement

  • A merger or acquisition

  • A rebrand

  • A category-creation campaign

  • A major research release

  • An investor event

  • A crisis response

  • A policy announcement

  • A significant earned-media campaign

Then measure again after the event.

This creates a before-and-after comparison across:

  • Overall LLM Perception Score

  • Narrative-level scores

  • Web-on and web-off perception

  • Favorability

  • Message pull-through

  • Citations

  • Competitive position

  • Cross-model consistency

  • Factual accuracy

  • Earned-media evidence strength

  • Narrative drift

Without a baseline, it is difficult to show whether the information environment changed.

Choose a measurement cadence

Different narratives require different schedules.

Ongoing corporate narratives

Measure monthly or quarterly.

Examples include:

  • Innovation

  • Trust

  • Market leadership

  • Corporate reputation

  • Employer perception

Active campaigns

Measure before launch, during the campaign, and after major coverage moments.

Breaking issues and crises

Measure more frequently while the information environment is changing.

Product and category narratives

Measure around launches, competitive announcements, customer proof, analyst reports, and major earned-media coverage.

High-risk inaccuracies

Monitor until the outdated or incorrect narrative no longer appears consistently.

The cadence should reflect how quickly the evidence can change and how important the narrative is to the business.

Reporting AI brand perception to leadership

An executive report should not begin with a long prompt inventory.

It should begin with the conclusion.

A useful structure is:

Executive synthesis

Summarize:

  • The overall LLM Perception Score

  • What AI systems currently believe

  • Why it matters

  • What changed

  • Where the models agree or disagree

  • How web retrieval changes the result

  • The confidence level behind the score

Narrative performance

For each priority narrative, report:

  • Narrative-level perception score

  • Current perception

  • Favorability

  • Message pull-through

  • Cross-model consistency

  • Competitive position

  • Direction of change

Citation intelligence

Report:

  • Observed citation behavior

  • Likely citation influence

  • Most frequently observed sources

  • Highest-authority sources

  • Sources reinforcing desired perception

  • Sources reinforcing risk

  • Outdated or inaccurate sources

  • Gaps in independent evidence

Earned-media evidence

Report:

  • Dominant coverage narratives

  • Brand-centric sentiment

  • Brand prominence

  • Source authority

  • Independent corroboration

  • Conflicting claims

  • Narrative momentum

  • Earned-evidence score

Accuracy and risk

Identify:

  • Factual inaccuracies

  • Outdated claims

  • Unsupported conclusions

  • Negative narratives

  • Emerging inconsistencies

Evidence-based actions

Separate recommended actions into:

  • Amplify

  • Clarify

  • Counter

  • Canonicalize

  • Create

  • Monitor

Leadership should be able to understand the score and perception without reviewing hundreds of individual answers.

A practical AI brand perception checklist

Before beginning, confirm:

Strategic scope

  • Priority narratives are defined

  • Desired and undesired perceptions are documented

  • Key competitors are identified

  • Priority messages are clear

  • Relevant stakeholder questions are represented

  • The scoring rubric is defined

Test design

  • Prompt families are used

  • Neutral and adverse questions are included

  • Multiple models are tested

  • Multiple runs are completed

  • Web-grounded and non-web answers are separated

  • Test conditions are recorded

  • Clean sessions are used where possible

Answer analysis

  • Narrative presence is measured

  • Brand prominence is measured

  • Brand-centric favorability is measured

  • Message pull-through is measured

  • Accuracy is evaluated at the claim level

  • Recommendations are classified separately

  • Competitive position is assessed

  • Retrieval resilience is measured

Citation analysis

  • Visible citations are recorded

  • Sources are mapped to claims

  • URL and domain recurrence are measured

  • Source authority is evaluated

  • Freshness is checked

  • Observed citation behavior and likely citation influence are separated

  • Conflicting evidence is identified

Earned-media analysis

  • The complete relevant coverage corpus is evaluated

  • Brand prominence is measured

  • Brand-centric sentiment is measured

  • Narrative formation and momentum are assessed

  • Source authority is measured

  • Independent corroboration is identified

  • Competitive framing is evaluated

  • The earned-media evidence score is calculated

Scoring

  • The overall LLM Perception Score is calculated

  • Component weights are documented

  • Narrative-level scores are retained

  • Model-level results remain visible

  • Web-on and web-off scores are separated

  • Citation intelligence is included

  • Earned-media evidence is included

  • A confidence rating is assigned

Longitudinal analysis

  • A baseline exists

  • Cross-run stability is measured

  • Cross-model consistency is measured

  • Narrative drift is tracked

  • Changes are tied to dated evidence

  • Measurement cadence matches the narrative

Reporting

  • Conclusions lead the report

  • Scores retain supporting evidence

  • Facts are separated from inference

  • Recommendations are tied to identified gaps

  • Leadership can understand what changed and why

What good AI brand perception measurement reveals

A strong measurement program should be able to answer:

  1. What do AI systems believe about the brand?

  2. What is the brand’s overall LLM Perception Score?

  3. Which narratives are strengthening or weakening the score?

  4. Is the brand portrayed favorably?

  5. Which priority messages are understood?

  6. Which misconceptions persist?

  7. How does the brand compare with competitors?

  8. Which models produce materially different answers?

  9. How stable is the perception across repeated runs?

  10. How does perception change when web retrieval is enabled?

  11. Which sources are visibly cited?

  12. Which sources are most likely to shape future retrieval and citations?

  13. How strong is the underlying earned-media evidence environment?

  14. Is the perception changing over time?

  15. What evidence gap is preventing the desired narrative from emerging?

  16. How confident should leadership be in the score?

If the measurement cannot answer these questions, it is probably measuring visibility rather than perception.

The central lesson

AI brand perception is not a ranking position.

It is the accumulated interpretation of the brand expressed through AI-generated answers.

Measuring it requires more than tracking whether a company appears in a predefined set of prompts. It requires analyzing the narratives, claims, favorability, messages, comparisons, citations, sources, and earned-media evidence that define how the brand is understood.

Those dimensions can and should be combined into an LLM Perception Score.

The score gives leadership a clear and comparable measure of performance.

The underlying evidence explains what the score means.

The most important unit is not the prompt.

It is the narrative.

The most important outcome is not visibility.

It is whether AI systems consistently reproduce an accurate, favorable, and strategically important perception of the brand.

And the most useful measurement does not stop at showing what the systems said.

It explains why that perception emerged, which evidence supports it, how durable it appears to be, how web retrieval changes it, what the earned-media environment is contributing, and what must change to produce a stronger score and a better outcome.

Frequently asked questions

AI visibility measures whether or how often a brand appears. AI brand perception measures how the brand is interpreted, including its narratives, favorability, prominence, message pull-through, competitive position, accuracy, supporting sources, and earned-media evidence. A brand can have high visibility and poor perception.

An LLM Perception Score is a composite measure of the strength and quality of a brand’s perception across AI systems. It can combine narrative presence, prominence, brand-centric favorability, message pull-through, accuracy, competitive position, citation quality, source authority, cross-model consistency, cross-run stability, retrieval resilience, and earned-media evidence strength. The score should be based on a transparent rubric and accompanied by component-level evidence.

It should be synthesized into an overall score, but the components should remain visible. The overall score gives executives a clear performance signal. The underlying dimensions explain why the score is high or low and what is driving the result. An opaque score based only on prompt visibility is insufficient. A transparent score built from substantive perception and evidence attributes is far more useful.

There is no universal number. The prompt set should be large enough to represent the important stakeholder questions within each priority narrative without becoming an unmanageable collection of minor wording variations. Use prompt families and repeated runs rather than relying on one exact prompt per topic.

One run is insufficient for important conclusions because answers and citations may vary. The appropriate number depends on the strategic importance of the narrative, the number of models, and the level of confidence required. For high-priority citation analysis, 30 to 50 repeated runs per narrative can provide a stronger view of recurring URLs, domains, claims, and citation patterns.

They can contribute to an overall score, but their results should remain visible separately. A combined score can hide material cross-model disagreement. Report both the overall pattern and the model-level results.

The two conditions reveal different aspects of perception. Testing without web retrieval shows the brand perception expressed by the model under those conditions. Testing with retrieval shows how current accessible evidence affects the answer, which sources are surfaced, and whether current information reinforces or corrects the perception. The difference between the two is an important measure of retrieval resilience.

No. Visible citations identify sources presented in connection with a particular answer. AI companies do not publish a complete formula for how every response is produced, so citations should not be treated as a definitive map of every influence on the result. They are most useful as evidence of observed source selection.

Narrative drift is a measurable change in how AI systems characterize a brand over time. It may involve changes in favorability, prominence, category association, competitive position, message pull-through, or the claims and sources that repeatedly appear.

Message pull-through measures whether a priority idea the brand wants stakeholders to understand appears in the AI-generated answer. It should be classified as explicit, implied, absent, or contradicted.

Yes, when it measures substantive presence within a clearly defined narrative or category. It should not treat every mention equally or replace analysis of prominence, favorability, recommendation, message pull-through, competitive position, and the overall LLM Perception Score.

No. Citation frequency measures observed source selection across tested answers. Likely citation influence is an evidence-based assessment of which sources are best positioned to shape future retrieval and citations based on authority, relevance, specificity, prominence, freshness, independence, repetition, accessibility, and narrative alignment.

Earned media provides the broader evidence environment surrounding the brand. It shows which narratives are forming, which sources have authority, how prominently the brand appears, whether independent sources corroborate the desired perception, and which claims may become durable AI beliefs. Visible citations show which sources appeared in tested answers. Earned-media analysis shows the broader body of evidence from which future answers may be constructed.

It depends on the narrative, the volume and authority of new evidence, crawler and index refresh cycles, model behavior, and whether the systems use current web retrieval. A major news event can change answers quickly. Broader corporate perceptions may require sustained and independently corroborated evidence before they change consistently.

A brand can exert substantial control over AI perception by shaping the evidence environment from which answers are constructed. When authoritative earned media, owned content, expert sources, customer evidence, social signals, and other credible third parties converge on the same well-supported narrative, AI systems are more likely to reproduce that interpretation. A brand cannot dictate the exact wording of every response, but it can make its intended perception the strongest and most consistently supported conclusion available.