GEO Guides AI Brand Perception

Why Prompt Monitoring Alone Is Not Enough for AI Reputation Intelligence

AI reputation is shaped by narratives, sources, and repeated model behavior, not a handful of prompts selected in advance.

Iceberg above and below the waterline

Most AI-monitoring tools work by running a fixed list of prompts against a handful of models and reporting what comes back. It is a reasonable starting point — and it is nowhere near a complete picture of how a brand is actually perceived.

The answer

Prompt monitoring only captures how models respond to a small, pre-selected set of questions at a single point in time. AI perception is actually shaped by the underlying narrative — the recurring claims, sources, and citation patterns a model has learned across many related questions. A durable measurement approach has to track that narrative layer directly, not just sample outputs against a fixed prompt list.

Key takeaways

  • A fixed prompt list only ever samples a narrow slice of how a model can be asked about a brand.
  • Model answers are downstream of a narrative layer — the claims and sources a model has converged on — not the other way around.
  • Citation behavior (which sources a model actually references) is a more stable signal than any single answer.
  • Earned media still matters — it is frequently the material a model is grounding its answers in.
  • Measurement that repeats across prompt variations is more reliable than any one-off test.

What prompt monitoring measures

To see the gap clearly, it helps to be precise about what a prompt-monitoring tool is actually doing when it reports a result.

A typical setup selects a small set of representative prompts — "What is [brand] known for?", "Is [brand] a trustworthy company?" — and runs them against a handful of frontier models on a fixed schedule. Whatever the model says back becomes the reported "AI sentiment" or "AI visibility" score for that week.

This works reasonably well as a smoke test. It is fast, it is easy to explain in a slide, and it will reliably catch the most obvious failure states, like a model confusing a brand with a competitor or repeating an outdated fact. The problem is that it treats the model's answer as the object of measurement, when the answer is really just one sample from something larger.

What it misses

  1. It samples too narrow a slice of the question space. Real users ask about a brand in hundreds of different phrasings, from hundreds of different angles — as a regulator would, as a journalist would, as a customer comparing options would. A fixed prompt list can only ever cover a handful of these framings.

  2. It treats every answer as independent. Model answers about the same brand are correlated — they tend to converge on the same handful of claims and sources. Measuring answers one at a time misses that convergence, which is usually the more important signal.

  3. It ignores where the model is getting its information. Two answers can sound similar on the surface but be grounded in very different source material — one in current, accurate reporting, the other in a stale or unreliable source. A prompt-level score can't tell these apart.

Common pitfall

A brand can score well on a fixed prompt list while an unrelated, higher-risk narrative is actively forming around it in coverage the prompt list never asks about.

A narrative cluster: the same underlying claim recurring across otherwise unrelated coverage and model answers.

Why narratives shape AI perception

Models don't reason about a brand from first principles each time they're asked. They draw on patterns learned from a large, overlapping body of text — which means the same underlying narrative tends to surface no matter how the question is framed.

This is the part that a prompt-by-prompt view can't show: ask the same underlying question ten different ways, and a model will usually converge on the same handful of claims and the same handful of sources, even when the wording, the persona asking, and the specific angle all change. That convergence — not any single answer — is the real unit of AI perception.

The unit of AI perception was never the single answer. It was always the narrative the model has learned to reach for.

— Handraise research team

In practice, this means two brands can look identical on a fixed prompt list and still be in very different positions. One brand's convergent narrative might be grounded in current, favorable reporting. The other's might be grounded in a single outdated article that keeps resurfacing — invisible to a tool that only checks whether the model's tone sounds positive.

~30

model queries typically run per narrative cluster before a signal is treated as stable

Handraise methodology

4

stakeholder personas used in adversarial testing — regulator, journalist, investor, customer

Handraise methodology

3

layers in the evidence hierarchy: earned media, grounded citation, no-retrieval behavior

Handraise methodology

Single-prompt view: five isolated answers, no visible pattern.
Narrative-cluster view: the same five answers, grouped by underlying claim.

The role of earned media and citation behavior

If narratives are the real unit of measurement, the next question is where a model's version of a narrative actually comes from.

Frontier models are frequently able to retrieve and cite live sources when answering brand-related questions, rather than relying purely on training data. That means earned media — the actual body of published coverage about a brand — is often doing more work in shaping a model's answer than any amount of prompt engineering. Three claims illustrate the pattern:

Claim: Grounded answers cite specific, datable sources.

When a model is asked a brand question it can retrieve for, it will frequently name the outlet and approximate date of the source it's drawing from — a strong signal that earned media is directly shaping the answer. Source: Model citation logs collected during Handraise's citation-cycle testing.

Claim: No-retrieval answers default to whatever narrative is most repeated in training data.

This is why a single outdated but widely-syndicated story can keep surfacing in model answers long after it has been corrected or superseded. Source: Comparative testing across retrieval-enabled and retrieval-disabled model configurations.

Claim: Coverage volume alone does not predict citation behavior.

A brand with fewer, more authoritative articles can be cited more consistently than a brand with a much larger volume of lower-authority coverage. Source: Cross-brand comparison of coverage volume vs. citation frequency.

We used to report on tone. Now we can tell you which three articles a model is actually pulling its answer from — and that changed which corrections we prioritized.

— VP, Corporate Communications · Enterprise financial services company

Question Prompt monitoring Narrative intelligence
What's measured Answers to a fixed prompt list The claims and sources a model converges on across many prompt variations
Coverage of question space Narrow — limited to prompts chosen in advance Broad — adversarial variation across stakeholder personas
Source visibility Not typically shown Sources and citation behavior tracked directly
Stability of signal Can shift with small wording changes Averaged across repeated runs per cluster
Citation frequency by source authority tier, aggregated across a sample of narrative clusters.

Data and methodology

Analysis period. Rolling 90-day window, refreshed continuously.

Sample. Narrative clusters run in batches of 10 queries, up to ~30 per cluster before a signal is treated as stable.

Data sources. Earned media corpus, model citation logs, no-retrieval baseline responses.

Method. Adversarial prompt variation across four stakeholder personas, compared against a no-retrieval control.

Limitations: Model behavior can change without notice as providers update retrieval and ranking systems; findings reflect the tested window and are refreshed on a rolling basis rather than treated as permanent.

A better measurement model

Put together, this points to a three-layer way of assessing AI perception instead of a single prompt-response score.

  1. Earned media corpus. The full body of published coverage a model could plausibly retrieve from — the raw material the rest of the model's behavior is measured against.

  2. Grounded citation behavior. What a model actually cites when it can retrieve — which sources it reaches for, how current they are, and how consistently they recur across prompt variations.

  3. No-retrieval behavior. What a model defaults to when it can't retrieve — a useful baseline for spotting when an outdated or one-sided narrative has settled into a model's default understanding.

None of these three layers is meaningful on its own. A large earned-media footprint with poor citation behavior just means the material exists but isn't being surfaced. Strong grounded citation with a troubling no-retrieval baseline means the model behaves well when it can check its work, and reverts to an old narrative when it can't. The combination is what tells a communications team where to focus.

Recommendation

Treat all three layers as one system. A change in any single layer is a signal worth investigating, not a full picture on its own.

What communications teams should do next

Why it matters. If AI answers are increasingly a first stop for stakeholders forming an opinion about a brand, the narrative a model has converged on is functionally part of the brand's reputation — whether or not anyone signed off on it.

What communications teams should do

  • Track citation behavior, not just answer tone, across a wide range of prompt framings.

  • Keep primary source material current — it is frequently what a model retrieves from directly.

  • Re-test after any correction to confirm the no-retrieval baseline has actually shifted, not just the grounded answer.

  • Review narrative clusters on a recurring cadence rather than treating any single test as final.

Amplify, clarify, counter, create

Amplify. Reinforce accurate, favorable narratives that are already gaining citation traction.

Clarify. Add current, authoritative material where a model's answer is technically correct but incomplete.

Counter. Directly address a narrative that is actively inaccurate or out of date before it becomes the default.

Create. Publish new primary material where no authoritative source currently exists for a model to cite.

What to watch next. As models rely more heavily on live retrieval, the gap between "what's published" and "what's cited" will likely keep narrowing — making earned-media accuracy and structured primary source material more directly consequential for AI perception, not less.

Frequently asked questions

No — it's a reasonable smoke test and can catch obvious errors quickly. The issue is treating it as a complete measurement system rather than one input into a broader view of narrative and citation behavior.

Handraise runs this on a rolling basis, since model behavior can shift as providers update retrieval and ranking systems. A one-time test only reflects that specific moment.

No. Earned media is one of the three layers in the model above — it's the raw material a lot of citation behavior is grounded in, so traditional monitoring stays essential, not optional.

In summary

Prompt monitoring measures answers. AI reputation is shaped by narratives — the claims and sources a model converges on across many ways of asking. A durable measurement approach tracks earned media, grounded citation behavior, and no-retrieval behavior together, on a recurring basis, rather than relying on any single prompt-response check.

Sources

  1. Handraise citation-cycle testing notes — Internal methodology, illustrative. Handraise analysis — not a customer case study.

  2. Comparative retrieval-enabled vs. retrieval-disabled testing — Handraise research, illustrative window. Our interpretation of observed citation patterns.

  3. Coverage volume vs. citation frequency comparison — Cross-brand sample, illustrative. Illustrative example for product education.