Measurement & Brand

Never Publish A Percentage Built From One Sample

Ask a generative platform the same question twice and the answers can differ. A figure produced by asking once, counting once and dividing does not reproduce.

XstraStar editorial team mark

XstraStar Editorial Team

GEO Research & Strategy

Never Publish A Percentage Built From One Sample
In this article
  1. The algorithm that breaks
  2. Why answers move
  3. The reproducible version
  4. Where one sample is enough

Ask a generative platform the same question twice and the answers can differ. A figure produced by asking once, counting once and dividing does not reproduce. It works for internal direction, and it does not belong in published material or in acceptance clauses.

The algorithm that breaks

mention rate = answers containing the brand name ÷ number of questions

In practice: 40 questions, one run each, 6 answers contain the brand, report 15%.

That figure breaks in three places.

PlaceWhy
Denominator40 is a count of questions, not of samples. With one run per question the two happen to be equal, which hides the distinction until someone adds a second run
NumeratorA single run is one draw from a distribution. Re-run the same batch and 6 becomes 4, or 9
PrecisionWith a denominator of 40, one answer is worth 2.5 percentage points, so a two-digit figure like 15% claims more resolution than the data holds

Why answers move

Four sources: decoding carries randomness, platforms return different retrieval results by session and region, the retrieval index keeps updating, and answer templates switch on small differences in phrasing.

None of these is a defect. All four are normal behaviour for this class of system, so the measurement method has to absorb them.

The reproducible version

mention rate = Σ(answers containing the brand) ÷ Σ(samples)      ← denominator counts samples
n ≥ 3 runs per question; weight by sample size across questions, never a plain average
label every report with: questions · runs per question · collection window · platform · region

Three rules travel with it:

  1. No plain averaging across platforms. Sample sizes differ, so a plain average over-weights the smallest set. Weight by sample size, and still list platforms separately.
  2. Daily-average basis, not cumulative. Cumulative coverage approaches 100% as the window grows, every vendor in a category converges there, and the metric stops discriminating.
  3. A day with an incomplete run is a gap, not a zero. Counting it as zero drags the average to a fraction of the real value.

Where one sample is enough

SituationIs one run enough
Internal direction-setting for the next content batchYes
Establishing whether a brand name ever appears for a questionYes
A published percentage or a baseline figure inside a proposalNo
A metric written into an acceptance clauseNo, and runs per question must be stated
Comparing two points in timeNo, and both points need identical sampling

⚠️ One more failure mode: platform-side matchers produce false positives. Answers with no brand string present have been flagged as mentions, so any externally used figure needs a human pass. A column where every verdict is identical (all "not mentioned" or all "mentioned") usually means the match never ran.


Key takeaway: Generative platforms can return different answers to the same question, so a mention rate built from one run per question does not reproduce and belongs only in internal direction-setting. The reproducible form takes at least 3 runs per question, divides by sample count instead of question count, weights across platforms by sample size instead of averaging, and labels every figure with questions, runs per question, collection window, platform and region. Report on a daily-average basis, record incomplete days as gaps instead of zeros, and have a human verify matcher output before any figure goes out, especially when every verdict in a column is identical.

Related: Prompt set designMention rateMeasurement disciplineFirst-party first

Sources: Algorithm and measurement rules from XstraStar monitoring practice (multiple projects across Chinese and English markets, first half of 2026). Matcher false positives from internal review records.

Last updated: 2026-08-27

Turn insight into growth

Find your next AI search growth opportunity

From brand visibility and citation sources to content strategy, the XstraStar team will help you map a clear, measurable GEO optimization path.

Talk to a GEO advisor
XstraStar editorial team mark

About the author

XstraStar Editorial Team

GEO Research & Strategy

The XstraStar editorial team studies AI search, generative engine optimization, and brand visibility, turning platform mechanics, field experience, and market shifts into practical growth guidance.

Related insights

Keep exploring GEO