AI Search Proof Lab

Measured vs. Estimated:A GEO Prioritization Model That Shows Its Work

Fifteen tracked AI prompts, scored twice. What the tracking layer measured, what I judged, and what to do when a source you do not own controls the answer.

By Ellen Tuckett · AI Search, AEO, GEO & SEO Strategy · August 2026 · v1.0
Scored against the four-week Cloudflare dataset · Otterly.AI · United States · June 12 to July 9, 2026
Where this data comes from. Between June 12 and July 9, 2026, I ran a four-week public GEO measurement experiment using Cloudflare as the test brand. The same fixed library of 15 buyer-intent prompts was tracked weekly in Otterly.AI across all engines in the United States, producing four weekly reports. This page takes that dataset and asks a different question: given the numbers, what should a team actually work on next? If you want the measurement build first, start with the Cloudflare GEO Measurement System.
The argument in one line

Every GEO priority score is part measurement and part judgment. Averaging them into one number makes the judgment half invisible, and teams then act on it as if it were data.

This page keeps them apart, scores all 15 tracked prompts, and ends with the part most prioritization frameworks skip: what to actually do when the answer belongs to a page you will never own.

Who this is for: SEO, GEO, and AEO practitioners who have to defend a content roadmap to stakeholders, and marketing leaders who want to know which half of a priority score they are trusting.

15
Prompts scored on both measured and estimated columns
8 of 15
Prompts where Cloudflare is mentioned but supplies under 2% of citations
4
Ways to win when a third party owns the answer, only one of which is content

Note: Measured columns come from Otterly.AI, United States, across a fixed 15-prompt library, comparing the Week 1 window (June 12 to 18, 2026) with the Week 4 window (July 3 to 9, 2026). Estimated columns are my judgment, labeled as judgment, with the basis stated. No revenue data was available for Cloudflare, which is the point.

The problem with a single priority score

Run a citation gap audit and you get a list. Every row looks like a problem. To make it actionable, most practitioners collapse it into a priority score: some weighting of visibility, intent, and revenue proximity, ranked high to low.

Half of that score's inputs are measured. Presence, position, which sources are cited, stability over time. Those are facts a tool produced. The other half is estimated. Buyer stage, purchase intent, how close a prompt sits to revenue, whether an answer is winnable at all. Those are judgment calls a person made.

Averaging them produces one confident-looking number in which nobody can tell which half is which.

The failure mode: a stakeholder reads the composite, anchors on the ranking, and makes a roadmap decision based mostly on the estimated half without knowing that is what they are doing. The number never says "this part was a guess," so nobody asks.

This matters more in GEO than in classic SEO, because the estimated half is weaker here. AI-referred traffic is badly instrumented. Much of it lands as direct. Someone who reads an AI answer about a brand on Tuesday and searches that brand on Friday shows up as branded search, not AI influence. Any revenue number attached to a citation gap is an inference and should look like one on the page.

The model: two columns, never combined

The fix is structural, not statistical. Keep measured inputs and estimated inputs in visibly separate columns, label the estimates as estimates, and refuse to produce a composite.

Measured inputs

Everything here comes from the tracking layer. If a tool did not produce it, it does not belong in this group.

InputSourceWhat it tells you
Brand mentionsPrompt trackingWhether the brand appears in the answer, and how often
Average positionPrompt trackingWhere the brand sits when it does appear
Owned citation ratePrompt citations divided by total citationsHow much of the answer the brand's own pages actually supply
Answer ownerPrompt citations by URLWhich specific page is shaping the answer
StabilityWeek-over-week comparisonWhether a gap persists or is noise in a single pull
Owned citation rate is the input most programs never calculate, and it is the one that separates presence from authorship. A brand can be named in every answer while contributing almost none of the material those answers are built from. Mentions alone cannot show that. This ratio can.

Estimated inputs

Everything here is judgment. Each carries a stated basis and a confidence flag, so a reader can disagree with the reasoning rather than the number.

InputBasisWhy it is an estimate
Buyer stagePrompt language and intentInferred from phrasing, not from observed buyer behavior
Purchase intentPrompt language and category normsA comparison prompt usually signals intent, but not always
Revenue proximityBuyer stage plus intentAI-referred traffic is poorly attributed, so this cannot be measured directly
DisplaceabilityWho owns the answer and why they winA judgment about whether an incumbent source can realistically be replaced
ConfidenceStrength of the aboveAn explicit signal about how much weight the estimate should carry

The rule

No composite. The output is a table a person reads, not a number a person sorts by. If a stakeholder wants a ranking, they get one built by hand from both columns, with the reasoning written down.

All 15 prompts, scored both ways

Measured columns are from Otterly.AI. Owned citation rate is Cloudflare-domain citations divided by total citations for that prompt in the Week 4 window. Estimated columns are my judgment.

Group 1: Cited and sourced

The brand is mentioned heavily and its own pages supply a meaningful share of the citations. There is no gap here. The correct action is to protect and monitor, which is worth saying out loud because high-scoring rows attract work they do not need.

PromptMentions W1 → W4 (M)Owned citation rate W4 (M)Revenue proximity (E)Action
Which platform offers unmetered DDoS and global CDN?8 → 2818.5%High, high confidenceProtect and monitor
Which Zero Trust service secures employees and SaaS apps?6 → 2218.4%High, high confidenceProtect and monitor
Which web application firewalls also include strong DDoS protection?8 → 259.5%High, medium confidenceProtect and monitor
Which enterprise solutions combine CDN, WAF, and DNS together?8 → 278.1%High, medium confidenceProtect and monitor

Group 2: Mentioned but not sourced

The brand is named in the answer, often first, while its own pages supply almost none of the citations behind it. Under a composite score these read as healthy, because mentions are high. The owned citation rate is what exposes them.

PromptMentions W1 → W4 (M)Owned citation rate W4 (M)Displaceability (E)Action
How can I reduce origin egress costs with edge caching?4 → 121.1%LowDistribution and third-party authority
What are the best global CDNs for high-traffic sites?8 → 270.7%MediumGet into the listicles that own the answer
What are the most reliable edge compute platforms for developers?8 → 220.7%MediumGet into the listicles that own the answer
What platform helps deploy low-latency global applications without servers?8 → 250.8%MediumOwned citation rate fell from 12 citations to 2. Investigate first
What edge compute platform charges only for CPU time?2 → 200.4%MediumPublish the pricing specifics the answer needs
How do leading SASE platforms compare on Zero Trust features?2 → 170.6%MediumBecome the source the comparisons cite
How do top Zero Trust access platforms replace legacy VPNs?4 → 161.2%MediumBecome the source the comparisons cite
How do I protect my production apps from massive DDoS?6 → 204.2%HighClosest to winnable in this group. Start here

Group 3: Absent

The brand barely appears or does not appear at all. These are the only true content gaps in the set, and only one of them is worth a quarter.

PromptMentions W1 → W4 (M)Total citations W4 (M)Revenue proximity (E)Action
What are the best platforms to secure AI agents?2 → 0263High, medium confidenceCategory entry. The clearest opportunity in the set
How can I secure users, devices, and data on one network?0 → 0252Medium, low confidenceAbsent both weeks. Diagnose before committing
What should I use if my VPN concentrator is overloaded?0 → 3278Low, low confidenceProblem-aware, low intent. Low priority

(M) = measured. (E) = estimated. Owned citation rate = Cloudflare-domain citations ÷ total citations for that prompt, Week 4 window.

What the split actually changes

Case 1: The row a composite score gets wrong

On How can I reduce origin egress costs with edge caching? in Week 4, Cloudflare was mentioned 12 times and ranked first. Under a blended score, that plus a purchase-adjacent prompt puts the row near the top of the list, and a content team writes a page.

The measured column tells a different story. Of the 280 citations behind those answers, Cloudflare's own domain supplied 3. The top cited page was egresscost.com at 20 citations, and two Fastly blog posts contributed 10 and 9. A direct competitor was supplying more than three times the source material Cloudflare was.

This is a distribution problem wearing a content gap's clothes. Publishing another owned page does not change which source the answer retrieves. And it is not just an independent site winning here: Fastly earned its position with two focused blog posts on exactly this question.
Otterly.AI prompt-level citation list for the origin egress prompt, showing egresscost.com at 20 citations and two Fastly blog posts above any Cloudflare page.
Prompt citations by URL for How can I reduce origin egress costs with edge caching? egresscost.com leads at 20 citations, followed by two Fastly posts at 10 and 9. Cloudflare's best-performing page is a blog post at 6.
Source: Otterly.AI, all engines, United States, July 3 to 9, 2026.

Case 2: The row that is genuinely a content gap

On What are the best platforms to secure AI agents? Cloudflare went from 2 mentions in Week 1 to zero in Week 4. Across 263 citations, not one came from a Cloudflare domain. Palo Alto Networks led the source set with a single cyberpedia page at 13 citations, followed by SourceForge, agentsecurity.com, and a string of vendor comparison posts.

Almost every cited page in that set carried no brand mention at all. The category conversation is being written without the brand in it.

Otterly.AI prompt-level citation list for the AI agent security prompt, showing eleven cited URLs with no Cloudflare domain and Brand mentioned marked No on nearly every row.
Prompt citations by URL for What are the best platforms to secure AI agents? No Cloudflare domain appears in the cited set, and nearly every source is marked Brand mentioned: No.
Source: Otterly.AI, all engines, United States, July 3 to 9, 2026.

This is the one row in the set where a composite score and the split agree on urgency, and they agree because the measured column is unambiguous: total absence, stable across both weeks, in a category with 263 citations of demand and no dominant incumbent.

Case 3: The row where doing nothing is correct

On Which platform offers unmetered DDoS and global CDN? Cloudflare was mentioned 28 times, ranked first, and its own domain supplied 50 of 270 citations. Five owned URLs appear in the top of the cited set, led by the DDoS for Web product page at 8 citations.

Otterly.AI prompt-level citation list for the unmetered DDoS and global CDN prompt, showing Cloudflare's DDoS for Web page first and multiple owned URLs in the cited set.
Prompt citations by URL for Which platform offers unmetered DDoS and global CDN? Cloudflare's DDoS for Web page leads the cited set at 8 citations, with four more owned URLs appearing below it.
Source: Otterly.AI, all engines, United States, July 3 to 9, 2026.

In a ranked list, high scores attract work. The split makes the right answer obvious: there is no gap. The action is to protect and monitor, and the measured column is what makes that defensible when someone asks why the highest-scoring row has no tasks attached to it.

What to do when a source you do not own controls the answer

Naming a distribution problem is only half a recommendation. Group 2 above contains eight prompts, and for all of them the instinct to publish another owned page is the wrong first move. There are four real options, and only one of them is content.

1. Get into the page that is already winning

If a listicle or independent guide owns the answer, the goal is to be represented favorably inside it. That means outreach: correct an error, supply current specs or pricing, offer data, ask to be included in a comparison. This is public relations and partnership work repurposed, and it is the fastest available win. It is also the option most search teams skip because it does not feel like search.

Best fit: best global CDNs, most reliable edge compute platforms. Both are dominated by roundups where inclusion is an editorial decision, not a ranking one.

2. Become the source that the winning page cites

Look at where the incumbent got its information. If a page is winning on benchmark numbers, pricing comparisons, or performance data, publish the definitive version of that data. You get cited by the citer first, and directly later. Original data is the strongest form of this, which is also why a measurement experiment like this one functions as an asset rather than a blog post.

Best fit: the CPU-time pricing prompt and the SASE comparison prompt. Both turn on specifics a vendor is uniquely positioned to publish.

3. Widen the source set rather than the page count

If one independent site owns the answer, one more owned page will not move it. Five credible third-party appearances might. Analyst coverage, community threads, documentation on partner sites, review platforms, a well-cited developer forum answer. AI answers assemble from a spread of sources, and the goal is to change the spread.

Worth noting from this dataset: youtube.com and reddit.com both appear in Cloudflare's top cited domains across all four weeks. Community and video surfaces are already shaping these answers.

4. Decide it is not winnable and spend the quarter elsewhere

Sometimes the incumbent is a neutral third party that is more credible on that question than any vendor can be, precisely because it is not a vendor. Buyers trust it for the same reason. In that case the honest call is to stop, log the reasoning, and move the effort to a prompt that can be won.

Best fit: the VPN concentrator prompt. Low intent, problem-aware phrasing, and no clear route to owning the answer. Three mentions in Week 4 is not a signal worth a roadmap slot.

The diagnostic underneath all four: ask why the incumbent is winning. If it wins on freshness, publish better data. If it wins on neutrality, get inside it rather than competing with it. If it wins on depth, either match that depth somewhere credible or walk away. That single question determines which of the four options applies, and it is exactly the judgment a composite score buries.

Where this model breaks

Three honest limits, since a prioritization model that claims no weaknesses is doing the same thing a composite score does.

LimitWhy it mattersCurrent workaround
Two windows is thin This scoring compares Week 1 with Week 4. Two points can show direction but not volatility, and one prompt in the set swung from 12 owned citations to 2 with no clear cause Score direction and absence confidently. Treat single-prompt swings as questions to investigate rather than findings
Revenue proximity is unverifiable Without clean attribution the estimate cannot be checked against outcomes, so it never improves on its own Label it, flag confidence, and revisit when downstream data exists
The judgment does not transfer Someone inheriting this table sees a finished artifact and follows it, missing the calls that produced it Keep a decision log of rows where the obvious read and the final action disagreed, with one sentence on why
The third limit is the real one. The columns are documentation. The judgment about which gaps are stable, which sources are displaceable, and when to override the obvious read is tacit, and tacit knowledge is where systems die when the person who built them leaves.

How to run this on your own data

The model is brand-agnostic and tool-agnostic. It needs a tracking layer that produces the measured columns and a person willing to state the estimates as estimates.

StepWhat you doOutput
1. Establish a baselineRun the same fixed prompt library across consecutive weekly windows, same market, same enginesA measured baseline that can distinguish a gap from variance
2. Fill the measured columnsPull mentions, position, total citations, and owned citations per prompt, then calculate owned citation rateFacts, with a source attached to each
3. Identify the answer ownerFor every low owned-citation-rate prompt, open the prompt-level citation list and find which page is winningThe difference between a content gap and a distribution problem
4. Fill the estimated columnsAssign buyer stage, intent, revenue proximity, displaceability, and confidence, writing down the basis for eachJudgment, visibly labeled as judgment
5. Refuse the compositeDo not average the columns. Read them together and write the action by handA decision a person made, that a person can defend
6. Log the overridesRecord every row where the obvious read and the final action disagreed, plus one sentence on whyThe beginning of transferable judgment

Related download

Weekly GEO Measurement Starter Kit

The measured columns in this model come from a weekly tracking cadence. The Starter Kit provides the template and instructions for producing that data on any AI visibility tool.

Download the Starter Kit (Word)

Why this matters

GEO is at the stage where measurement is improving faster than the judgment applied to it. Tools now produce coverage, share of voice, position, and citation data reliably. What they do not produce, and cannot, is the decision about which gaps are worth an organization's next quarter.

That decision is where the value sits, and it is also where the dishonesty creeps in. A composite score is a way of making a judgment call look like a measurement, and it works right up until someone asks how the number was built.

Keeping the two apart costs a slightly less impressive slide and buys a defensible one. In this dataset it also changed the answer: eight of fifteen prompts that look healthy on mentions are being written by somebody else, and the fix for most of them is not content at all.

Related Proof Lab work

This model connects to the measurement work that produces its inputs and the workflows that act on its outputs.

FAQ

What is the difference between measured and estimated in a GEO priority score?

Measured inputs come from the tracking layer: brand mentions, average position, owned citation rate, which page owns the answer, and stability across pulls. Estimated inputs are judgment: buyer stage, purchase intent, revenue proximity, and whether an incumbent source can realistically be displaced. Both are necessary. Combining them into one number is what causes the problem.

What is owned citation rate and why does it matter?

Owned citation rate is the share of a prompt's total citations that come from your own domains. It matters because it separates being mentioned from being the source. In the Cloudflare dataset, eight of fifteen prompts showed strong mention counts alongside an owned citation rate under 2%, meaning the brand was named in answers built almost entirely from other people's pages.

How do you tell a content gap from a distribution problem?

Compare mentions with owned citation rate. High mentions plus a low owned citation rate is a distribution problem: the brand is in the answer but is not supplying it, and publishing another owned page rarely changes which source gets retrieved. Low mentions plus low citations is a genuine content gap. The two require different teams and different budgets.

What should you do when a third-party page owns the AI answer?

There are four options and only one is content. Get represented inside the page that is already winning through outreach. Become the source that page cites by publishing the underlying data. Widen the source set through analyst, community, partner, and review surfaces. Or decide the answer is not winnable and spend the quarter elsewhere. The choice depends on why the incumbent is winning: freshness, neutrality, or depth.

Why is revenue proximity always an estimate in GEO?

Because AI-referred traffic is poorly instrumented. Much of it arrives as direct, and delayed branded search from an AI answer is attributed to branded search rather than AI influence. Any revenue figure attached to a citation gap is an inference and should be presented as one.

Can this prioritization model be used with any AI visibility tool?

Yes. The measured columns require a tracking layer that reports brand mentions, position, total citations per prompt, and citations by URL. Otterly.AI, Profound, and comparable platforms produce those inputs. The model itself is tool-agnostic.