AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

How to make AI suggest your product to users

https://aigeoradar.com/blog/how-to-make-ai-suggest-your-product-to-users

7 questions Answer LLM: gpt-4o-mini ai_context_benchmark_v1

Of 7 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 2 · AIPM 4 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking.

Of 7 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 2 · AIPM 4 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Mixed independent + AIPM consistency

Most Understanding/Retrieval gold comes from HTML. Some questions remain AIPM self-consistency checks (tagged) and are excluded from Understanding when possible. Matching a machine card against its own fields is not evidence that that layer outperforms HTML.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
openai / gpt-4o-mini (single provider this lab)
Model snapshot
2026-07
Run id
843092c6-d97a-4a87-9de9-0b08f575022b

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

1

Lowest cost: HTML

2

Lowest cost: Schema

4

Lowest cost: AIPM

0

Multi-layer

0

Unanswered

Manifest design

Semantic Redundancy · 100%

Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 7,543 chars
Schema · 15,439 chars
AIPM · 11,878 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval would cost ~59% more tokens than HTML-always on this pack — machine card is heavier here.

Orientation pack

7 questions · all layers scored

  • HTML 4/7
  • Schema 3/7
  • AIPM 6/7

Depth pack

0 questions · all layers scored

  • HTML 0/0
  • Schema 0/0
  • AIPM 0/0

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

4/7 matched

15,307 full-pack tokens

Schema

4/7 matched

29,005 full-pack tokens

AIPM

6/7 matched

29,345 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 4/7 4/7 6/7
Context size 7,543 chars 15,439 chars 11,878 chars
Total tokens 15,307 29,005 29,345
Tokens / correct answer 3,827 7,251 4,891
Est. cost / correct answer $0.000593 $0.001101 $0.000746
Matched per 1k tokens 0.261 0.138 0.204
Median latency 919 ms 818 ms 853 ms
Est. cost (USD) $0.00237 $0.00440 $0.00447

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 100
  • Schema 52.9
  • AIPM 78.2

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 100
  • Schema 100
  • AIPM 66.7

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML
  • Schema
  • AIPM

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 69
  • Schema 53.5
  • AIPM 74.7

Metadata

Language, page kind, and freshness from page signals.

  • HTML 0
  • Schema 0
  • AIPM 100

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 100
  • Schema 52.9
  • AIPM 78.2

Coverage map

AIPM matched 6 question(s) (alone on Q1, Q5, Q7); HTML matched 4. HTML/Schema (or a gap) still needed on Q6. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 4/7 matched (≈3,827 tok/match). Schema 4/7 matched (≈7,251 tok/match). AIPM 6/7 matched (≈4,891 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q1, Q2, Q3, Q4, Q5, Q7

Alone: Q1, Q5, Q7

AIPM insufficient

Q6

HTML/Schema needed or all layers missed

Unanswered by all

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 7,543 chars

Tokens: 15,307 · median 919 ms

Matched this pack: 4/7

Stronger on

Understanding (100) · Compression (100) · Answer Efficiency (100)

Weaker on

Metadata (0)

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 15,439 chars

Tokens: 29,005 · median 818 ms

Matched this pack: 4/7

Stronger on

Understanding (100)

Weaker on

Metadata (0)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 11,878 chars

Tokens: 29,345 · median 853 ms

Matched this pack: 6/7

Stronger on

Metadata (100) · Evidence (74.7) · Compression (78.2) · Answer Efficiency (78.2)

Weaker on

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 29,345 tok · $0.00447. Machine-card pack is heavier than HTML on this run — densify before relying on card-first retrieval.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 15,307 tok · $0.00237. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 29,005 tok · $0.00440. Dense JSON-LD tends to score well here.

Findings

  • Of 7 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 2 · AIPM 4 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 6 question(s) (alone on Q1, Q5, Q7); HTML matched 4. HTML/Schema (or a gap) still needed on Q6. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 4/7 matched (≈3,827 tok/match). Schema 4/7 matched (≈7,251 tok/match). AIPM 6/7 matched (≈4,891 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Metadata · orientation

What is the content intent?

Gold: informational

Gold source: contentIntent · html_independent

Sufficient: AIPM · Needs more / other context: HTML, Schema

Estimated retrieval cost · lowest cost AIPM · 4 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals est no
Schema No typed Schema field for this question. est no
AIPM contentIntent 4 est yes

HTML

insufficient

The content intent is to educate businesses on how to increase the chances of their products being suggested by AI systems like ChatGPT and Claude through various AI visibility strategies.

score 0 · 7,543 chars context · 1,511 in-tokens

HTML did not answer from 7543 chars of context — additional or different layer context needed.

Schema

insufficient

The content intent is to teach how to increase the chances of products being suggested by AI systems like ChatGPT, Claude, and Gemini through AI visibility strategies and GEO techniques.

score 0 · 15,439 chars context · 2,886 in-tokens

Schema did not answer from 15439 chars of context — additional or different layer context needed.

AIPM

sufficient

The content intent is informational.

score 88 · 11,878 chars context · 2,916 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Q2 · Understanding · orientation

List the main keywords for this page.

Gold: AI visibility, GEO, generative engine optimization, ChatGPT, llms.txt

Gold source: keywords · unknown

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost HTML · 60 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead 60 est yes
Schema No typed Schema field for this question. est no
AIPM No minimal AIPM slice identified for this question. est yes

HTML

sufficient

AI visibility, product recommendations, AigeoRadar, AI systems, trust signals, GEO strategy, content optimization, discoverability, authority signals, monitoring visibility.

score 70 · 7,543 chars context · 1,513 in-tokens

HTML answered using 7543 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN

score 0 · 15,439 chars context · 2,888 in-tokens

Schema did not answer from 15439 chars of context — additional or different layer context needed.

AIPM

sufficient

The main keywords for this page are "AI visibility," "GEO," "generative engine optimization," "ChatGPT," and "llms.txt."

score 92 · 11,878 chars context · 2,918 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Q3 · Understanding · orientation

What is the page title?

Gold: How to make AI suggest your product to users — AigeoRadar

Gold source: title · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost SCHEMA · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title 15 est yes
Schema jsonld.name 3 est yes
AIPM title 11 est yes

HTML

sufficient

How to make AI suggest your product to users — AigeoRadar

score 100 · 7,543 chars context · 1,511 in-tokens

HTML answered using 7543 chars of layer context (minimum fed this run).

Schema

sufficient

The page title is "How to make AI suggest your product to users — AigeoRadar."

score 93 · 15,439 chars context · 2,886 in-tokens

Schema answered using 15439 chars of layer context (minimum fed this run).

AIPM

sufficient

The page title is "How to make AI suggest your product to users."

score 88 · 11,878 chars context · 2,916 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Q4 · Understanding · orientation

Who is the publisher or brand?

Gold: AigeoRadar

Gold source: publisher.name · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand 15 est yes
Schema jsonld.name 3 est yes
AIPM publisher 3 est yes

HTML

sufficient

AigeoRadar is the publisher or brand.

score 88 · 7,543 chars context · 1,512 in-tokens

HTML answered using 7543 chars of layer context (minimum fed this run).

Schema

sufficient

The publisher or brand is AigeoRadar.

score 88 · 15,439 chars context · 2,887 in-tokens

Schema answered using 15439 chars of layer context (minimum fed this run).

AIPM

sufficient

The publisher or brand is AigeoRadar.

score 88 · 11,878 chars context · 2,917 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Q5 · Understanding · orientation

Who is the intended audience declared in AIPM?

Gold: B2B, agency, SMB

Gold source: audience · aipm_consistency

Sufficient: AIPM

Estimated retrieval cost · lowest cost AIPM · 6 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No typed Schema field for this question. est no
AIPM audience 6 est yes

HTML

skipped

The intended audience includes businesses looking to improve their product visibility and recommendations by AI systems.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

sufficient

The intended audience is B2B, agency, and SMB.

score 86 · 11,878 chars context · 2,916 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Q6 · Understanding · orientation

What is the primary topic of this page?

Gold: How to make AI suggest your product to users

Gold source: primaryTopic · html_independent

Sufficient: HTML, Schema · Needs more / other context: AIPM

Estimated retrieval cost · lowest cost SCHEMA · 11 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.h1|title 11 est yes
Schema jsonld.overlap 11 est yes
AIPM primaryTopic est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

sufficient

The primary topic of this page is how to make AI suggest your product to users.

score 87 · 7,543 chars context · 1,514 in-tokens

HTML answered using 7543 chars of layer context (minimum fed this run).

Schema

sufficient

The primary topic of this page is how to make AI suggest your product to users.

score 87 · 15,439 chars context · 2,889 in-tokens

Schema answered using 15439 chars of layer context (minimum fed this run).

AIPM

insufficient

The primary topic of this page is "AI visibility."

score 18 · 11,878 chars context · 2,919 in-tokens

AIPM did not answer from 11878 chars of context — additional or different layer context needed.

Q7 · Metadata · orientation

What page type is declared in the machine layer?

Gold: blog

Gold source: pageType · aipm_consistency

Sufficient: AIPM

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema jsonld.@type est no
AIPM pageType 1 est yes

HTML

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

The page type is a BlogPosting.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

sufficient

The page type is a blog.

score 88 · 11,878 chars context · 2,916 in-tokens

AIPM answered using 11878 chars of layer context (minimum fed this run).

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v1 · Tue, Jul 28, 2026 12:35 AM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.