AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

Dişli Malzeme Galvaniz

https://kapadokyagalvaniz.com.tr/disli-malzeme-galvaniz/

8 questions Answer LLM: gpt-4o-mini ai_context_benchmark_v1

Of 8 questions: lowest measured retrieval cost among sufficient layers — HTML 2 · Schema 1 · AIPM 3 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking.

Of 8 questions: lowest measured retrieval cost among sufficient layers — HTML 2 · Schema 1 · AIPM 3 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 64% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Mixed independent + AIPM consistency

Most Understanding/Retrieval gold comes from HTML. Some questions remain AIPM self-consistency checks (tagged) and are excluded from Understanding when possible. Matching a machine card against its own fields is not evidence that that layer outperforms HTML.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
openai / gpt-4o-mini (single provider this lab)
Model snapshot
2026-07
Run id
329665c9-117f-40b1-ab7c-8f79a29b6285

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

2

Lowest cost: HTML

1

Lowest cost: Schema

3

Lowest cost: AIPM

0

Multi-layer

2

Unanswered

Manifest design

Semantic Redundancy · 64%

Semantic Redundancy 64% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%
purpose keyFacts 46%
abstract keyFacts 46%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 6,497 chars
Schema · 1,579 chars
AIPM · 7,779 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval is roughly token-neutral vs HTML-always on this pack.

Orientation pack

8 questions · all layers scored

  • HTML 4/8
  • Schema 3/8
  • AIPM 5/8

Depth pack

0 questions · all layers scored

  • HTML 0/0
  • Schema 0/0
  • AIPM 0/0

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

4/8 matched

21,693 full-pack tokens

Schema

3/8 matched

5,662 full-pack tokens

AIPM

5/8 matched

23,678 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 4/8 3/8 5/8
Context size 6,497 chars 1,579 chars 7,779 chars
Total tokens 21,693 5,662 23,678
Tokens / correct answer 5,423 1,887 4,736
Est. cost / correct answer $0.000844 $0.000316 $0.000736
Matched per 1k tokens 0.184 0.530 0.211
Median latency 1,093 ms 782 ms 1,003 ms
Est. cost (USD) $0.00338 $0.00095 $0.00368

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 34.7
  • Schema 100
  • AIPM 39.8

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 25
  • Schema 50
  • AIPM 50

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML
  • Schema
  • AIPM

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 69.4
  • Schema 49.4
  • AIPM 62.3

Metadata

Language, page kind, and freshness from page signals.

  • HTML 100
  • Schema 0
  • AIPM 50

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 34.7
  • Schema 100
  • AIPM 39.8

Coverage map

AIPM matched 5 question(s) (alone on Q3); HTML matched 4. HTML/Schema (or a gap) still needed on Q1, Q4, Q5. No layer matched gold on Q1, Q5. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 4/8 matched (≈5,423 tok/match). Schema 3/8 matched (≈1,887 tok/match). AIPM 5/8 matched (≈4,736 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q2, Q3, Q6, Q7, Q8

Alone: Q3

AIPM insufficient

Q1, Q4, Q5

HTML/Schema needed or all layers missed

Unanswered by all

Q1, Q5

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 6,497 chars

Tokens: 21,693 · median 1,093 ms

Matched this pack: 4/8

Stronger on

Metadata (100)

Weaker on

Understanding (25) · Compression (34.7) · Answer Efficiency (34.7)

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 1,579 chars

Tokens: 5,662 · median 782 ms

Matched this pack: 3/8

Stronger on

Compression (100) · Answer Efficiency (100)

Weaker on

Metadata (0)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 7,779 chars

Tokens: 23,678 · median 1,003 ms

Matched this pack: 5/8

Stronger on

Weaker on

Compression (39.8) · Answer Efficiency (39.8)

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 23,678 tok · $0.00368.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 21,693 tok · $0.00338. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 5,662 tok · $0.00095. Dense JSON-LD tends to score well here.

Findings

  • Of 8 questions: lowest measured retrieval cost among sufficient layers — HTML 2 · Schema 1 · AIPM 3 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 64% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 5 question(s) (alone on Q3); HTML matched 4. HTML/Schema (or a gap) still needed on Q1, Q4, Q5. No layer matched gold on Q1, Q5. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 4/8 matched (≈5,423 tok/match). Schema 3/8 matched (≈1,887 tok/match). AIPM 5/8 matched (≈4,736 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Understanding · orientation

What is the purpose of this page?

Gold: Dişli Malzemeler İçin Galvanizleme Uygulamaları Dişli malzeme galvaniz, özellikle şu alanlarda yaygın olarak kullanılır.

Gold source: purpose · html_independent

Partial: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema jsonld.description est no
AIPM purpose est no

Root cause · Partial semantic match

Tighten primaryTopic/purpose/abstract so wording overlaps independent HTML gold.

HTML

partial

The purpose of this page is to provide information about galvanizing services for gear materials offered by Kapadokya Galvaniz.

score 55 · 6,497 chars context · 2,142 in-tokens

HTML was only partially aligned with gold from 6497 chars of context.

Schema

insufficient

UNKNOWN.

score 0 · 1,579 chars context · 544 in-tokens

Schema did not answer from 1579 chars of context — additional or different layer context needed.

AIPM

partial

The purpose of the page is to provide strong and durable solutions for Dişli Malzeme Galvaniz used in various industries such as machinery, automotive, and energy.

score 55 · 7,779 chars context · 2,339 in-tokens

AIPM was only partially aligned with gold from 7779 chars of context.

Q2 · Understanding · orientation

List the main keywords for this page.

Gold: Dişli Malzeme Galvaniz, disli, malzeme, galvaniz

Gold source: keywords · unknown

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 181 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead 181 est yes
Schema No typed Schema field for this question. est yes
AIPM No minimal AIPM slice identified for this question. est yes

HTML

sufficient

Dişli Malzeme, Galvaniz, Kapadokya Galvaniz, Korozyon Direnci, Uzun Ömür, Düşük Bakım Maliyeti, Güçlü ve Dayanıklı, Galvanizleme Süreci, Makine Endüstrisi, Otomotiv, Enerji Sektörü, Tarım Makineleri.

score 100 · 6,497 chars context · 2,142 in-tokens

HTML answered using 6497 chars of layer context (minimum fed this run).

Schema

sufficient

Dişli Malzeme Galvaniz, Kapadokya Galvaniz, sıcak daldırma galvaniz, santrifüj galvaniz, metal yüzey kaplama, bağlantı elemanları galvaniz, savunma sanayi galvaniz, enerji sektörü galvaniz, tamburlu yüzey hazırlama, PLC otomasyon sistemleri.

score 100 · 1,579 chars context · 544 in-tokens

Schema answered using 1579 chars of layer context (minimum fed this run).

AIPM

sufficient

The main keywords for this page are "Dişli Malzeme Galvaniz," "disli," "malzeme," and "galvaniz."

score 75 · 7,779 chars context · 2,339 in-tokens

AIPM answered using 7779 chars of layer context (minimum fed this run).

Q3 · Understanding · orientation

Who is the intended audience declared in AIPM?

Gold: Metal Kaplama & Galvaniz hizmetleriyle ilgilenen firmalar, şahıslar

Gold source: audience · aipm_consistency

Sufficient: AIPM

Estimated retrieval cost · lowest cost AIPM · 18 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No typed Schema field for this question. est no
AIPM audience 18 est yes

HTML

skipped

The intended audience includes industries such as machinery, automotive, energy, and agriculture that require galvanizing services for gear materials.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

sufficient

The intended audience includes "Metal Kaplama & Galvaniz hizmetleriyle ilgilenen firmalar" and "şahıslar."

score 82 · 7,779 chars context · 2,337 in-tokens

AIPM answered using 7779 chars of layer context (minimum fed this run).

Q4 · Metadata · orientation

What language is this page in?

Gold: tr

Gold source: inLanguage · html_independent

Sufficient: HTML · Needs more / other context: Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 2 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lang 2 est yes
Schema jsonld.block est no
AIPM inLanguage est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

sufficient

The page is in Turkish.

score 88 · 6,497 chars context · 2,141 in-tokens

HTML answered using 6497 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN

score 0 · 1,579 chars context · 543 in-tokens

Schema did not answer from 1579 chars of context — additional or different layer context needed.

AIPM

insufficient

The page is in English.

score 0 · 7,779 chars context · 2,338 in-tokens

AIPM did not answer from 7779 chars of context — additional or different layer context needed.

Q5 · Understanding · orientation

Summarize the page in one sentence.

Gold: Dişli Malzemeler İçin Galvanizleme Uygulamaları Dişli malzeme galvaniz, özellikle şu alanlarda yaygın olarak kullanılır.

Gold source: abstract · html_independent

Partial: HTML, Schema, AIPM

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema jsonld.description est no
AIPM abstract est no

Root cause · Partial semantic match

Tighten primaryTopic/purpose/abstract so wording overlaps independent HTML gold.

HTML

partial

Kapadokya Galvaniz offers galvanizing services for gear materials, enhancing their durability and corrosion resistance, crucial for various industries such as automotive and energy.

score 55 · 6,497 chars context · 2,144 in-tokens

HTML was only partially aligned with gold from 6497 chars of context.

Schema

partial

Kapadokya Galvaniz, founded in 2023 by Mahmut Şimşek and Mustafa Gümüş in Kahramankazan, specializes in hot dip centrifugal galvanizing and serves the Turkish market.

score 55 · 1,579 chars context · 546 in-tokens

Schema was only partially aligned with gold from 1579 chars of context.

AIPM

partial

Kapadokya galvaniz Dişli Malzeme Galvaniz, güçlü ve dayanıklı çözümler sunarak dişli malzemelerin makine, otomotiv ve enerji sektörlerinde kullanımını vurgulamaktadır.

score 55 · 7,779 chars context · 2,341 in-tokens

AIPM was only partially aligned with gold from 7779 chars of context.

Q6 · Metadata · orientation

What is the content intent?

Gold: commercial

Gold source: contentIntent · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals 100 est yes
Schema No typed Schema field for this question. est no
AIPM contentIntent 3 est yes

HTML

sufficient

The content intent is to promote galvanizing services for gear materials and highlight the benefits of galvanization.

score 88 · 6,497 chars context · 2,140 in-tokens

HTML answered using 6497 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,579 chars context · 542 in-tokens

Schema did not answer from 1579 chars of context — additional or different layer context needed.

AIPM

sufficient

The content intent is commercial.

score 88 · 7,779 chars context · 2,337 in-tokens

AIPM answered using 7779 chars of layer context (minimum fed this run).

Q7 · Understanding · orientation

Who is the publisher or brand?

Gold: Kapadokya Galvaniz

Gold source: publisher.name · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 5 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand 11 est yes
Schema jsonld.name 5 est yes
AIPM publisher 5 est yes

HTML

sufficient

Kapadokya Galvaniz.

score 100 · 6,497 chars context · 2,141 in-tokens

HTML answered using 6497 chars of layer context (minimum fed this run).

Schema

sufficient

Kapadokya Galvaniz.

score 100 · 1,579 chars context · 543 in-tokens

Schema answered using 1579 chars of layer context (minimum fed this run).

AIPM

sufficient

The publisher is Kapadokya Galvaniz.

score 88 · 7,779 chars context · 2,338 in-tokens

AIPM answered using 7779 chars of layer context (minimum fed this run).

Q8 · Understanding · orientation

What is the page title?

Gold: Dişli Malzeme Galvaniz – Kapadokya Galvaniz

Gold source: title · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost SCHEMA · 5 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title est no
Schema jsonld.name 5 est yes
AIPM title 6 est yes

HTML

insufficient

UNKNOWN.

score 0 · 6,497 chars context · 2,140 in-tokens

HTML did not answer from 6497 chars of context — additional or different layer context needed.

Schema

sufficient

The page title is "Dişli Malzeme Galvaniz – Kapadokya Galvaniz."

score 91 · 1,579 chars context · 542 in-tokens

Schema answered using 1579 chars of layer context (minimum fed this run).

AIPM

sufficient

The page title is "Dişli Malzeme Galvaniz."

score 75 · 7,779 chars context · 2,337 in-tokens

AIPM answered using 7779 chars of layer context (minimum fed this run).

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v2 · Wed, Jul 29, 2026 9:58 AM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.