Information Gain Scorecard & Prompt Framework

I tried to explain the need for a report like this in my “Crawled – Currently Not Indexed” Is an Economics Problem, Not Just a Quality Problem article and finally had some time to adapt it for a wider audience.

Why This Matters: The Economics of Source #11 to N

AI has changed search optimization forever. While many are fighting over the proper acronym, others are looking at this through more mature lenses. In nearly every enterprise earnings report and especially those of the Magnificent 7 and pre-IPO data leaks, everyone is talking about the costs of generative AI. Despite being complex and resource-intensive, traditional search engine crawling and indexing content is pennies compared to the tokens necessary to process and infer responses for AI results. It should not be a surprise that all of the LLMs and search engines are being more selective on what they ingest into their system.

The old model: cost was mostly storage and crawl bandwidth. Cheap enough that indexing source #11, #50, #200 saying roughly the same thing about a topic was a rounding error. Links then sorted which of those thousand pages surfaced on page one — but nearly all of them still got a seat in the index.

The new model: both Google and AI labs now run pages through synthesis — AI Overviews, AI Mode, RAG pipelines, LLM-based answer generation. That’s a fundamentally more expensive operation than storing a URL. It means reading the page, extracting what it adds, and deciding whether it changes the synthesized answer at all. Once a search engine or AI lab has already built a solid answer from sources #1–10, source #11 saying the same thing contributes zero marginal information to that answer — but ingesting, embedding, and periodically re-crawling it still costs the same as it always did. That’s a bad trade, repeated at web scale, across every commodity topic ever written about.

Old Model (Link Economy)New Model (Synthesis Economy)
Primary costCrawl bandwidth + storageCrawl + ingest + synthesis/inference
Marginal cost of source #11Near zeroReal (embedding, re-crawl, evaluation)
Primary sorting signalLinks (proxy for trust)Information gain (direct value signal)
What links still doRank the pageEstablish trust/authority
What decides index inclusionMostly technical/crawl budgetIncreasingly: does this change the answer?

Google has been trying to evangelize the concept of Commodity Content at several Search Central Live events and in its AI Optimization guide. Google defines commodity content as generic, easily replicable information that anyone, or an AI model, can write, contrasting it with non-commodity content, which requires real human experience, unique data, and first-hand expertise.

Why commodity content specifically gets throttled, not just low-quality content: a commodity page isn’t necessarily bad. Marie Haynes makes this point directly — these are often good, competently written articles. The problem is redundancy, not quality. From a synthesis-economics standpoint, a well-written page that says nothing new is worse than a mediocre page that adds one new data point, because the well-written redundant page still costs the same to process and yields nothing the system didn’t already have.

Information gain is the signal that survives this shift because it’s the one variable that determines whether ingesting a page changes the output. Links answer “is this trustworthy.” Gain answers “is this worth the compute.” As indexing and ranking increasingly happen through a synthesis layer rather than a link-graph layer, gain becomes the gatekeeping variable — not because links stopped mattering, but because the marginal cost curve for redundant content got steeper for everyone doing synthesis, not just Google.

My Contribution To the Nonsense

More than 20 years ago at PubCon, I showed how one of my enterprise clients had converted their simple glossary into a topical knowledge base. Not just defining what cloud computing is, but understanding searchers’ interests. We found that a simple question often forked into multiple questions: security, types of cloud, and literally, what is it and what does it mean for my business? Within a month, multiple nimble competitors cloned this approach, albeit half-assed, spinning out glossary-type content that did not work well. Then in 2019, when Google launched BERT, this deep definition content went insane. Since it was deep and descriptive, Google demoted everyone who was not simply answering the question. The problem no one was adding anything to it.

Fast forward to 2023 when an agency reached out to Hreflang Builder to deploy hreflang to solve their cannibalization problem. Their client has spent “millions” creating this type of glossary content and put it out to nearly 100 markets in English. We called this “aspirational hreflang.” There was nothing additive; in fact, when we researched it, 89 to 95% was found on multiple other sites as written. None of it was ranking. As part of our tests on why Hreflang was not working, we found it was 100 versions of 100% duplicate, and even the mighty hreflang could not help it. That is exactly what Google is talking about. Cloning content just because there are search queries. And no or minimal context to your business is a significant waste of resources. I fear the new crop of AI content tools will do the same thing.

Caution: Why Content GEO Gap Tools Can Undermine This Goal

Most GEO/AEO tools follow the same playbook: pull the sources currently cited in an AI Overview or answer-engine response, diff your content against them, and generate a “gap” list of topics, entities, or phrases you’re missing. The recommendation is almost always to add that content to close the gap.

Why this backfires under an information-gain model: the gap analysis is built entirely from what’s already cited — meaning it optimizes you toward matching the existing consensus, not adding to it. If ten cited sources all mention “recommended chlorine levels,” and your page doesn’t, the tool flags it as a gap. You add it. Now you match the consensus perfectly — but you haven’t added anything the system didn’t already have. You’ve become source #11 by design. The tool is effectively automating the production of commodity content, because it defines “gap” as absence relative to competitors rather than absence relative to what’s true and useful in the world.

The deeper issue: these tools are built on the wrong objective. A gap tool answers “what do the cited sources say that I don’t say,” essentially a consensus-matching objective. The objective that actually matters for indexation and citation now is “what does the world know, or what experience do I have, that no cited source says” — an information-gain objective. Those are nearly opposite instructions. One pushes you toward convergence with competitors; the other pushes toward divergence from all of them.

The scale problem: if every site in a niche runs the same class of gap tool against the same cited sources, everyone converges toward the same expanded checklist of subtopics. You get more comprehensive pages — but comprehensive in exactly the same way as everyone else. This is the fan-out risk Marie’s article flags directly: anticipating and covering every related query looks like thoroughness, but reads to the algorithm as more surface area of the same commodity content, not depth.

The fix: don’t discard gap analysis — sequence it correctly. It’s useful for confirming you’re not missing baseline coverage or accuracy (a consensus floor you still need to clear to be a legitimate answer). But it should never be the source of your differentiation strategy. Real differentiation has to come from something a gap tool structurally cannot surface: first-party data, direct experience, a contrarian but defensible read, or specificity no competitor bothered to capture. A gap tool can tell you what’s missing from the room everyone’s already standing in. It can’t tell you what’s outside the room.

Content Evaluation Framework

Grounded directly in the criteria Google’s own Creating Helpful Content documentation lists, Google’s New AI Content guide, and the methodology Marie Haynes uses in practice: comparing your page against the pages currently earning the citation/ranking, not against an abstract rubric.

The four questions (Google’s actual criteria):

  1. Does the content provide original information, reporting, research, or analysis?
  2. Does it provide insightful analysis or interesting information beyond the obvious?
  3. If it draws on other sources, does it avoid simply rewriting them — adding substantial value and originality instead?
  4. Does it provide substantial value compared to other pages already in search results?

Information Gain is the primary, load-bearing pillar — a page can be fresh, complete, and pleasant to read and still fail if it adds nothing new. Freshness, Completeness, and Page Experience are secondary checks: necessary to clear, but not sufficient on their own to earn a high score.

Practical scoring guide (for internal triage, not a Google-published rubric):

ScoreLabelWhat it means
8–10Primary / Non-CommodityOffers original data, direct experience, or a perspective competitors structurally can’t replicate
5–7Useful but at riskSolid, accurate, well-written — but largely matches existing consensus coverage
1–4CommodityRehashes what’s already indexed; an AI Overview could fully replace it

Step 0 — Rule Out Technical Causes First

Before auditing content quality, confirm the page isn’t stuck in “crawled-currently not indexed” for a technical reason. Use GSC’s URL Inspection tool → “Test Live URL” and check the rendered HTML. Common culprits: robots.txt blocking parameterized URLs the theme relies on (CSS/JS), JS-rendering failures, or a recent migration. If the live test shows Google genuinely can’t see the content, fix that first. A content audit won’t help a page Google can’t read. If the live test shows full content, proceed to the audit below.

The Merged Prompt: Comparative 3-Pillar Delta Audit

This combines Marie Haynes’ core method, auditing your page against what’s actually ranking right now, not in isolation, with three useful additions (Freshness, Completeness, Page Experience) and her explicit two-stage discipline: diagnose before prescribing.

Step 1 — Establish the live baseline. Search the target query. Open the AI Overview and/or top-ranking pages, and load them into Gemini (via Chrome sidebar tabs), or paste their URLs/content into Claude/ChatGPT. Run:

You are auditing content against Google’s helpful content standards and current indexation dynamics, where AI-driven synthesis makes redundant content costly to keep indexed.

Evaluate these currently-ranking/cited pages on four pillars:

1. INFORMATION GAIN (primary pillar)

   – Does the content provide original information, reporting, research, or analysis?

   – Does it provide insightful analysis beyond the obvious?

   – If it draws on other sources, does it add substantial value and originality rather than just rewriting them?

   – Does it provide substantial value compared to other pages already in search results?

2. FRESHNESS

   – Is the content temporally current?

   – Are cited facts, data, policies, or specs up to date?

3. COMPLETENESS

   – Does it fully resolve the core query intent?

   – Does it address the logical follow-up questions a reader would otherwise have to search separately?

4. PAGE EXPERIENCE

   – Is the actual content easy to access, or buried behind ads, interstitials, excessive filler, or slow-loading elements?

   – Would a user’s real experience of this page match the quality of the text itself?

For each page, note specifically what makes it non-commodity on each pillar — first-hand experience, proprietary data, a unique angle, etc. Also note if the source appears to be an established authority in this space, since authoritative sources can carry more “commodity-ness” without being penalized.

Step 2 — Diagnose your page against that same live bar (no fixes yet).

Now analyze this page against the same four pillars and the same live comparison set. This page is not ranking/indexing well — it is our client’s page. Please share specifically where it is lacking on each pillar. Also flag if any “completeness” in this page looks like shallow fan-out coverage (many thin subtopics) rather than genuine depth, since that pattern carries its own scaled-content risk. No need to suggest improvements yet.

[PASTE CLIENT PAGE / URL]

Keep Steps 1 and 2 as separate turns, and keep diagnosis separate from ideation. Asking for gaps and fixes in the same pass tends to pull the model toward generic advice (“add more headings”) instead of a clean-eyed read of what’s actually missing.

Output format to request (once both turns are complete, ask the model to summarize):

## Content Quality & Information Gain Audit

**Overall Non-Commodity Score:** [1–10]

**Indexation Risk:** [Low / Medium / High]

**Classification:** [Primary Expert Analysis | High-Quality Editorial 

Commodity | Thin Commodity Content]

### Pillar 1: Information Gain — [Pass/Borderline/Fail]

### Pillar 2: Freshness — [Pass/Borderline/Fail]

### Pillar 3: Completeness — [Pass/Borderline/Fail]

### Pillar 4: Page Experience — [Pass/Borderline/Fail]

**Commodity elements:** [generic, rehashed points with zero new value]

**Unique delta (gain):** [explicit unique insight, proof, or data found]

Step 3 — Idea Generation (Separate Follow-Up Turn)

Only after diagnosis is complete, use a prompt like this:

Given where this page is lacking, give me 20 ideas that draw from our first-hand experience to make this article substantially better than anything else that exists on this topic on the web.

This intentionally comes after diagnosis and is meant to surface ideas rooted in the site owner’s own experience — not a consensus gap list pulled from competitor content. If the output starts reading like a topic checklist rather than experience-based ideas, that’s a sign the model has drifted into gap-matching mode; redirect it explicitly toward first-hand knowledge, proprietary data, or a defensible original take.

Quick standalone check, useful for a lighter first pass before running the full audit:

Is this content likely to be considered commodity content? Explain your reasoning against Google’s helpful content standards.

[PASTE CONTENT]

This framework leans heavily on Marie Haynes’ methodology, “Why Your Pages Are Stuck In Crawled-Currently Not Indexed & What To Do About It,” Search Engine Journal, July 2026, and her deeper article Why Your Pages Are Stuck In Crawled-Currently Not Indexed. And what to do about it.