Google's information gain score counts one thing: data the searcher hasn't already seen. Sharper prose, wider coverage, a fresher date on the byline, none of it registers, because none of it adds information.
That turns ranking into an inventory question. What do you hold that the pages above you don't?
For most content teams the honest answer is nothing, because the standard refresh workflow starts by reading what ranks. You study the top results, list what they cover, and rewrite the same material with better craft and a current date. Every team feeding on the same inputs produces the same output. Google holds a patent on scoring exactly that overlap, which means the teams refreshing hardest are often the ones the system was built to demote.
Google Scores Pages on New Information
The mechanism comes from a patent for "Contextual estimation of link information gain", granted in 2022. The language leaves little room:
"An information gain score for a given document is indicative of additional information that is included in the given document beyond information contained in other documents that were already presented to the user."
Google tracks what a searcher has already read and grades each next result against it. A page that restates the previous three results grades low. A page carrying data the searcher hasn't met grades high.
The patent spells out what happens to the low graders: "one or more documents may be excluded (or significantly demoted) from the search results based on the new information gain scores." A page that adds nothing can fall out of the results entirely.
Google's Helpful Content documentation points the same direction with a question worth answering honestly: "Does the content provide original information, reporting, research, or analysis?" Then a sharper one: "If the content draws on other sources, does it avoid simply copying or rewriting those sources, and instead provide substantial additional value and originality?" And the Quality Rater Guidelines tell human raters to mark regurgitated content as lowest quality. Patent, rater instructions, self-assessment questions: three documents, one signal.
Comprehensiveness Is Table Stakes
For years the winning move was coverage. Write longer, add subtopics, build the most complete result on the query. That era suited the refresh workflow, because completeness is something you can copy.
Now picture the SERP you're trying to rank in: 10 pages covering the same 15 subtopics, citing the same borrowed statistics. Thoroughness separates none of them. The one that pulls ahead holds something the other nine can't show.
Content Refresh Adds Zero Information Gain
Here's the moment the problem lives in. A post that peaked eight months ago has shed 30% of its traffic. The playbook says refresh. You open the top five results, note what they cover, mark the gaps, start rewriting. In another time zone, a competitor's content lead has the same five tabs open.
That's the content freshness lie running at full speed. When step one is reading what ranks, everything downstream inherits the same inputs, and the outputs can't diverge. The publish date moves. The substance stands still.
Every round of this piles onto your content debt. Real editorial hours go into reproducing pages that already exist, and rankings hold flat because the page says nothing the searcher hasn't read twice already.
Zombie Statistics Set the Floor
A conversion benchmark from 2019. A market projection revised three times but still circulating at the original figure. A user behavior stat lifted from a report that lifted it from another report. Zombie statistics keep spreading long after the reality underneath them changed, and they spread without provenance.
When five domains cite the same Gartner number within the same quarter, the scorer sees five pages carrying identical information. The stat entered the ecosystem once. Everything after it is a copy, and the pages built on copies score zero on the only delta being measured.
Each refresh that pulls the same stat from the same shared source drops that floor a little lower. The page reads as current. The data inside it was collected by someone else, years ago.
Scoring 961 SaaS Blog Posts on Originality
Theory says the refresh workflow converges. I wanted the size of the convergence, so we measured it.
We extracted 5,034 claims from 961 SaaS blog posts across 46 domains and tagged each by origin: Original, backed by the publisher's own data; Sourced, attributed to a named external source; or Unattributed, stated with no citation at all. 31% came back Unattributed, and 65.5% were borrowed from third-party research. The Sourced pool was narrow, too: Gartner, McKinsey, HubSpot, and a short list of industry reports recurring across domains.
A separate source verification study traced 3,299 third-party claims toward their origins. 70% offer no external link in the claim's own paragraph that a reader could follow to the data. Of the 30% that do carry a verified link, about 20% point at pages that are dead, gated, or broken. The report got named, the URL went missing, or the source sits behind a paywall that ate it.
That chart is the delta drawn out in public. When the same unverifiable statistics recur across dozens of posts, the novelty each post contributes shrinks toward zero. Orphaned data, cut off from its source and shared too widely to set any page apart.
Charts are claims. So are statistics. When both are borrowed from a shared pool, no amount of prose around them creates information the scorer can credit.
Dividing Original claims by total claims gives each domain an Originality Score, and the spread is stark. GitHub scores 91%. Twilio scores 100%. Both teams write from their own products, their own engineering decisions, their own usage data. The domains at the bottom, scoring 12% to 21%, run on Unattributed claims borrowed from the industry at large. I read every domain in that study, and the split rarely comes down to writing talent. It tracks whether the publishing system generates data or borrows it, which is the exact property Google patented a way to score.
Infrastructure Generates Information Gain
Three kinds of publishing infrastructure produce new information as a side effect of running. Polls collect zero-party data from readers on the page. Living content blocks turn those responses into prose that Google indexes on its next crawl. Charts wired to live sources update without a re-export or re-upload. A reader votes, a paragraph rewrites itself, and the crawler finds content that wasn't there last week.
A scoring system that punishes shared inputs leaves one scalable response: publish from data your own pages generate. LiquiChart is one implementation of that loop.
Go back to the question in Google's Helpful Content documentation: "Does the content provide original information, reporting, research, or analysis?" A living poll that generates its own data answers it with every response collected. The data is original because your readers supplied it. The analysis is original because no other page holds those responses.
JSON-LD Dataset schema on the embed hands Google a machine-readable signal that the page carries structured original data. The AI Insights tab turns the response distribution into analysis no other page can run.
The Number Nobody Tracks
What share of the data claims on your site rest on your own research? Almost no team can name that figure, and that blind spot is the information gain problem sitting inside your own CMS. A poll that creates a dataset where none existed is the purest form of the fix, because the aggregate answer will say something no page in the SERP can. So answer it for your own team, right here.
Every response above joins a distribution that exists nowhere else on the web.
What Teams Report
Self-reported numbers are the only kind available before the poll reaches scale, and they carry their own limits.
As readers weigh in above, the distribution will sharpen. The measurement gap is already clear: most publishing teams have never quantified how much of their published data originated from their own systems versus borrowed sources. That ratio determines whether a page generates information gain or reproduces it. Infrastructure that collects proprietary data through reader interaction makes the ratio measurable for the first time. Until teams know the number, they cannot change it.
Measure Your Own Originality Score
The poll above measures the industry. Your risk is specific to your URLs: which posts share the most claims with the competitors ranked beside them, and which claims carry no source at all. Scan a published page and the Content Health Scanner extracts its claims, tags each as Original, Sourced, or Unattributed, and returns the same Originality Score behind every domain in the study. The diagnosis takes under a minute, and knowing the gap comes before closing it.
Information Gain Is the New Unit of Differentiation
A page holding 500 reader responses contains data a competitor can't obtain by reading it. A paragraph that rewrites on live input says something new each week. A chart on a live source shows the current quarter instead of the one it was exported from.
Every crawl scores the delta between pages like that and pages assembled from the shared pool, and every crawl widens it. Teams that build publishing systems capable of generating their own data will own the delta. Everyone else will split it.
The delta never shrinks.