Somewhere in your blog there's a post that cites a statistic. The link under it still works. The page loads fast and returns 200 OK. And at some point between your publish date and today, the number left that page. Nothing in your stack noticed, because every signal a link checker reads says that citation is healthy. That's a hollow citation, and this study counts them.
We measured citation chain depth across 46 SaaS blog domains: how many hops sit between a published claim and whatever source waits at the end of its reference trail. Of 1,006 verified linked citations followed all the way down, 17.2% reached a primary source. The rest ended at live pages missing the number, at bot walls and login gates, or at secondary sources citing secondary sources citing nothing at all.
Two companion studies frame the scale of the problem. Across the same kind of corpus, about two-thirds of SaaS blog claims are borrowed from third-party research, and when a borrowed claim does carry an external link, about one in five resolves to a dead, gated, or broken page. Those studies asked whether citations exist and whether their links resolve. This one asks where the trail ends when the link works.
How We Measured Citation Chain Depth
A citation chain starts at a published claim and follows links downstream. Your post cites a number and links to a page. Maybe that page cites its own source and links again. Every link followed is one hop, and the chain ends where a page either supplies the original dataset or names nothing deeper.
LiquiChart's claim extraction pipeline pulled every monitorable claim from each post: statistical, comparative, source-citation, and temporal. A percentage qualifies, and so does a ratio, a count, or a dated assertion, because a claim about "a 2024 survey" can be checked against reality just as hard as a claim about "67%". A production tracer then walked each chain: open the cited page, check whether the asserted value or date is actually on it, and if it is, look for a deeper source to follow.
The census covered 961 posts on 46 domains and traced 1,006 citations. That number is smaller than the claim count for a reason. The corpus held 3,299 borrowed claims, and only 30% of them carried a followable external link once a genuine citation was separated from a number that just happens to sit near a URL. The 1,006 are that 30%: every case with a claim, a link, and at least one page to judge at the other end. First-party operational claims stayed out of scope, since a vendor quoting its own uptime figure has no upstream trail to walk.
Each terminus landed in one of six buckets:
- Primary source: the chain reaches the original dataset, study, or methodology.
- Claim not found: the cited page loads but the specific number is absent, or the page is the wrong one.
- Access blocked: the page exists but can't be read: a bot wall, a robots block, a login gate, a redirect loop, an unparsed PDF, or a client-side shell that never renders the value.
- Named without link: the post names a source by title but provides no URL.
- Broken: the URL returns a 4xx or 5xx error.
- Circular: the chain loops back to a page already visited.
How Far Does Your Team Trace a Citation
Behind those 1,006 chains sit thousands of small editorial decisions about when a source is checked enough. Before you see how the chains ended, log your own team's practice:
The gap between the first two options is the gap between a chain that stops at layer one and a chain that touches the ground. By default, that standard goes unwritten, which means three writers on the same blog can publish at three different verification depths without anyone deciding that.
Where 1,006 Citation Chains End
Six possible endings, one very lopsided distribution.
173 chains (17.2%) reached a primary source [15.0%, 19.7%]: the researcher's dataset, the government database, the survey methodology itself. On these chains, a reader who wants to check the claim can put their hands on the data that produced it.
The biggest bucket was absence. 374 chains (37.2%) ended at a live page that no longer contained the claimed number, or that turned out to be the wrong page for the claim. Some sources had shipped a newer edition and dropped the old figure. Others had restructured the page and cut the section. Either way, the citing post now leans on a page that stopped holding it up, and from the outside everything still looks fine: blue link, fast load, healthy status code.
Absence is easy to overcount, so every chain in this bucket got a second pass. A full-page browser render read the complete text, not just the first fetch, before any chain was scored as claim-not-found. The 37.2% is what survived that re-check: the number genuinely missing, or the page genuinely wrong.
333 chains (33.1%) hit a page that exists but can't be read. The dataset splits this bucket into bot detection that turns away an automated reader (11.8%), redirect loops (7.2%), robots directives that disallow the path (5.5%), PDFs that never parse (5.2%), and client-side shells that render no text to a fetch (3.5%). Every one of these pages is alive, which is exactly why no uptime tool will ever flag them.
71 chains (7.1%) named a source without linking to it: "according to a leading analyst firm," full stop, no URL. 49 chains (4.9%) were broken links. Six chains (0.6%) looped back on themselves or dead-ended at a page asserting the number with no source of its own.
Sit with that 4.9% for a second, because broken links are the entire category your current tooling can see. A link checker audits the 4.9%. The other 77% fail while every check reports green.
Mean Citation Chain Depth Is 1.20 Hops
How deep do the trails actually go?
Depth one: 823 chains. Depth two: 170. Depth three: 12. Depth four: one. Averaged across all 1,006, the mean chain depth is 1.20 hops, and 82% of chains are a single hop long. The post links to one page, and whatever that page is, that's where the trail stops: original dataset, secondary writeup, or a page that never held the number at all.
Depth also predicts where a chain ends. Of the 823 single-hop chains, 165 (20.0%) reach a primary source: the post linked the original research directly. Of the 170 two-hop chains, eight (4.7%) reach one. None of the 13 chains that ran three or more hops made it; the longest passed through several secondary sources before dying at a gated report with no methodology section.
Each added hop routes the claim through one more secondary page, and each secondary page is one more chance for the trail to end at a wall, a login, or a missing number. A chain that starts at a middleman almost never recovers.
I've watched a writer drop a hyperlink behind a statistic and move on, sourcing done. In that workflow, the link itself is the verification step, and whether the page at the far end contains the original data is a question nobody's job description includes.
Gated Reports End More Chains Than Broken Links
The largest slice of unreadable endings wasn't dead URLs. It was live pages an automated reader can't enter: research aggregators, consultancy reports, analyst portals, industry-body PDFs. These sites compile numbers from elsewhere and park them behind registration walls, bot challenges, or paywalls.
The shape repeats across the corpus. A SaaS post cites a statistic and links to an aggregator. A reader clicks through and meets a login form; a crawler meets a bot challenge. The chain dies at the gate, and whatever source the aggregator itself leaned on stays invisible to everyone downstream.
Each of those terminations mints a zombie statistic: a number that keeps circulating with no verifiable path back to the research that produced it. The blog states it as fact. The aggregator frames it as a data point. Neither page shows a methodology or a dataset, and the number survives on repetition alone. In my own audits, the aggregator link is the one I brace for, because it almost always ends at a tollbooth.
Three Citation Failure Modes That Pass Review
The Statistic That Left the Page
One chain started at a claim about analytics data lost to cookie-consent denial. The citation pointed to a marketing-agency post that was live, current, and well maintained. The figure was nowhere on it. Click the link today and you'll land on a perfectly healthy page that simply no longer says what the citing post needs it to say.
That's a frozen liability: a claim bolted to a source that moved on without it. The number has become orphaned data, present downstream and absent upstream, with nothing connecting the two ends. Of all six outcomes, this one was the most common in the dataset.
The Source Named With No URL
71 chains (7.1%) died at prose. The citing post links to a second page, the second page says "according to our research," and the trail stops there with no URL to follow.
One of these began at a claim about daily search volume and led to a vendor post that name-dropped businesses and creators but never linked to where its figure came from. The chain ended at an authoritative domain publishing the number with no upstream path. This failure mode clears a dead-link check and a paywall check both: the URL resolves, the page loads on a trusted domain, and a reader who takes the domain's authority as the answer never scrolls far enough to see the trail dissolve into assertion.
Three Hops Into a JavaScript Wall
One chain followed a platform user-count claim through two secondary pages to the originating announcement. The announcement returns 200 OK. To the tracer, it rendered exactly one word of body text, because everything else loads client-side. Three hops, three pages, each one wearing the authority of the page above it, and at the top a source that no static fetch, search crawler, citation auditor, or training pipeline can read. The number may well sit behind the JavaScript. Nothing that audits at scale will ever confirm it.
Citation Quality Means Provenance
In this dataset, 83% of chains ended before touching the research that generated the number, and of all third-party claims in the corpus, only 5.2% trace to a primary source. A companion study across the same 46 domains found one in five posts carry data two or more years out of date, and that share grows the longer a post sits unmaintained.
Each unverified chain is content debt accruing without a signal: the data behind the link shifts and nothing in the publishing workflow fires. Counting citations tells you nothing about SaaS blog citation quality, because 15 linked statistics with no verified path to original research carry exactly the provenance of five unsourced claims.
Provenance also decides what Google's information gain score actually rewards. A post that traces a claim to its primary source brings something no other result in the SERP has. A post citing the same secondary aggregator as 10 competitors recycles a reference everyone already holds.
LiquiChart's claim-level monitoring watches the upstream URL beyond its HTTP status, so a number vanishing from a live page propagates to every post that cited it. That difference between link health and citation health is where published claims lose their footing, and if you're building systems to detect when published data goes stale, start where this study found the exposure: the third of chains that resolve, load, and no longer contain the number.