About 1 in 5 Linked SaaS Citations Is Dead or Broken

46 domains. 961 posts. 3,299 third-party citations.

Daniel SmithApr 1, 2026Living Content8 min read

For years the sourcing problem in SaaS content was a missing link. A statistic sat in the prose with no name and no URL, and a reader who wanted to check it had nowhere to go. The industry heard that complaint and applied the obvious fix: name the study, link the report.

We traced 3,299 third-party claims across 961 posts on 46 SaaS domains to measure how far the fix reached. Only 30% of those claims carry an external link in the claim's own paragraph, the spot a reader actually scans when a number makes them pause. The pages themselves are dense with links. Links that belong to the specific claim are rare.

And the links that do exist have their own failure rate. When we followed every verified source link through a real browser, about one in five came back dead, gated, or otherwise unreadable.

Sourcing Across 46 SaaS Blogs

The corpus is 5,034 claims extracted from 961 posts on 46 SaaS domains, the same one behind the companion attribution study. Removing first-party claims, the numbers a company reports about itself, leaves 3,299 third-party statistics. Every one of them rests on someone else's research, so every one of them needs a trail a reader can follow.

Here's how they split:

  • 1,006 claims (30%) carry an external link in the claim's own paragraph. A reader can click through to the origin.
  • 752 claims (23%) name a source without a verified external link. Some cite an authority by brand with no URL. Many link only to the publisher's own pages, which loops the reader back into the blog instead of out to the evidence.
  • 1,541 claims (47%) cite nothing. The number sits in the prose as if everyone already knew it.

Seven in ten borrowed claims give a reader no path to the origin. Attributing has spread faster than the discipline of making the attribution checkable.

Part of that 70% is an artifact of proximity, and it shaped how this audit counts. A link earns credit here only when it sits in the claim's own paragraph and survives tracking-link filters. Under a looser rule, crediting any external link near the claim, roughly a third of the seemingly linked citations turned out to be borrowing a neighbor: a navigation item, a footer, or the citation attached to a different paragraph's claim.

The unverifiable claims matter on their own. A sentence like "a named research firm reports that 75% of B2B organizations will shift to a composable architecture by 2027" hands the reader a brand and a number with no report title, no year, and no way in. That claim is orphaned the day it ships. What we wanted to know next is what happens when the link does exist and you follow it.

We re-checked all 1,006 verified third-party source URLs through a real browser, the same rendered-fetch path our scanner uses: request the link, escalate to a headless browser when a bot-wall blocks the first attempt, classify what comes back. 801 resolved to a readable page. 205 did not.

That's 20.4% of the external citations. 20 were dead outright, a 404 or 410 where the page used to be. 102 were gated, a login wall or paywall waiting on the far side of the click. 83 returned a server error, a redirect loop, or a rate-limit block that even a real browser couldn't get past. Every one of those 205 citations cleared the only check most editorial workflows run, which is whether a link exists and points somewhere plausible.

The link is the promise. The page is the proof.

A claim in that state keeps presenting its number as fact. A reader who clicks lands on a 404, a login form, or a timeout, and the sentence loses the only footing it had. I've shipped citations like that myself: live the day they went in, dead a year later, with nothing in place to notice the change.

The 2.0% genuinely-dead slice deserves its own line, because dead links are the one failure existing tools catch. The remaining 18%, the gated pages and redirect loops and rate-limited servers, return statuses that no broken-link checker treats as a problem. If one of these belongs to you, fixing link rot in your citations starts with separating a source that's truly gone from one hiding behind a login wall.

One methodology choice moves this number a lot, so it's worth stating plainly. A naive checker that fires raw requests with no browser flags reports closer to three in ten (28.3%), because it counts every bot-wall, consent-redirect loop, and rate-limit block as a dead source. Escalating those through a real browser, the path a reader's own browser takes, clears about eight points of them. We report the browser number because it matches what a reader experiences, and because a checker weaker than the pages it audits will overcount the damage.

Reachability is the lower bar. A link that resolves can still anchor a claim to data that a newer edition has since replaced.

Among the external citations that resolved and carried a readable publication or update date, 224 were under a year old, 76 were one to two years old, and 157 were at least two years old. More than a third of the datable links point at data older than most readers assume a cited statistic to be.

Age alone doesn't condemn a source. A 2024 benchmark cited in a 2024 post holds up fine. The same benchmark cited in the present tense in a post a reader opens today asks them to take aging data as current, and nothing on the page signals that the number may have moved since.

More than half of the linked sources carried no machine-readable date at all. That's a finding in itself: for most citations, even the freshness of the source can't be determined from the outside. The page behind the link might be from last month or from five years ago, and the reader has no way to tell.

What Teams Check Before Citing

The findings so far describe what gets published. The step before publish is worth measuring too, so here's the question put directly to you: when someone on your team drops a statistic into a draft, how far does the check go?

Living Content

Most content teams treat sourcing as a formatting step: find a number, paste a link, move on. Whether that link points to the original dataset or to another post that also pasted a link rarely enters the workflow. The distinction matters because a secondary citation can go stale without the citing team ever knowing. Only the primary source change triggers a correction. Everything downstream inherits the error.

A citation whose link nobody has opened since the day it was added looks identical, in the published post, to one somebody read last week. The reader can't tell them apart, and most editorial workflows can't either.

Sourcing Tracks Subject Matter

The 46 domains don't share a standard. The unsourced share runs from 6% on the most careful blogs to past 40% on the least, with a median near 21%. I've read across that whole range, and writer diligence explains less of the spread than you'd expect.

Subject matter explains more. Blogs built on their own platform data, usage metrics, internal experiments, transparency reports, verify by default, because the company is the source and there's no third-party trail to lose. Blogs that synthesize industry trends borrow most of their numbers, and each borrowed number is one more link that has to exist, resolve, and stay current.

Teams writing about their own products get verifiability as a byproduct. Teams synthesizing the wider industry inherit a portfolio of links that decay on someone else's schedule.

Monitoring Citations as Dependencies

The 70% unverifiable rate measures absence: claims that never had a checkable trail. The 20.4% unhealthy-link rate measures decay: claims that had one and lost it after publication. Only the second can be repaired after the fact, because only the second leaves something to watch.

Link-level monitoring treats each citation as a dependency. When a source URL starts returning a 404, a 403, or a timeout, every claim citing it gets flagged. When the page still loads but the cited number no longer appears on it, that gets flagged too, which is what claim verification catches that a link check cannot: a link that resolves cleanly to a page that never contained your number. Monitoring the full citation chain carries the same check past the first hop, to the sources your sources cite.

None of this works for the 70%. A claim with no external connection to its origin offers nothing to watch. Detecting when published data goes stale needs a trail, and living content needs a known origin. Both start with a recorded link that still resolves.

Across this dataset, the citation supply chain those tools would watch is mostly built, and about one link in five has already given way. The scan behind this study runs on any page through the content health scanner.

Your readers find the failures one click at a time. Check the pages behind your own links before they do.

Source-level tracking is in early access. You can join the waitlist or browse the claim registry.

Trace Every Stat Back to Its Source

Hop-by-hop citation tracing. See exactly where each number originates — and which chains end at a paywall, a broken link, or a value that drifted along the way.

Supporting Data & Claims

Every anchor below is first-party. Polls are live. Claims are monitored. Experiments are dated.

Related Posts

The Citation Monoculture: How Few Sources SaaS Blogs Share

1,006 citations. 516 domains. The weight behaves like 104.

Jul 10, 2026

Why Unsourced Stats Are Round Numbers (2,473-Claim Analysis)

The last digit of a percentage tells you whether anyone counted it.

Jul 10, 2026

Undatable Data: Most Web Numbers Cannot Be Age-Checked

We scanned 5,034 data claims across 46 domains and found that only 11.8% carry a date you can actually check.

Jul 10, 2026