Two-Thirds of SaaS Blog Claims Are Borrowed Data

46 domains. 961 posts. 5,034 claims.

Daniel SmithApr 1, 2026Living Content7 min read

Most of the numbers in a SaaS blog post came from somewhere else.

We pulled 5,034 statistical claims out of 961 posts across 46 SaaS domains and sorted every one by where its data originated. 34.5% are first-party: the company's own product data, its own surveys, its own customers. The other 65.5% are borrowed, a figure from someone else's report or study restated in the prose.

Every blog borrows. Covering an industry means citing people who measured things you never measured yourself. What matters is what a borrowed number carries with it when it ships, because a publisher can't vouch for data it never generated, and 70% of those borrowed claims give the reader no external link to follow back to a source. Each one is content debt taken on the day the post goes live.

How the 5,034 claims break down

Every claim in the scan landed in one of four buckets.

34.5% are first-party. Product telemetry, the company's own survey, its own customer metrics. SaaS companies generate plenty of real data, and when a claim comes from it, the publisher and the source are the same entity. These are the numbers a blog can stand behind.

20% cite a third party and link out. A reader can click through, find the report, and judge the methodology for themselves. That's the floor for attribution that holds up, and fewer than one in three borrowed numbers clears it.

15% name a source and stop there. Some attribute a figure to a named analyst firm with no URL anywhere on the page. Many more link only within the publisher's own domain, so the citation loops back into the blog instead of out to the research. Either way the reader gets a brand name where a path should be.

31% are unsourced. A number sits in the prose as if everyone already knew it. "72% of B2B buyers prefer self-service." Whose survey? What year? How many respondents? The page has no answer.

The risk pools in that gap between naming and linking. Seven in ten borrowed claims give the reader nothing to click, and even the linked share only stays verifiable as long as the pages behind those links stay up.

Where borrowed claims lose the trail

Of the 3,299 third-party claims in the dataset, 1,006 carry an external link in the claim's own paragraph. That's 30.5%.

The remaining 69.5% split two ways. 752 claims name a source without a verified external link, and 1,541 attribute nothing at all. Together that's 2,293 numbers a reader can only check by leaving the post and running their own search.

Picture the sentence "HubSpot reports that 60% of marketers prioritize blog content." The reader gets a brand and a figure. Which report, what year the data was gathered, what the sample looked like, all of it stays out of reach. The claim ships as an orphan, and it stays one unless somebody circles back and attaches the link by hand.

A name is a label. A link is a path.

Teams that name their sources tend to believe they're citing them. But a name asks the reader to trust the writer, the writer to trust a memory of the original, and nobody along that chain can tell whether the figure mutated across three retellings before it reached the page.

What the scan sees is the published output. The habits that produce it live upstream in the drafting doc, and only the people writing can report on those.

Living Content

Most content teams track publishing cadence, not claim attribution. The two require different infrastructure. As readers weigh in above, the gap between what teams intend and what they actually do at the claim level will sharpen.

Whatever your team's answer, the missing links in this dataset look like a workflow gap rather than a care gap. Every one of the 46 domains publishes regularly and updates often; the sourcing step just never existed in their process. I've shipped that gap myself: a statistic remembered from something I read, pasted without a link, or lifted from an older post of ours that never linked the original either. The number keeps traveling long after anyone could say where it started.

What high first-party blogs have in common

First-party share ranges from 0% on blogs that write entirely about the wider industry to 65% on blogs that write mostly about themselves, with a median near 29%. That spread says more about editorial subject matter than about quality.

The blogs at the top all write about data they own. Stripe, MailerLite, and Clari lead the dataset, each attributing more than 60% of their claims to internal product data, transparency reports, and customer metrics. A reader can check those numbers against the company itself, because the company produced them.

A blog covering a field it generates no data about has to borrow, and every borrowed figure comes with upkeep: link it, verify it, keep it current. Blogs built on their own data skip that burden entirely, while the ones synthesizing the field inherit every bit of it. What separates the two groups is whether the publishing system records a source at the moment a claim is created, so the number enters the world with a trail attached instead of as one more figure the reader takes on faith.

Why freshness audits miss attribution

Freshness measures when a post was last touched. Attribution measures whether anyone can check the numbers inside it. Those are different axes, and audits only run along the first one.

A post updated last week can lean on a borrowed figure whose source now returns a 404. A two-year-old post can tie every claim to a named, dated study. Bumping the publish date and confirming that links resolve is freshness theater: the ritual passes while the claims themselves were never examined. Across 46 companies with very different editorial operations, that held every time.

AI-assisted drafting compounds it. A statistic that arrives in a model's draft has no provenance at all; it sounds plausible and traces to nothing, and without a record of the citation supply chain from published claim back to primary source, there's no separating what's real from what's approximate from what's generated. Any system that detects when published data goes stale depends on the data being traceable to begin with, and the first requirement of living content is a recorded origin for every claim.

How we built the dataset

The scan ran on LiquiChart's claim extraction and source verification infrastructure, the same engine behind the Content Health Scanner. Each post was parsed for statistical assertions, and each assertion was classified by attribution type: first-party, sourced with a link, sourced without a link, or unsourced. The classifier recorded provenance and made no judgment about quality.

We selected 46 SaaS domains with active blogs and collected recent posts from each. A programmatic filter removed statistics roundups, annual benchmark compilations, product announcements, and quote-heavy listicles, and a hand review of the remaining titles dropped posts carrying fewer than two data claims. The 961 posts left are ordinary argumentative SaaS content: a case, data behind it, a conclusion.

One question drove classification: can a reader verify this number? First-party claims qualify because the company is the source. Linked third-party claims qualify because the reader can click through. Named-without-link and unsourced claims fail, because the path ends at the prose. A link counts only when it appears in the claim's own paragraph and survives tracking-link filters; an earlier pass credited any external link near the claim, which inflated the linked share by treating navigation and neighboring-paragraph links as citations. Source URLs were then checked for availability and, where a date was present, for freshness. The full methodology is published for reproducibility.

The gap is structural. The posts are recent and the statistics are recent, yet for 65.5% of the claims the source sits on someone else's server, and for seven in ten of those, no route back to it exists on the page.

Our own blog was one of the 46 domains, and until the scan ran I couldn't have vouched for our numbers either. The Content Health Scanner runs the same extraction on any URL. Run it on one of yours.

Trace Every Stat Back to Its Source

Hop-by-hop citation tracing. See exactly where each number originates — and which chains end at a paywall, a broken link, or a value that drifted along the way.

Supporting Data & Claims

Every anchor below is first-party. Polls are live. Claims are monitored. Experiments are dated.

Related Posts

The Citation Monoculture: How Few Sources SaaS Blogs Share

1,006 citations. 516 domains. The weight behaves like 104.

Jul 10, 2026

Why Unsourced Stats Are Round Numbers (2,473-Claim Analysis)

The last digit of a percentage tells you whether anyone counted it.

Jul 10, 2026

Undatable Data: Most Web Numbers Cannot Be Age-Checked

We scanned 5,034 data claims across 46 domains and found that only 11.8% carry a date you can actually check.

Jul 10, 2026