Undatable Data: Most Web Numbers Cannot Be Age-Checked

We scanned 5,034 data claims across 46 domains and found that only 11.8% carry a date you can actually check.

Daniel SmithJul 10, 2026Living Content9 min read

Open the source behind the last statistic you cited and look for its date. For most numbers on the web there is nothing to find: not a stale figure, not a current one, just a number with no attached measurement date to check its age against. When we scanned 5,034 data claims across 961 SaaS blog posts for our State of Content Decay corpus this July, 88.2% carried no discoverable date at all. That is undatable data: a figure you cannot age-check because no measurement date was ever attached to it. Every freshness audit you have ever run worked off the other 11.8%.

Every conversation about content decay argues about how old the data is. This report is about the layer underneath: how much of it can be dated at all. The undatable majority is the part of your archive you can least stand behind, and the claims resting on undated sources go stale at roughly twice the rate of the ones you can age-check. You have been auditing the numbers that look old, and the riskier ones never gave you a date to flag.

Undatable Data: Why 88.2% of Numbers Have No Age to Check

Undatable data is a published statistic with no discoverable measurement date. You cannot find when the underlying figure was collected, so you cannot verify whether it is still current. In our July 2026 scan of 5,034 claims across 46 domains, 88.2% carried no date at all, which makes most web numbers impossible to age-check.

The word describes the metadata around a number and says nothing about whether the number is right. A figure can be perfectly accurate and still undatable: the value holds, but nothing on the page records when it was measured, so you have no way to test whether it has moved since. Unknown is its own category. It does not mean the number is wrong, and it does not mean the number is fresh.

There is no clock to check.

How We Measured Datability

This number comes straight out of the instrument. Content Health assigns a freshness_tier to a claim only when it can find a real time anchor for the underlying figure, so 88.2% is simply the share of the corpus where there was no date to score.

We ran it across the same State of Content Decay corpus, version five: 46 domains, 961 posts, 5,034 individual claims, each scored as its own unit, separate from the page it sits in. That distinction carries the finding. A page can wear a tidy publication date while the numbers printed inside it carry none, and we counted the numbers.

Nothing here calls a single figure wrong. It records only that, for 88.2% of them, there was no date to check.

Can You Trust a Statistic Without a Date

Trust in a statistic is really trust in its currency, the question of whether it still holds today. A date is what lets you return in a year and ask whether the number survived. Take the date away and that question has no answer. The figure keeps whatever level of faith you extended it the day you cited it, and nothing in your workflow will ever prompt you to check it again.

The Datable Minority Is Already Aging

Turn to the datable minority. I expected the part we could date to be the part in good shape, and the news gets no better. Of the 5,034 claims, 595 carry a real date, 11.8% of the corpus. That is the entire population your freshness process has ever been able to touch, because a review needs a date to act on, and only these 595 have one.

Look at how that slice breaks down. 98 of the 595 read as fresh. 209 are expiring, close enough to their shelf life that a careful editor would requeue them. 288 are already aged. So even inside the sliver you can check, the fresh claims are the smallest group and the aged ones almost outnumber the rest.

The part of your archive you can audit is mostly overdue.

None of this is a decay curve. We are not tracking how these numbers age over time, only photographing the composition of the checkable slice on one day. The State of Content Decay report already tells the age-over-time story. This is a narrower cut: the shape of the 11.8% you were relying on to stand in for the whole.

Now trace the citations themselves, and the picture repeats. Of 1,006 traced citations in the corpus, 549 return no usable freshness signal at all, 54.6%, even when the link resolves to a live page.

This is a different axis from the one our other studies measure, and the difference is the whole point. A source-verification pass asks whether a link still resolves, which is reachability. The citation-provenance work asks whether the chain reaches a primary source, which is provenance. This asks a third question both of them leave open: once you arrive, can you tell how old the thing you arrived at is.

A working link tells you where a number came from. Whether that source can be placed in time is a question the link was never built to answer. More than half the time, the source behind a citation is as undatable as the claim it supports. This is the water these numbers swim in. Borrowing a figure means inheriting its missing date along with its value.

Undatable Data Is the Risk You Never See

The blind spot turns into a bill right here. Among the traced, linked claims in the corpus, the ones resting on an undatable source went stale at 10.3%. The ones backed by a fresh, datable source went stale at 4.7%. Roughly twice the rate.

The method is standard and worth saying plainly. A two-proportion z-test returns p=0.015, with Wilson confidence intervals of 8.0 to 13.2% for the undatable group (n=526) and 2.6 to 8.5% for the fresh group (n=212). Pool the undatable links against every datable bucket and the gap holds: 10.3% versus 5.5%, p=0.0065. Needs-review claims were held out as uncertainty and excluded from the stale count, and born-stale claims were exempted by the usual cut. The intervals are wide and they nearly touch, so read the gap as directional: the effect is real and it replicates, but the exact multiple will move as the sample grows.

The instinct inverts here. A number stamped 2023 announces its age, trips every audit, and gets corrected, which is what makes it the safer figure to carry. A number with no date reads as eternally current and never enters the queue at all, because you cannot schedule a review for a figure you cannot age-check.

The honest-old stat is the one you fix; the dateless one just accrues.

One caution on the number itself. The State of Content Decay report also prints a 10.3% figure, for a completely different quantity: among posts that are themselves two to three years old, the share of their cited stats that are two or more years old. Same two digits, unrelated measurement. The 10.3% here is a stale rate among linked claims, split by whether their source can be dated at all.

Why a Fresh Timestamp Cannot Reach an Undatable Number

This is why a refresh pass so often changes nothing that matters. You update a post, the modified date jumps to today, and the page reads as freshly maintained. Underneath it, the undatable numbers sit exactly where they were, because a page-level timestamp has no idea which figures inside the post were ever checked against a source.

Freshness theater is the gap between the stamp and the substance: the refresh touches the page's date and leaves every undatable figure inside it untouched. And the page date itself is often unreliable. If you cannot fully trust the one date a page advertises, you certainly cannot recover the dates the numbers inside it never carried. The problem sits one level below the timestamp: the figure was never age-checkable, so there is nothing for a fresh stamp to move.

Dated Data as a Feature

There is a way out of the undatable set, and it runs through first-party data. When you measure a number yourself, its date comes free.

A poll vote lands at a recorded timestamp. A poll of your own audience closes on a day you can name. A chart wired to a sheet refreshes at a known hour. Each of those figures lands in the datable 11.8% by construction, because the act of collecting it also recorded when.

A poll vote knows exactly when it happened.

That is the whole appeal of dated data: you do not chase the date down later, because owning the measurement already recorded it.

Test it against your own habit. Pick the last statistic you published. You can almost certainly find where it came from. The harder question is whether you wrote down how old it already was.

Content Health runs this same check across a whole archive. It reads each cited page for its real publication and modified dates and takes max(published, modified) as the effective one. The scanner extracts every statistical claim and scores its staleness risk. It flags them. You decide what to fix.

I built it to stop there on purpose. It will not manufacture a date the source never published, which is the honest edge of the whole method: an undatable number stays undatable until someone measures it fresh. So the durable move is to keep publishing research you generated yourself, so more of your archive carries a date by birth and less of it leans on sources that never had one.

The Debt That Never Trips an Audit

The undatable set is the invisible balance on your books, a form of content debt that no audit ever flags, because an audit needs a date to act on and these numbers refuse to hand one over. Nothing about it reads as a problem, which is exactly how it grows.

The 88.2% is a corpus average. Yours is a number you could actually find. Since running this scan, I ask one thing before I cite anything now, and it is the question I would leave you with: how old is this, really, and could you still stand behind it if someone made you show the date?

How Fresh Is Your Content?

Paste any URL and find out which data points have gone stale.

Related Posts

The Citation Monoculture: How Few Sources SaaS Blogs Share

1,006 citations. 516 domains. The weight behaves like 104.

Jul 10, 2026

Why Unsourced Stats Are Round Numbers (2,473-Claim Analysis)

The last digit of a percentage tells you whether anyone counted it.

Jul 10, 2026

How to Turn Vague Stats Into Your Own Data

Every 'most' in your archive marks a number you never measured, and the people who could supply it are already reading the sentence.

Jun 30, 2026