The word "most" holds a seat for a percentage you never had. Somewhere in your archive sits a sentence like "most professionals use HubSpot," shipped on a deadline with a guess where a measurement belonged, and the people it guesses about are the ones reading it. To turn vague stats into your own data, you ask them: put the question to the readers the claim already describes, and the figure that replaces "most" becomes a statistic that exists on your page before it exists anywhere else.
The reflex repair runs the other direction. A claim feels thin, you search for a study, you cite it, and your reader now knows which other site to trust for the number. Polling the audience keeps the measurement where the claim lives, and every soft quantifier in your back catalog marks a page where that trade is available this week.
What It Means to Turn Vague Stats Into Your Own Data
Turning a vague stat into your own data means replacing the guess inside a sentence like "most professionals use HubSpot" with a number your audience reported about themselves. You publish a short poll on the page carrying the claim, readers answer from their own behavior, and the tally becomes a first-party statistic with your page as its point of origin. The sentence already supplied the question and named the population; the poll is the part nobody ran.
A Vague Quantifier Is a Skipped Measurement
You reach for "most" at the exact moment the data fails you. The deadline holds, the sentence needs its count, the count is missing, so a word absorbs the job a number should have done. Once published, nothing flags it. Editors read past it. It looks like style.
The monitoring tools read past it too. LiquiChart's claim extractor scans a page for statements it can track over time, and tracking needs an anchor: a number, a date, a source. Hand it "72% of teams invest in content marketing" and it has a value to watch. Hand it "most professionals use HubSpot" and there's nothing to hold, so the sentence falls straight through.
I kept running into this in our decay studies: the sentences that most need measuring are the ones every scanner discards before measuring begins, so any freshness audit built on extraction undercounts the whole category. Our study of claims that cite no source covers a neighboring failure, where a number is present and its backing is missing. A vague quantifier fails one step earlier. The number never arrived at all, so no instrument graded on numbers will ever surface it.
Your Readers Already Hold the Missing Number
The recovery path sits inside the claim's own grammar. "Professionals" names a group, and the group reading a post about HubSpot workflows is made of professionals weighing HubSpot. Write a sentence about what your audience does and you've written a question your audience can answer. You framed that question the day you drafted the sentence; nobody has put it to them.
Seen that way, the back catalog becomes a worklist. Each "most" and "many" marks a page where one question would swap an assertion for a measurement, and a page carrying its own measurement earns what Google's systems reward as original data: information the page adds rather than repeats. The mechanics of writing and embedding the question live in how to create a poll for a blog; the judgment upstream of that is choosing which sentence deserves it.
Think about the last time a draft needed a statistic you didn't have, because writers split four ways at that moment.
Each of those four paths sets a different ceiling on how far a reader can trust the sentence. The population that sentence describes is already here, reading it. No one has put the question to them yet.
Whichever way you split, only one of the four ends with a number that starts on your page.
Which Vague Stats Convert Into Poll Data
Not every "most" survives the trip. Ask readers which CRM they use and they answer from their own week. Ask them whether most teams use one and they guess about strangers, and a pile of guesses measures what your audience believes about other people rather than the thing your sentence asserts. Only the first kind of question rebuilds the missing figure.
The bar is concrete: a vague stat converts when readers can self-report their own behavior, preference, intent, or belief, and when those reports sum to the exact subject of the claim. "Do you use a formal content calendar?" clears it, because each reader answers from experience and the tally is the statistic. "Do you think most teams use a formal content calendar?" polls a belief about a population the reader has never seen, which is the difference between self-reported data and opinion.
There's a mechanical check. Rewrite the vague sentence as a question aimed at one reader about themselves. If it reads like something you'd ask a colleague and the answers would add up to your missing figure, convert it. If the only natural phrasing asks the reader to estimate what others do, predict the future, or rank options they haven't tried, leave the sentence alone, because forcing the poll hands you a number with nothing behind it.
Why Citing a Study Costs You the Number
The borrowed stat feels responsible. It's checkable, it's fast, and it lets the draft move. It also files your page into the long line of pages pointing at the same source, the pattern behind why AI cites third-party sources instead of the sites repeating them. Every citation is a signpost toward the domain that did the measuring, and once you run the poll the signposts turn around: the next writer who needs that figure has to link to you.
A borrowed stat makes your page a stop on the way to the number. A measured one makes it the destination.
The fair objection is sample quality, since readers who vote in an embedded poll chose themselves. A poll journalists will cite survives that objection through disclosure: name the audience, state the count, and claim only what the sample supports. "Readers of this page report X" is bounded, checkable, and something no other site can say.
How the Originality Engine Finds Poll Candidates
Detection runs in two passes over a page you own. The first pass is a deterministic sweep for the telltale shape, a soft quantifier governing a population with no number anywhere in the sentence, and "most professionals use HubSpot" is the canonical hit. This pass is deliberately loose, so it forwards more candidates than deserve to survive.
The second pass is a veto. It takes each flagged sentence and asks one thing of it: could a reader of this page answer from their own behavior, and would those answers aggregate into the figure the sentence lacks? The veto starts from yes and hunts for a reason to decline. A claim that would poll guesses about strangers dies here, and so does anything the model can't call confidently, since below a fixed confidence threshold the resolution is a skip.
Why the Engine Stays Silent Most of the Time
What survives both passes arrives as a single recommendation, "Back this claim with your own poll," carrying a proposed question written in the second person plus a starter set of mutually exclusive options. Accepting it opens poll creation with the question pre-filled, and you shape the options and publish. The engine's work ends at the suggestion. It never creates the poll, never edits your draft, and never invents the figure, which is why "76% of professionals use HubSpot" appears here only as a worked example of what your audience might eventually report.
I designed the veto around a cost asymmetry: a recommendation that fires wrongly costs more trust than a real candidate the engine lets pass. A suggestion you can act on every single time beats a longer list you'd have to triage, so when the engine can't confirm a claim is self-reportable, you see nothing at all.
Turning Your Back Catalog Into a Worklist
Point the engine at a URL you've already published. It ingests the page as a Monitored Page, runs detect and veto across every sentence, and returns the convertible quantifiers as a short list of proposed poll questions. Pick one, publish the poll, embed it back into the post. Once votes arrive, the result becomes a first-party claim the platform tracks, and your Originality Score increases with the share of claims backed by your own data.
For a first look before you convert anything, the Content Health Scanner reads any URL without a login and returns its unattributed and stale claims. Those findings are the other half of the problem: sentences that carry a figure but have lost the backing behind it, sourceless numbers and citations pointing at dead pages. A scan graded on numbers and sources has nothing to grip in a sentence containing neither, and that remaining half belongs to the originality engine.
The Number Stays Current After You Measure It
Conversion would be a poor trade if the new figure aged like the old guess. The poll result sits in the post inside a living content block, and when later votes move the leading answer, the block rewrites its own sentence to match. The gap you closed this quarter stays closed for as long as readers keep voting.
The Page Where the Number Starts
Every "most" you leave in the archive is a figure you're donating to whoever measures it first. The readers who could supply it show up on schedule, read the sentence that guesses about them, and leave without being asked. The same logic reaches forward into whatever you write next, where the stronger opening move is building utility content instead of chasing keywords: generate the data the topic needs.
Run one poll against one vague sentence this week. When the votes land, a statistic will exist with exactly one address on the internet, and the address is yours.