Nostos GEO
GEO · CONTENT AND DATA

Why having lots of content isn't enough: the role of verifiable data in AI visibility

By Fabián Torres  ·  June 24, 2026  ·  7 min read

There's a widely held idea in content marketing: more content equals more visibility. For years, that logic worked reasonably well for SEO. But when the discovery channel is an AI — ChatGPT, Gemini, Perplexity, Claude — the equation changes fundamentally.

+40%
increase in source visibility when citations, quotations from relevant sources and verifiable statistics are added to content

That number doesn't come from a marketing agency or an SEO blog. It's the result of an experimental benchmark published at KDD 2024 by Princeton University researchers. What it measures is concrete: the effect of adding citations, quotations from relevant sources, and verifiable statistics to content. Together, those three methods exceed a 40% visibility increase in the benchmark. On a real engine —Perplexity— the measured improvement was up to 37%.

The conclusion is direct: for AIs, word count matters less than the density of verifiable evidence. A site with 20 articles backed by real data competes better than one with 200 opinion pieces, generic guides, or unsourced lists.

Why LLMs prefer data over words

Language models don't read content the way a human does. A human can follow an argument with no data and evaluate it by internal coherence or style. An LLM, on the other hand, looks for patterns associated with trustworthiness: dates, figures, institution names, verifiable references, authorship with real credentials.

When a model decides which source to cite in an answer, it isn't choosing "the most complete article" or "the best-known blog." It's choosing the content that has the most signals that someone verified what it says. A sourced number is one of those signals. A percentage with an academic citation is another. A statistic with a year and institutional authorship is another.

"We find that including citations, quotations from relevant sources, and statistics can significantly boost source visibility, with an increase of over 40% across various queries." Aggarwal et al., Princeton University — Proceedings of KDD 2024 (verbatim)

Put another way: AIs function like academic reviewers. They don't just read, they evaluate. And what they evaluate is whether the content has an empirical basis or not.

The problem with unanchored content

Most web content is written to persuade, not to verify. Phrases like "SEO is essential for digital growth" or "customer experience is key to retention" are claims anyone could make, at any time, with no data behind them. For a traditional search engine, that kind of content can rank well if it has the right keywords. For an LLM, it provides no trustworthiness signals at all.

We call this unanchored content: text that floats without a data point tying it to reality. It can be well written, well structured, and cover a relevant topic. But in the eyes of a language model, it's indistinguishable from anyone's opinion.

The problem isn't writing badly. The problem is writing without evidence. An LLM can't verify a claim with no source, and what it can't verify, it tends not to cite.

What counts as verifiable data, exactly

Not every number is verifiable data. "We increased sales by 300%" with no context or source gives an LLM nothing. What does add value:

Statistics with institutional origin. A percentage published by a university, a chamber of commerce, an official body, or an academic paper carries weight because there's an institution backing it and a date that contextualizes it.

References to research with author and year. Citing "Aggarwal et al., Princeton, 2024" is different from saying "according to recent studies." The first is verifiable; the second is noise.

Proprietary data with an explicit methodology. "We analyzed 150 Argentine websites with our tool between January and June 2026 and found that 78% had no Schema.org implementation" is verifiable data if you explain how it was obtained. It doesn't require an academic paper — it requires methodological transparency.

Measurable before-and-after results. A case study with real metrics — initial score, final score, actions taken — is worth more to an LLM than five paragraphs of argument with no numbers.

Case study — what happens when there's no data

Case study — Independent consultant · June 2026

The site of a marketing consultant with more than 60 published articles, good navigation structure, and an active blog with monthly posts. No article included sourced statistics. Every claim was of the "most companies," "experts agree," "it's estimated that" variety. No visible publication date on the posts.

42
Initial score
71
Projected score

The Data density criterion came in at 15/100. The Freshness and updating criterion came in at 20/100, because no article had a visible date or time marker. Content volume did not compensate for the lack of verifiable data. The work plan prioritizes adding sourced statistics to the 10 highest-traffic articles and adding visible publication dates to all existing posts.

What changes when content has an anchor

The change isn't just quantitative. When a site incorporates sourced, verifiable data, it changes how an LLM "reads" it. The model can associate the content with a verified knowledge corpus — papers, institutions, dates — which increases the likelihood it will use it as a reference in its answers.

Data also has a structural advantage: it's specific. Specificity is a citability signal. "40% of Argentine companies have no website" is more citable than "many Argentine companies still aren't online." The first can be cited with precision; the second can't.

That's why adding verifiable statistics to existing content — what Aggarwal et al. call Statistics Addition — is one of the three interventions the paper identifies as able to significantly boost a source's visibility, alongside including citations and quotations from relevant sources. It doesn't require rewriting the entire site. It requires anchoring existing claims to real data.

What this means for your content strategy

If you're producing content with AI visibility in mind, the question you need to ask for each article isn't "is it well written?" or "does it have the right keywords?" The question is: is there at least one sourced, verifiable data point in this piece?

If the answer is no, that article is invisible to LLMs, regardless of its length or editorial quality.

The good news is this is technically fixable. Existing content doesn't get thrown out — it gets anchored. You identify the central claims in each piece and replace vague statements with statistics from a verifiable source. It's work, but it's work with a clear direction and a measurable result.

Do you know how much verifiable data your site has?

Nostos's free analysis evaluates your site against 6 GEO criteria — including Data density — and tells you exactly where you stand and what to change.

Analyze your site for free