Because when a language model decides which source to believe, it doesn't evaluate tone, references, or perceived authority — it evaluates how directly the text answers the question it was asked. A University of California, Berkeley study (ACL 2024) measured this by comparing 2,208 real evidence paragraphs against 238 controversial questions across 144 categories: stylistic changes — adding scientific references, sounding more objective, adding more information — had a neutral or negative effect on how convincing the text appeared to the model. Changes that increased direct relevance to the question, on the other hand, substantially improved that persuasiveness.
Of the eight criteria Nostos measures, this is the one that most directly decides whether a model cites you when someone asks something you can answer.
For 238 contentious questions across 144 categories, they compared pairs of real web paragraphs — one answering "yes" and one answering "no" — and measured which one convinced five different models more (GPT-4, Claude, LLaMA-2, Vicuna, WizardLM). Semantic similarity between the question and the paragraph strongly predicted the outcome across nearly every model. Text readability and the number of unique words predicted nothing.
Neutral or negative. Adding scientific references, adding unrelated information, sounding more objective or more confident — none of those stylistic changes improved the odds that the model would pick that paragraph as the correct answer. What did work was a relevance change: prefixing the paragraph with a sentence that restated the question — "The following text is about the question: [question]" — was enough to substantially improve its win-rate against the opposing paragraph, while barely changing the content.
That the section answering your own question needs to say so in the first lines, using the question's own words, before giving context or background. That's exactly what this criterion measures in Nostos: if a section's heading asks something, the answer belongs in that section's first paragraph, not the third.
In yesterday's analysis (08/14/2026), Nostos measured 70/100 in Direct Question-Answer Relevance: most of the content does answer the question implied by its own heading, but several sections still open with context or narrative before reaching the answer.
The room for improvement is tactical, not structural: move the answer that's already written to the start of each section, without rewriting the content.
Check every H2 on your site and ask whether the first sentence below it answers the heading above it. If not, move the answer to the top and push the context down. Three places where this fails most often:
Get your GEO score across the eight criteria backed by academic research, including your current Direct Question-Answer Relevance.
Run the analysis at nostosgeo.app