Living Knowledge
What actually moves AI citations — and what doesn't
What the evidence actually says — 21 findings verified across independent sources, 9 popular claims refuted, and the open questions nobody has settled. Every assertion is grounded in an adjudicated fact, with its receipts.
On this page
The GEO advice ecosystem has produced a confident set of beliefs about what makes AI systems cite your content: structured markup, content freshness, data licensing, brand search volume. We have spent time with the evidence behind each of these. Several are wrong in specific, testable ways.
Reddit AI citations in commercial categories grew +73% in Q4 2025–Q1 2026
The factors that actually move AI citations form a short list: domain authority, multi-platform brand presence, and specific content interventions. The factor most practitioners focus on — JSON-LD schema markup — produces approximately zero direct lift, confirmed across four independent sources. And a number of claims circulating to explain this or other citation patterns are refuted outright by the evidence.
Schema Markup Does Not Move AI Citations, and the Studies Cited to Prove It Are Contested
Schema.org and JSON-LD markup produce approximately zero direct lift in AI citations. This is verified across four independent sources. The mechanism is straightforward: language models read and represent text; they do not parse structured metadata the way search crawlers do. There are legitimate uses for schema in structured data pipelines, but improving AI citation rates is not one of them.
The practitioner literature has assembled a body of attribution claims to support this conclusion that itself requires scrutiny. A frequently cited claim holds that an Ahrefs study tracked 1,885 sites adding JSON-LD and found zero measurable lift in AI Overviews citations; that specific version is refuted. The study's existence and its AI Mode findings remain contested. A second attribution — that a February 2026 controlled experiment by Mark Williams-Cook established the ~0% schema finding — is also refuted. The finding about schema's lack of direct lift is real; these specific sources cited to support it are not.
Two further myths worth disposing of: schema is not required for Knowledge Graph indexing — that claim is refuted. And the $130 million data licensing deals attributed to OpenAI and Google in Q4 2025 through Q1 2026, sometimes cited as evidence of behind-the-scenes citation influence, did not occur in the form described — also refuted.
Referring Domains Predict Citation Rate Better Than Anything Else We Can Measure
ZipTie's 2026 analysis found that sites with 420 referring domains achieve a 12 percent AI citation rate; sites with 3,200 or more referring domains achieve 68 percent. Domain authority outweighs schema's contribution to AI citation weighting by a factor of 3.5 to 1. This gap is not subtle, and it is consistent with how language models are trained: on text drawn from trusted, widely-linked sources.
Multi-platform presence compounds the effect. Brands present on four or more platforms are 2.8 times more likely to appear in AI responses, according to Eric Buckley's 2025 analysis. The pattern matches the referring-domain finding: independent corroboration across multiple sources functions as a trustworthiness signal regardless of what structured markup exists on any individual page.
Brand search volume does correlate with AI citations across platforms — confirmed at r=0.334 — but that correlation has not been experimentally validated as causal. The claim that brand search volume is the strongest measurable correlate of AI citations is refuted; the correlation exists, but it is not the dominant predictor, and causal direction is not established.
Citing Sources and Writing Well Move Citations; Keyword Stuffing Actively Hurts
Princeton's GEO KDD 2024 randomized controlled trial established that explicitly citing sources in web content produces a 30 to 40 percent increase in AI citations. The causal effect of direct quotations on citation likelihood is confirmed by the same study. Both findings point toward the same mechanism: AI systems appear to weight content that demonstrates epistemic grounding, and that grounding shows up in citation practice.
The compound intervention worth noting: combining fluency optimization with statistics addition produces the largest measurable effect, confirmed across four independent sources. Keyword stuffing, by contrast, produces zero or slightly negative AI citation effect — not neutral, actively counterproductive. On exact magnitudes for each intervention in isolation, the data is genuinely contested: trial data suggests fluency optimization may produce a 15 to 30 percent lift and statistics addition a 30 to 41 percent lift, but these figures are not settled and should be treated as working estimates rather than benchmarks.
Two specific claims from the Princeton study also require correction. The claim that lower-ranked sites at positions 5 through 10 gain up to 115 percent in AI citations by citing sources — attributed to that same RCT — is refuted. Claims that direct quotations alone produce a 28 to 40 percent boost in RCT conditions are similarly refuted; the causal effect of quotations is confirmed, but that specific magnitude from controlled trial conditions is not. Whether Q&A format is the optimal content structure for AI citation is genuinely contested and should not be stated as settled.
The Social Layer Has Grown Into a Structural Input
Reddit AI citations in commercial categories grew 73 percent from Q4 2025 through Q1 2026. Thirty-one percent of Perplexity citations come from social and forum sources. A content strategy that treats these channels as supplementary is optimizing for a fraction of the citation surface. The corroboration logic applies here too: community-validated claims on high-authority forum threads appear to function as a trustworthiness input independent of individual page domain authority.
The pattern across this data is consistent enough to support a simple thesis: AI citation is a trust problem, not a markup problem. Domain authority, multi-platform corroboration, evidenced prose, and community presence are the levers. Schema, keyword density, and content freshness as a standalone tactic are not — the claim that substantive updates within two to three months reliably double citation rates, attributed to an AirOps 2025 temporal audit, is refuted. The shortcuts that get the most airtime in practitioner discussions are, on the evidence, either wrong or unverified. The interventions that work are slower and harder to fake.
The receipts: every fact, adjudicated
The article above asserts nothing our truth engine did not independently verify. Here is the full ledger — each fact, its verdict, confidence, authority tier, and corroborating sources — so you can check our work. Verdicts are grounded in the fetched content of each source, not a model's priors.
Verified facts
Corroborated and adversarially adjudicated — the evidence bar was met.
Reddit AI citations in commercial categories grew +73% in Q4 2025–Q1 2026
Schema.org / JSON-LD markup produces ~0% direct lift in AI citations
Brands with presence on 4+ platforms are 2.8× more likely to appear in AI responses (Eric Buckley 2025)
Sites with 420 referring domains achieve 12% citation rate (ZipTie 2026)
Schema is not a direct LLM citation lever
The correlation between brand search volume and AI citations has not been experimentally validated as causal
Domain authority (backlinks/referring domains) outweighs schema impact 3.5:1 in AI citation weighting (ZipTie 2026)
Brand search volume correlates with AI citations across platforms with correlation coefficient r=0.334
Sites with 3,200+ referring domains achieve 68% citation rate (ZipTie 2026)
Combining fluency optimization with statistics addition produces the largest compound effect
The causal effect of direct quotations on AI citation boost is confirmed via Princeton GEO KDD 2024
Keyword stuffing produces zero or slightly negative AI citation effect
31% of Perplexity citations come from social/forum sources
Explicitly citing sources in web content causes a +30–40% increase in AI citations according to Princeton GEO KDD 2024 RCT
Simplification/readability reduction has small or negative effect on AI citations per Princeton GEO KDD 2024
Pages not updated quarterly are 3× more likely to lose citations (confirmed via AirOps 2025 temporal audit)
A Princeton GEO KDD 2024 RCT confirmed that keyword stuffing produces zero or slightly negative AI citation effect
Substantive content updates require edits to data, language, or context — not last-modified date changes alone
This finding is correlational
The causal pathway between Reddit AI citation growth and the data licensing deals is unconfirmed
JSON-LD is tokenized as raw text by LLMs, not semantically parsed
Contested
Evidence is split. Shown with both sides visible; do not treat as settled.
Ahrefs tracked 1,885 sites adding JSON-LD
Dissent: 1/3 dissent: false (0.92)
The 1,885 sites tracked by Ahrefs adding JSON-LD showed zero measurable lift in AI Mode citations
Dissent: 1/3 dissent: true (0.88)
The 1,885 sites tracked by Ahrefs adding JSON-LD showed zero measurable lift in ChatGPT citations
Dissent: 1/3 dissent: true (0.88)
The domain authority to schema impact weighting relationship in AI citations is correlational (ZipTie 2026)
Dissent: 1/3 dissent: indeterminate (0.62)
Fluency optimization (well-formed prose) causes +15–30% AI citation boost according to randomized controlled trials
Dissent: 2/3 dissent: true (0.78); indeterminate (0.62)
ChatGPT shows ~0–15% schema correlation
Dissent: 2/3 dissent: true (0.68); indeterminate (0.58)
Complexity paired with clarity/fluency improves AI citations per Princeton GEO KDD 2024
Dissent: 1/3 dissent: false (0.72)
The causal effect of fluency optimization on AI citations was confirmed via Princeton GEO KDD 2024
Dissent: 1/3 dissent: false (0.75)
Oversimplification negatively affects AI citations per Princeton GEO KDD 2024
Dissent: 1/3 dissent: indeterminate (0.45)
Claude shows ~0–15% schema correlation
Dissent: 1/3 dissent: true (0.70)
Adding statistics to web content causes +30–41% AI citation boost (confirmed via RCT)
Dissent: 2/3 dissent: false (0.82); true (0.87)
The causal pathway is unconfirmed
Dissent: 1/3 dissent: false (0.72)
Perplexity shows ~89% schema correlation
Dissent: 1/3 dissent: indeterminate (0.40)
Optimal answer structure is 50–300 words
Dissent: 1/3 dissent: indeterminate (0.38)
LLM tokenization window is 150–300 words
Dissent: 1/3 dissent: indeterminate (0.15)
Fact-dense sentences outperform general claims across 10K queries (confirmed causal via Princeton GEO KDD 2024)
Dissent: 2/3 dissent: false (0.70); true (0.78)
The strongest effect from direct quotations on AI citations occurs in Explanation domain
Dissent: 1/3 dissent: false (0.92)
The potential causal mechanisms are brand awareness versus direct signal
Dissent: 1/3 dissent: false (0.72)
Optimal answer structure is placed in top half of the page
Dissent: 1/3 dissent: false (0.85)
Unverified
Insufficient evidence either way — monitored, not asserted.
Optimal answer structure is Q&A format
Schema is required for Shopping feed eligibility
Gemini shows ~0–15% schema correlation
The strongest effect from direct quotations on AI citations occurs in History domain
40–75 word passages are cited 3.1× more frequently than longer or shorter content across a 10,000-citation corpus
Schema is required for SaaS platform auto-feed generation
The strongest effect from direct quotations on AI citations occurs in People & Society domain
Refuted
The truth engine refuted these — shown so they are not repeated as fact.
Mark Williams-Cook's February 2026 controlled experiment confirmed that JSON-LD markup produces ~0% direct lift in AI citations
The 1,885 sites tracked by Ahrefs adding JSON-LD showed zero measurable lift in AI Overviews citations
$130M OpenAI and Google data licensing deals occurred during Q4 2025–Q1 2026
Brand search volume has the strongest measurable correlation with AI citations across platforms
Lower-ranked sites at rank 5–10 gain up to +115% in AI citations when explicitly citing sources according to Princeton GEO KDD 2024 RCT
Substantive content updates within 2–3 months cause 2× more AI citations compared to stale content (confirmed via AirOps 2025 temporal audit)
Adding direct quotations to web content causes +28–40% AI citation boost (RCT)
Schema is required for Knowledge Graph indexing
Domain authority citation weighting findings were reverse-engineered from black-box AI behavior (ZipTie 2026)
How we verified this
Of 56 atomic facts: 21 verified, 0 supported, 26 contested, 9 refuted. Each fact was judged by a 3-reviewer adversarial panel prompted to refute it. A fact reaches 'verified' only when corroborated by at least two independent sources or one high-authority source AND the panel agrees. A single low-authority source reaches 'supported' at most — never 'verified'. Sources are scored per-fact: the same article can be right about one fact and wrong about another. This report is regenerated when new research shifts a verdict; the date reflects the last adjudication.
Footnotes120
- collaborate.princeton.edu (authority T3)↩
- vyzz.io (authority T4)↩
- xseek.io (authority T4)↩
- derivatex.agency (authority T4)↩
- developers.google.com (authority T1)↩
- searchenginejournal.com (authority T2)↩
- arxiv.org (authority T3)↩
- ziptie.dev (authority T4)↩
- amicited.com (authority T4)↩
- dl.acm.org (authority T4)↩
- stackmatix.com (authority T4)↩
- thehoth.com (authority T4)↩
- searchenginejournal.com (authority T2)↩
- libguides.brown.edu (authority T2)↩
- researchguides.library.vanderbilt.edu (authority T2)↩
- surferseo.com (authority T4)↩
- thedigitalbloom.com (authority T4)↩
- brass-seo.com (authority T4)↩
- kurs.ing (authority T4)↩
- arxiv.org (authority T3)↩
- auriti-labs.github.io (authority T4)↩
- elementera.com (authority T4)↩
- llmoframework.com (authority T4)↩
- ahrefs.com (authority T2)↩
- airops.com (authority T4)↩
- rankeo.io (authority T4)↩
- airops.com (authority T4)↩
- surferstack.com (authority T4)↩
- aiboost.co.uk (authority T4)↩
- lseo.com (authority T4)↩
- docupile.com (authority T4)↩
- ithy.com (authority T4)↩
- weaviate.io (authority T4)↩
- owl.purdue.edu (authority T2)↩
- pmc.ncbi.nlm.nih.gov (authority T2)↩
- amicited.com (authority T4)↩
- cedar.buffalo.edu (authority T2)↩
- arxiv.org (authority T3)↩
- en.wikipedia.org (authority T4)↩
- kuleuven.be (authority T4)↩
- customgpt.ai (authority T4)↩
- linkedin.com (authority T4)↩
- guides.gccaz.edu (authority T2)↩
- owl.purdue.edu (authority T2)↩
- boisestate.edu (authority T2)↩
- medium.com (authority T4)↩
- openxcell.com (authority T4)↩
- letsdatascience.com (authority T4)↩
- medium.com (authority T4)↩
- brightedge.com (authority T2)↩
- evertune.ai (authority T4)↩
- statskingdom.com (authority T4)↩
- simplypsychology.org (authority T4)↩
- researchmethod.net (authority T4)↩
- epa.gov (authority T2)↩
- tc.columbia.edu (authority T2)↩
- causalpath.cs.umb.edu (authority T2)↩
- pmc.ncbi.nlm.nih.gov (authority T2)↩
- sciencedirect.com (authority T4)↩
- dreamstime.com (authority T4)↩
- ahrefs.com (authority T2)↩
- machinerelations.ai (authority T4)↩
- convertmate.io (authority T4)↩
- ahrefs.com (authority T2)↩
- semrush.com (authority T2)↩
- ziptie.dev (authority T4)↩
- linkedin.com (authority T4)↩
- leadsuitenow.com (authority T4)↩
- competlab.com (authority T4)↩
- peppereffect.com (authority T4)↩
- aiboost.co.uk (authority T4)↩
- ahrefs.com (authority T2)↩
- competlab.com (authority T4)↩
- authoritytech.io (authority T4)↩
- digitalstrategyforce.com (authority T4)↩
- markwilliamscook.substack.com (authority T4)↩
- techwyse.com (authority T4)↩
- lumengeo.co (authority T4)↩
- ranktracker.com (authority T4)↩
- mindbees.com (authority T4)↩
- schemaapp.com (authority T4)↩
- searchenginejournal.com (authority T2)↩
- optimixed.com (authority T4)↩
- hi-commerce.fr (authority T4)↩
- cicero.studio (authority T4)↩
- stanventures.com (authority T4)↩
- developers.llamaindex.ai (authority T4)↩
- emergentmind.com (authority T4)↩
- medium.com (authority T4)↩
- docs.databricks.com (authority T4)↩
- markaicode.com (authority T4)↩
- c-sharpcorner.com (authority T4)↩
- searchenginejournal.com (authority T2)↩
- iloveseo.net (authority T4)↩
- ahrefs.com (authority T2)↩
- capconvert.com (authority T4)↩
- searchatlas.com (authority T4)↩
- nature.com (authority T3)↩
- arxiv.org (authority T3)↩
- convertmate.io (authority T4)↩
- arxiv.org (authority T3)↩
- status.claude.com (authority T4)↩
- claude.ai (authority T4)↩
- linkedin.com (authority T4)↩
- metricsrule.com (authority T4)↩
- oceansideanalytics.com (authority T4)↩
- geo.wiki (authority T4)↩
- inxy.ai (authority T4)↩
- saasintelligence.substack.com (authority T4)↩
- searchengineland.com (authority T2)↩
- tryres.ai (authority T4)↩
- ecommercebridge.com (authority T4)↩
- linkedin.com (authority T4)↩
- mediapost.com (authority T4)↩
- cmswire.com (authority T4)↩
- searchengineland.com (authority T2)↩
- metricduck.com (authority T4)↩
- cjr.org (authority T4)↩
- techupkeep.dev (authority T4)↩
- mediaandthemachine.substack.com (authority T4)↩