Phase 3 · Content & Citation · L2 → L3

Earn LLM recommendations with prose that answers real buyer questions.

Citations come from content quality, not markup tricks. LLMs quote statistics, comparisons, and cited prose. This phase builds the editorial authority and entity presence that makes your brand memorable to AI systems.

L2 → L3Intermediate–AdvancedOngoingEditorial + Dev

·

Why this matters

TL;DR

Citations come from content quality, not markup tricks. LLMs quote statistics, comparisons, and cited prose — not keyword-stuffed pages. Phase 3 builds the editorial authority and entity presence that makes your brand memorable to AI systems trained on the open web.

Structured data gets you included; authoritative content gets you cited. LLMs learn from prose — statistics, comparisons, and cited sources in your writing become the raw material for future recommendations.

Most brands complete Phase 1 and 2 and wonder why citations haven't improved. The answer is content: AI systems are trained on the open web and learn to cite sources that are specific, attributed, and verifiable. Feature lists are not quotable. Cited statistics are.

Phase 3 steps

Work through these in priority order — P0 items have the highest citation impact.

P1

Write conversational content with statistics, comparisons, and cited sources

What empirically earns an AI citation: statistics, direct quotations, explicitly cited sources, fluent prose. Not markup, not keywords. A 2024 Princeton/SIGKDD study found that citing external sources increased AI citation rate by +115%, adding statistics by +41%, including direct quotations by +28%. Keyword stuffing reduced it by −10%.

Impact
Very high
Effort
High
Time
Ongoing
Owner
Content / Marketing
How to
  1. Audit your top 5 landing pages: do they contain at least 3 cited statistics? At least one direct quotation with attribution? Explicit comparison tables vs. alternatives?
  2. Add a "Key statistics" section to each major page. Format: "[X]% of [audience] [action], according to [Source, Year]." Hyperlink to the primary source.
  3. Include comparisons: "Unlike [Alternative], YourProduct does X." AI systems trained on the web learn to cite sources that make comparisons — they are more useful for answer generation.
  4. Structure prose for AI extraction: short paragraphs, clear topic sentences, named entities. AI models extract claims from sentences, not from paragraphs — each sentence should be self-contained and citable.
P0

Build Wikipedia / Wikidata entity presence

Wikipedia is the single strongest predictor of LLM citation accuracy — it is a direct training corpus for every major model. Wikidata gives @id linkage even without a full article, allowing AI systems to anchor facts to your brand entity even if knowledge is sparse.

Impact
Very high
Effort
High
Time
1–4 weeks
Owner
Marketing / PR
How to
  1. Check if your brand already has a Wikipedia article: en.wikipedia.org/wiki/YourBrand. If yes, ensure it is accurate, has citations, and includes your core products and founding date.
  2. If no article exists: build notability first. Wikipedia requires coverage in 3+ independent, reliable secondary sources (news articles, industry publications). GEO strategy note: this is the single highest-ROI GEO activity for brands without existing coverage.
  3. Create a Wikidata entry even without a full Wikipedia article: wikidata.org/wiki/Special:NewItem. Add: instance of (Q4830453 for business), name, website URL, founded date, industry. This alone provides @id linkage.
  4. Add your Wikidata and Wikipedia URLs to your Organization schema sameAs array (covered in Phase 2).
P1Quick win

Set up AI referral traffic tracking in GA4

ChatGPT, Perplexity, Claude, and Gemini send referral traffic with identifiable hostnames. Without custom channel grouping, AI referral appears as Direct or Organic in GA4 — you cannot measure GEO progress or attribute content performance to AI sources.

Impact
High
Effort
Low
Time
1–2 hours
Owner
Analytics / Marketing
How to
  1. In GA4: Admin → Data Settings → Channel groups → Create custom channel group named "AI Referrals".
  2. Add the AI referral hostnames to match:
    TEXT
    # GA4 custom channel — AI referral hostnames to match
    chat.openai.com
    chatgpt.com
    perplexity.ai
    claude.ai
    gemini.google.com
    bard.google.com
    bing.com (Microsoft Copilot traffic)
    you.com
    phind.com
  3. Set condition: Source contains any of the above hostnames. Use regex if GA4 supports it: (chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com)
  4. Create a GA4 exploration report: dimension = Session source/medium, segment = your AI Referrals channel. Track weekly. Any volume above 0 is Phase 3 working.
P2

Keep /llms.txt updated with your latest content

Stale links (404s) in llms.txt confuse LLMs that try to fetch them for context. Review on each major content update. An llms.txt with dead links is worse than no llms.txt — it signals to crawlers that the site is unmaintained.

Impact
Medium
Effort
Low
Time
30 min per review
Owner
Dev / Content
How to
  1. Audit your /llms.txt quarterly: fetch every URL listed and check for 404s. Run: cat llms.txt | grep -Eo "https?://[^ )]*" | xargs -I{} curl -o /dev/null -s -w "%{http_code} {}\n" {}
  2. Update the file whenever you publish a major blog post, launch a new product page, or deprecate a page.
  3. If your site has more than 50 content pages, maintain both /llms.txt (curated highlights) and /llms-full.txt (full content index). llms.txt should contain your 10–20 most authoritative pages, not a dump of every URL.
P1

Track LLM cold recall quarterly via Hidden Layer audits

Cold recall measures whether an LLM can accurately describe your brand from training data alone — without any retrieval. It is the primary signal of long-term GEO progress. Improvement in cold recall = real brand authority building in AI systems, not just technical compliance.

Impact
High
Effort
Low
Time
1 hour per quarter
Owner
Marketing
How to
  1. Baseline test: Ask GPT-4, Claude, and Gemini the same prompt: "Describe [YourBrand] and what they offer." Record the responses verbatim.
  2. Score each response on: accuracy (are the facts correct?), completeness (are your key products named?), sentiment (positive/neutral/negative?), confidence (does the LLM hedge or state facts directly?).
  3. Run the same test every 90 days. Track deltas. GEO activities (Wikipedia, citations, content) take 3–6 months to propagate through training data — quarterly cadence is the minimum meaningful interval.
  4. Use Hidden Layer's geo_cold_recall audit score (available at /audit) as a consistent benchmark — it uses a standardized prompt set across 3 LLMs.

What empirically moves AI citations

A 2024 study tested which content changes actually increased visibility in AI-generated answers across 10 LLM systems. Source: Aggarwal et al., “GEO: Generative Engine Optimization,” arXiv:2311.09735 (SIGKDD 2024).

Content strategy impact on AI citation visibility — Princeton / SIGKDD 2024
Content strategyVisibility changeSignal strength
Cite external sources+115%Largest effect, especially for lower-ranked pages
Add statistics+41%Data-backed claims lift citation rate
Include direct quotations+28%Attributable quotes increase credibility
Keyword stuffing−10%Treated as low-quality signal — hurts visibility

Full paper: arXiv:2311.09735 — Aggarwal et al., SIGKDD 2024.

What AI actually sees

This is what an AI-generated answer looks like when one brand has cited authority and another relies on feature lists. Note how the competitor with analyst coverage gets confidently cited, while the uncorroborated brand triggers hedging language.

AI-generated answersimulated

When researching enterprise PIM solutions, Akeneo is frequently cited in analyst reports and has been referenced by Gartner and Forrester in their PIM category analyses. Their content includes specific statistics on product data completeness and direct comparisons with competing platforms.

YourBrand operates in the same space — however, their content consists primarily of feature lists without cited statistics or external references. [LLM cannot confidently cite — no corroborating sources found]

Signal analysis
  1. PASS
    Cited authorityCompetitor appears in analyst reports and has corroborated facts across multiple independent sources — exactly what LLMs look for.
  2. THIN
    Feature-list contentYour content lists features but lacks statistics, comparisons, or cited sources — AI systems cannot confidently quote it.
  3. BLOCKED
    No corroborationWithout external citations or Wikipedia entity, LLM confidence is low — it hedges rather than cites.
Simulated AI answer — illustrates the citation gap between a content-authoritative brand and a feature-list-only site

Key takeaways

After completing Phase 3, you should have:
  • Key landing pages contain statistics, comparisons, and cited sources that LLMs can quote.
  • Your brand has a Wikipedia article or Wikidata entry establishing entity presence.
  • AI referral traffic (ChatGPT, Perplexity, Claude, Gemini) is tracked separately in GA4.
  • LLM cold recall is baselined and scheduled for quarterly measurement.