Hidden Layer/Research/Conversational vs Structured Attributes: What Humans Read vs What AI Reads
RENDERDOMSTYLESCHEMANETWORKGEO FUNDAMENTALS

GEO Fundamentals

Conversational vs Structured Attributes: What Humans Read vs What AI Reads

Schema markup gates inclusion. Prose attributes earn citations. They are not two artifacts — they are two renderings of the same Golden Record. Here is what the evidence says about which layer does which job, and what the practical implications are for product content strategy.

Ask most content teams whether they optimize their product attributes for AI, and they will tell you about their schema markup. Ask the same teams whether their product descriptions are written to be cited in a generative answer, and you will get a blank stare. Both answers reveal a misconception: that the machine-facing layer and the human-facing layer are different things. They are not. They are the same underlying record, rendered differently — and understanding the difference is what determines whether a product is included in AI results at all, and whether it earns a citation once included.

The confusion persists because structured data and prose attributes look like alternatives. They appear in separate systems — schema markup in a JSON-LD block, product copy in a CMS field — and they serve visually different purposes on a page. The evidence, however, is unambiguous about what each layer actually does in AI retrieval: structured attributes gate machine-layer inclusion; conversational prose attributes gate citation quality. One is a binary prerequisite. The other is the competitive variable. Neither substitutes for the other.

This post covers the two-layer model for product content. For the citation-mechanics deep-dive (why schema does not move AI citation rates and what the controlled evidence shows), see Schema.org Is Infrastructure, Not a Citation Signal.

The two-layer model: what each attribute type actually does

The difference between structured and conversational attributes is not stylistic — it is functional. They operate at different layers of the AI discovery stack and are evaluated by different signals at different moments in the retrieval pipeline.

Structured / machine attributesConversational / prose attributes
ExamplesGTIN, price, availability, currency, dimensions, category taxonomy, Merchant feed fieldsUse-case descriptions, comparative claims, FAQ Q&A, compatibility notes, causal/benefit copy
Primary readerLLM crawlers, Shopping agents, Merchant feed parsers — machines onlyHumans reading search results; LLMs synthesizing conversational answers
Crawl mechanismServer-rendered JSON-LD or HTML — LLM crawlers (GPTBot, ClaudeBot, PerplexityBot) do NOT execute JS; client-rendered schema is invisibleServer-rendered prose in page body; same render gap applies — must be in static HTML
What it gatesINCLUSION — binary gate for Shopping/Knowledge Graph/rich results. Missing GTIN = excluded from AI Shopping. 5CITATION QUALITY — what earns a mention in a generative answer, once included
Effect on AI citations≈0% direct citation lift. Confirmed by Ahrefs 1,885-page study (May 2026) + searchVIU direct-fetch + StanVentures. 17Statistics +32.8–37%, quotations +42.6%, cited sources +27.7–31.4% citation lift (Princeton RCT, causal). 2
Effect of keyword stuffingN/A — structured fields are controlled vocabulary, not proseZero or negative. Princeton RCT: no lift. ACM 2026: 99.78% blocked in AI retrieval. 23
Dominant quality signalCompleteness + consistency (feed matches on-page schema; Google cross-validates both 8)E-E-A-T r≈0.81 dominant predictor; domain authority r≈0.18 (~3% of citation variance explained). 4
Maintenance failure modeGTIN absent or feed mismatched → disapproval, Shopping exclusionGeneric copy, keyword-stuffed prose → AI retrieval ignores or down-weights

Schema is plumbing, not pressure

The clearest finding in the recent structured-data literature is also the most misunderstood: schema markup does not measurably increase the probability that a page is cited inside a generative AI answer. An Ahrefs study of 1,885 pages published in May 2026 found approximately 0% direct citation lift attributable to schema. 1 The result was independently corroborated by a searchVIU direct-fetch test and StanVentures analysis. 7 This is not a contested claim — it has been tested across multiple methodologies and the results are consistent.

This finding does not mean schema is unimportant. It means schema is infrastructure, not a lever. For product pages, schema.org Product/Offer markup — especially when paired with a complete Google Merchant Center feed containing GTIN, price, availability, and currency — is the binary gate for Shopping eligibility. 5 A product missing its GTIN does not appear in AI Shopping surfaces regardless of prose quality. A product with complete structured data but generic copy may appear and then fail to be cited. Both failure modes are real; they occur at different layers.

Schema markup is the admission ticket. It gets the product into the room. Prose quality determines whether the product is spoken about once it is there. Conflating the two is the most expensive content strategy error in AI commerce.

Hidden Layer research synthesis

The render gap compounds the stakes. LLM crawlers — GPTBot, ClaudeBot, PerplexityBot — are HTTP-only agents. They do not execute JavaScript. Server-rendered JSON-LD is the agent's read path; JSON-LD injected client-side via React or a tag manager is invisible to these crawlers. The same is true for prose: product descriptions loaded via API or lazy hydration after page paint are not read. Both the structured attributes and the conversational copy must be present in server-rendered HTML to be in scope at all.

What actually earns a citation

The most rigorous evidence for what moves AI citation rates comes from a 2024 Princeton RCT (Aggarwal et al., KDD 2024) that tested content modification tactics against live AI engines in controlled conditions. 2 The causal findings are clear: statistics embedded in prose, direct quotations from primary sources, and explicit citation of external sources each produced significant citation lift. Fluent prose rewrites added further lift. Keyword stuffing produced zero to negative results.

Citation Lift by Content Tactic (Princeton GEO RCT)
Position-adjusted citation word count lift vs control
Quotations added
+42.6%causal, RCT
Statistics with sources
+35%range 32.8–37%, causal
Cited external sources
+29.5%range 27.7–31.4%, causal
Fluent prose rewrite
+22.5%range 15–30%, causal
Keyword stuffing
0%zero or negative
Source: Aggarwal et al., KDD 2024 / Princeton GEO study. [2] Treat specific magnitudes as study-reported figures, not settled constants — direction and ranking of tactics are the durable takeaway.

The correlation data from Clairon / Wellows (April 2026) contextualizes the citation signal hierarchy. 4 E-E-A-T — expertise, experience, authoritativeness, trustworthiness as Google evaluates it — shows a correlation of r≈0.81 with AI citation probability. Domain authority, the most commonly cited SEO proxy metric, shows r≈0.18, explaining roughly 3% of citation variance. The practical implication: domain authority matters as a floor condition, not as the primary lever. E-E-A-T-signaling content — specific claims, named experts, primary sources, cited evidence — is what the citation layer responds to. This replaces any "DA vs schema" framing with a more accurate picture: E-E-A-T dominates, schema explains none of it, and DA is a modest contributor.

VERIFIED 3SUPPORTED 2

Schema markup ≈0% direct AI citation lift (eligibility-only function)

VERIFIED0.92authority T13 sources

Ahrefs May 2026 (1,885 pages) + searchVIU + StanVentures. Nuance: untested for cold-start pages with zero existing visibility.

Keyword stuffing produces zero or negative AI citation lift

VERIFIED0.95authority T12 sources

Princeton KDD 2024 RCT (causal) + ACM Web Conference 2026 (99.78% blocked). Double-refuted.

E-E-A-T r≈0.81 dominant citation predictor; DA r≈0.18

SUPPORTED0.78authority T21 source

Clairon / Wellows Apr 2026. Single study — correlational. Direction strongly supported by mechanism.

~83% of ChatGPT Shopping carousel sourced from Google Shopping organic top-40

SUPPORTED0.82authority T21 source

Search Engine Land / Peec AI, 43k products, Mar 2026. Single study but large sample.

GTIN is a binary gate for AI Shopping inclusion; missing GTIN = exclusion

VERIFIED0.95authority T12 sources

Google Merchant Center docs + AthosCommerce. "60% of catalogs missing GTIN" figure is single-vendor sourced — not independently verified; treat as directional.

One record, three renderings: the Golden Record model

The most common implementation error in AI-optimized product content is maintaining two separate records: one for "AI" or "feeds" (structured, clean, controlled vocabulary) and one for "SEO" or "humans" (keyword-rich, prose copy). This fork is the anti-pattern. It introduces inconsistency between the feed and on-page schema — a mismatch that Google explicitly cross-validates and treats as a disapproval signal. 8 It also creates a maintenance burden that compounds every time a price, availability status, or specification changes.

The correct model is a single Golden Record in a Product Information Management system, rendered three ways:

  1. Server-rendered JSON-LD (schema.org Product/Offer) — the machine read path for LLM crawlers and Google's Knowledge Graph. Must be in static HTML. Contains GTIN, price, availability, currency, brand, category. Zero marketing copy.
  2. Merchant feed push to Google Merchant Center (and, where applicable, the separate OpenAI Merchant Feed) — the Shopping eligibility gate. Must match on-page schema exactly; Google cross-validates both. 8 Triggers the 3-step chain: GMC feed → Google Shopping organic rank → ChatGPT Shopping carousel sourcing. 6
  3. Human-readable prose — the same factual attributes (specifications, use cases, compatibility, comparisons) expressed as natural language that answers the questions a buyer actually asks. This layer is what earns citations when the product is retrieved. It is also what humans read.

These three renderings share a single source of truth. The prose in rendering (c) should contain the same facts as the structured fields in (a) and (b), expressed in answer form. When a product description says "fits all standard 26mm bar clamps, rated to 120kg load" it is not diverging from the structured attributes — it is expressing the same data in a form that can be quoted by a generative AI engine answering "what is the best bar clamp for heavy-duty woodworking."

Golden Record → Three Renderings → Two AI Outcomes
PIMgolden recordfeedJSON-LDChatGPTPerplexityGoogle AICopilotAmazoninclusion decidedat ingestion —not query time
PIM → JSON-LD (machine read path, inclusion gate) + Merchant feeds (Shopping eligibility chain) + Prose (citation quality, human reading). All three must be consistent; the Merchant feed and on-page schema are cross-validated by Google.

Why keyword-optimized copy fails both humans and AI

Classic keyword-optimized product copy — repeating target terms, front-loading category keywords, stuffing synonyms — was already declining in effectiveness in traditional search as Google's ranking signals shifted toward entity understanding and content quality. In AI retrieval it is not declining: it has been tested and shown to be actively counterproductive.

The Princeton RCT found keyword stuffing produced zero or negative citation lift. 2 A separate ACM Web Conference 2026 paper found that 99.78% of keyword-stuffed content was blocked from AI search engine retrieval. 3 The mechanism is straightforward: LLMs are trained on embeddings, not term frequency. Repetitive term density is a quality signal in the wrong direction — it resembles spam training patterns, not authoritative content. The same prose characteristics that signal authority to a human reader (specificity, evidence, cited sources) are the signals that move AI citation probability.

What to do, in order

The two-layer model collapses into a concrete priority sequence. These are not parallel workstreams — they are dependencies.

  1. Audit structured completeness first. Every product needs a GTIN. The Offer block needs price, availability, and currency. The Merchant feed must match on-page schema exactly — Google cross-validates and a mismatch triggers disapproval. 58 Without this, prose quality is irrelevant: the product is not eligible for inclusion in Shopping surfaces.
  2. Verify server-render. Confirm JSON-LD is present in view-source, not injected by JavaScript after load. LLM crawlers do not execute JS. If your schema is client-rendered, it is invisible to GPTBot, ClaudeBot, and PerplexityBot — and so is any prose loaded dynamically.
  3. Rewrite product descriptions as answered questions. Replace keyword-dense copy with natural-language answers to the questions buyers ask: what problem does this solve, what is it compatible with, what differentiates it from the alternatives. Include verifiable specifications as prose claims. This doubles as the citation layer for AI and the reading experience for humans.
  4. Add statistics, direct quotes, and cited sources where applicable. These are the three tactics with the strongest causal citation lift from the Princeton RCT. 2 For product content, that means: cite the test standard that your load rating meets, quote the verification lab, include the specific dimension that confirms compatibility with a named system.
  5. Do not maintain two records. Any time a price, availability, or specification changes, it must propagate to all three renderings simultaneously. A feed that shows "in stock" and a page that shows "available for order" is a Google disapproval and an AI retrieval inconsistency. One source of truth, three renderings.
  6. Build domain authority through E-E-A-T signals, not schema tags. The dominant citation predictor (r≈0.81) is expertise-authoritativeness-trustworthiness as signaled in content — named authors, cited sources, institutional credibility, depth of coverage. 4 Domain authority (r≈0.18) matters as a floor; E-E-A-T is the ceiling variable.

Two renderings of the same truth

The schema-vs-prose framing is a false dichotomy in one sense and a useful distinction in another. It is false to treat them as alternatives — both are required, they serve different layers, and they must be kept consistent. It is useful to distinguish them operationally, because they fail in completely different ways: structured attributes fail through omission (missing GTIN) or inconsistency (feed-page mismatch); prose attributes fail through generic copy that cannot be cited.

The practical resolution is the Golden Record model: a single managed record rendered as machine-readable schema, a complete Merchant feed, and human-readable prose. What AI reads and what humans read are not separate concerns. They are the same underlying data, shaped for different readers. The prose layer serves both audiences simultaneously — and given that E-E-A-T dominates AI citation probability, the prose that earns a human reader's trust is precisely the prose that earns an AI citation.

The content teams closest to getting this right are the ones that stopped asking "what does the AI want to read" and started asking "what question does our buyer actually have, and what is the most credible answer we can give." That reframe — from keyword optimization to question-answer coverage — is the single change that aligns human readability, AI citability, and content maintainability in one move.

GEOSchemaProduct ContentGolden RecordStructured Data

Footnotes8

  1. Ahrefs (Williams-Cook, May 2026): 1,885-page study — schema markup shows ~0% direct AI citation lift; eligibility-only function confirmed
  2. Aggarwal et al. (KDD 2024, Princeton): GEO RCT — statistics +32.8–37%, quotations +42.6%, cited sources +27.7–31.4% citation lift; keyword stuffing 0% or negative
  3. ACM Web Conference 2026: keyword stuffing has "no influence on any LLM search engine"; 99.78% of keyword-stuffed content blocked from AI retrieval
  4. Clairon / Wellows (Apr 2026): E-E-A-T r≈0.81 dominant predictor of AI citation; domain authority r≈0.18 (~3% variance explained)
  5. Google Merchant Center documentation: GTIN required for Shopping eligibility; price/availability/currency required for Offer completeness; on-page schema cross-validated against feed
  6. Search Engine Land / Peec AI (Mar 2026): ~83% of ChatGPT Shopping carousel sourced from Google Shopping organic top-40; base-64 Google Shopping params found in ChatGPT carousel source
  7. searchVIU direct-fetch test + StanVentures: structured data markup confirmed as eligibility plumbing, not citation lever, corroborating Ahrefs finding
  8. Google Merchant Center documentation: mismatched on-page schema vs feed data triggers disapproval and quality downgrade signals
ShareLinkedInXEmail

Cite this article

Full
Harshak Patel. “Conversational vs Structured Attributes: What Humans Read vs What AI Reads.” Hidden Layer, 17 June 2026. https://hidden-layer-blogs.pages.dev/post/conversational-vs-structured-attributes
In line
Hidden Layer (2026)

Reuse

Republish, translate, excerpt or adapt the article text and the figures Hidden Layer drew under Creative Commons Attribution 4.0 International (CC BY 4.0), provided you credit Hidden Layer and link to the original. Read the CC BY 4.0 terms.

The grant covers our own words and charts only. It does not extend to data and figures quoted from other organisations, which stay with their owners; to trademarks and logos, ours and everyone else’s; or to audit reports and customer data produced by the product.

Author

HP
Harshak PatelFounder & Head of Research, Hidden Layer
Harshak Patel runs Hidden Layer, where the work is auditing how AI systems surface — or refuse to surface — brands and products. Background in enterprise product data and catalogue intelligence. The publishing rule here is simple: every article ships with its sources, its per-fact confidence, and the claims that were cut. The methodology is public and reproducible, and that, not the byline, is the credential.

Next

See how your domain scores against these checks.

Run a free audit

GEO Week — every Friday

Weekly brief on AI discoverability, agent readiness, and what shipped in the GEO space. No fluff.

We'll never spam you. Unsubscribe anytime. GDPR-compliant double opt-in.