LLM SEO: 9 Changes That Actually Move Citations

02. September 2026

LLM SEO is the practice of changing pages so large language models cite them, and only three of the nine changes usually recommended survive controlled testing.

The rest are either conditional, correlational, or measured at zero. This article ranks all nine by the strength of the evidence behind them, names the study and sample size for each, and says plainly where the studies disagree. Two of the results contradict advice you will find on almost every checklist.

1. Rank well in classic search

This is the strongest and least fashionable finding in the field.

AirOps analysed 16,851 queries, 50,553 responses and 353,799 pages in The Fan-Out Effect and measured citation rate by retrieval position.

Retrieval position Citation rate
1 58.4%
2 54.4%
3 35.5%
4 29.9%
6 24.6%
10 14.2%

The drop from position two to position three is the cliff. Sprinklr's controlled study, 252,000 pairwise trials across six LLM systems, found odds ratios above 2,000 for position one against position two, which is as close to a decisive factor as anything in this literature gets.

There is no LLM SEO tactic that outruns being retrieved first. Everything below this line operates on pages that already made the candidate set, which is also the working definition of AI content optimization: fixing pages that are close rather than pages that are absent.

2. Match the title and headings to the question asked

Ahrefs analysed 1.4 million ChatGPT prompts and found mean cosine similarity between prompt and title of 0.602 for cited pages against 0.484 for pages that were not cited. For cited pages, maximum similarity between a fan-out query and the title reached 0.656.

AirOps measured the same relationship on headings and found citation rates rising from 30.2% below 0.50 similarity to 41.0% at 0.90 or above. Among already top ranked pages the effect was larger still: 55.9% below 0.60 similarity, 75.3% at 0.90 or above.

Both are observational, so neither proves that editing a title alone causes the lift. But two independent large samples pointing the same way, plus the fact that the mechanism is obvious, makes this the cheapest high confidence change on the list.

The fan-out part matters more than the headline number. Questions generated by the model on your behalf never appear in a keyword tool, so a title tuned only to your seed keyword misses most of what actually retrieves you. Query fan-out analysis for AI Mode is how those sub questions become visible, and prompt research for AI search is how you decide which of them are worth a page.

3. Put the answer in the first third of the page

Kevin Indig's Growth Memo analysis of roughly 1.2 million AI answers and 18,012 verified citations found that 44.2% of citations came from the first 30% of the source text.

That is a distributional finding rather than a controlled experiment, and nobody has run a clean "move the answer upward and measure" test. AgentGEO identifies buried or truncated content as a distinct failure mode and includes content relocation among its repairs, but its headline result bundles several fixes together.

Treat this as strong circumstantial evidence with an obvious mechanism: a passage selector reading the top of your page cannot select what is at the bottom of it.

4. Be specific, with numbers, prices and specifications

Sprinklr's controlled trials found "specifications versus no specifications" odds ratios between 8.63 and 243 across six models, and an explicit price was one of only four factors significant in every system tested.

A separate 2026 measurement by Zhang, He and Yao, covering 602 prompts, 21,143 search layer citations and 18,151 fetched pages, found that pages containing numbers or statistics had mean influence 0.1171 against 0.0725 for pages without, a relative difference of 61.55%. Direct quotations scored 0.1633.

The caveat is important, and it is where most advice goes wrong. Quantitative evidence is a reliable positive correlate. Bolting statistics onto a page is not a reliable intervention. See section nine.

5. Keep it current, but new is not automatically better

Freshness looks decisive in controlled testing. Sprinklr's recent versus old timestamp comparison produced odds ratios of 14.4 to above 10,000 across six models.

The observational picture is stranger and more useful.

Page age Citation rate
Under 30 days 25.3%
30 to 89 days 32.8%
90 to 179 days 32.4%
1 to 2 years 32.0%
2 to 5 years 27.5%
Over 5 years 27.6%

The relationship is not monotonic. Pages under 30 days old were cited less often than pages between one and three months old, and a two year old page did about as well as a three month old one. AirOps could detect a publication date for 41% of pages, so read this as a large partial sample.

The practical reading: publish, then let a page settle. And when the freshness effect did appear, it was weak on pages whose query match was poor, which puts section two ahead of this one in priority.

6. Name entities explicitly and consistently

Sprinklr found "query terms present versus missing" odds ratios of 5.99 to 40.0, and "consistent versus contradictory" content of 1.74 to 4.09.

The supporting observational work is striking but weaker. Sarkhedi's 2026 study of 187 standardised queries across 83 professionals and three platforms found that people with 50 or more cross source co-occurring domains appeared in 61% of answers against 14% for those below 15 domains, and that having all six entity signal categories corresponded to 73% appearance against 11% for two or fewer. The author states plainly that the design is correlational.

What survives from all of this is a writing rule rather than a growth hack. Write "Finseo launched post purchase attribution in 2026", not "the company launched it". A selected passage may begin mid page with no subject in view, and a pronoun there refers to nothing.

7. Organise the whole document around the topic

GEO-SFE, from Yu and colleagues at the University of Tokyo, tested structure across 200 articles, 377 queries and six generative engines, 2,400 test cases in total. Citation rate rose from 45.0% to 52.8%, a 17.3% relative improvement, p below 0.001, Cohen's d of 0.64.

The split inside that result is the actionable part. Macro structure, meaning how the document as a whole is organised, contributed 44.9% of the gain. Micro structure contributed 15.4%.

Reorganising a page beats reformatting a paragraph by roughly three to one. That also explains why the formatting advice in section nine performs so poorly.

8. Use natural language URL slugs

Ahrefs found that natural language slugs were cited 89.78% of the time against 81.11% for slugs that were not.

Small, cheap, and one of the few findings here with no serious counter evidence. It costs nothing on a new page and is rarely worth a redirect on an old one.

9. Ship schema, but stop expecting citations from it

This is the finding that contradicts nearly every checklist, and it now rests on two independent studies.

Ahrefs tracked 1,885 pages that added schema against 4,000 matched controls, comparing 30 days before and after across Google AI Overviews, AI Mode and ChatGPT. The difference in differences estimates were −4.6% for AI Overviews, +2.4% for AI Mode and +2.2% for ChatGPT. Only the AI Overviews decline was statistically distinguishable from zero, and Ahrefs cautioned against attributing even that to schema.

Kurt Fischman's SSRN preprint analysed 730 AI citations from 75 commercial queries across 1,006 unique pages and found nulls across the board: schema presence odds ratio 0.678 (p = .296), entity richness 1.001 (p = .833), schema to query alignment 1.068 (p = .626).

There is one exception worth keeping. Attribute rich Product and Review schema was associated with 61.7% citation against 41.6% (p = .012) in Fischman's data, which is a subgroup association rather than a randomised addition, and plausibly reflects that pages with real product attributes have real product content.

Meanwhile the raw association still looks positive: AirOps observed 38.5% citation for pages with JSON-LD against 32.0% without. Both can be true. Sites that ship clean structured data tend to be sites that also ship clean information architecture, fast pages and accurate entities. Schema is a symptom of that discipline more than a cause of citations. Ship it because machines have to parse your page, and stop expecting it to move this particular needle.

What is measured not to work

Technique Evidence
Keyword stuffing Princeton GEO: at or below baseline, roughly 10% worse in the Perplexity test
Adding jargon or unusual vocabulary Princeton GEO: no reliable improvement
A more authoritative tone Princeton GEO: no significant improvement
Generic JSON-LD as a citation lever Ahrefs matched study and Fischman: null

Why "add statistics" is not the same as "have evidence"

Section four said quantitative evidence correlates strongly with citation. Section nine warned that adding it may not work. Both are correct, and the gap between them is the most useful thing on this page.

Princeton's original GEO study measured Statistics Addition at roughly 30 to 40% aggregate improvement. Broken out by the page's original Google rank, the same technique produced −20.6% at rank 1 and +97.9% at rank 5. Quotation Addition behaved identically: −22.9% at rank one, +99.7% at rank five.

FeatGEO then tested the same interventions across three simulated engines and found Statistics Addition below baseline in all three: 12.21% against a 13.34% baseline on GPT-4o-mini, 5.54% against 8.89% on Gemini 2.5 Flash, and 3.72% against 5.20% on Qwen-plus.

So: a page that genuinely contains evidence is more likely to be cited. A page that has evidence sprinkled onto it late is not reliably improved, and if it was already the cited source, the edit can cost it share. Write the evidence in because the piece needs it, not because a study reported 40%.

What this looks like across 8.9 million answers

Two things from our own tracking put the whole list in context.

First, where you spend effort matters as much as what you change. Across 731,000 answers in a 14 day window, a Perplexity answer carried 13.69 citations on average against 4.76 for ChatGPT. Same page, same work, roughly three times the number of slots, which is why Perplexity and ChatGPT are worth planning as separate surfaces.

Second, the field is wider than a search results page. Across 8.85 million cited sources, the ten most cited domains hold 6.2% of all citations, and it takes 5,000 domains to reach 57.9%. On a classic results page ten results hold all of page one. That long tail is the reason a mid authority site can compete here at all, and seeing who currently holds the slots you want is what competitor analysis in AI search is for.

Knowing which of your own pages are retrieved but never cited is the part no study can give you. That is an observation problem: run a fixed prompt set, record the cited sources per answer, and compare. AI visibility tracking across nine answer engines exists for exactly this, and the pages it flags as retrieved but uncited are usually the cheapest wins on the list.

How to test any of this yourself

  1. A fixed prompt set. Twenty to fifty questions per topic, unchanged between runs.
  2. A baseline period. Two to four weeks before you edit anything, so you know your normal variance.
  3. One change at a time, on some pages and not others. A holdout group is what separates a result from a coincidence.

Every number above comes from someone else's sample. The only sample that decides your roadmap is yours. We ran this exercise on the published literature in nine GEO techniques tested, and the same caveats apply to our numbers as to everyone else's.

Method note

Our own figures come from Finseo tracking data, aggregated across accounts, with no customer, project or private domain identifiable in any number. The engine comparison covers a 14 day window and states the answer count per engine. The concentration figures cover the full source corpus. Prompt sets are chosen by the customers who run them, so the corpus skews toward the categories those customers sell in.

Every external figure is attributed to its study with sample size and year. Several of the 2026 sources are preprints or vendor analyses rather than peer reviewed work, and some use simulated retrieval rather than production endpoints. Where studies disagree, notably on schema and on statistics, both results are shown rather than the convenient one.

FAQ

What is LLM SEO? LLM SEO is optimising content so large language models retrieve and cite it. It overlaps heavily with classic SEO, because retrieval in most systems runs on ordinary search ranking.

What is the difference between LLM SEO and GEO? Mostly vocabulary. Generative engine optimization, answer engine optimization and LLM SEO describe the same work. The mechanisms and the studies are the same.

Is SEO dead now that AI answers exist? No. Retrieval position is the single strongest measured factor: pages at retrieval position one were cited 58.4% of the time against 14.2% at position ten. Classic ranking is an input to the new surface.

Does schema markup help me get cited? On the evidence available, no. A matched study of 1,885 pages adding schema found effects between −4.6% and +2.4% depending on the surface, and a separate analysis found nulls across schema presence, entity richness and alignment. Ship schema for parsing, not for citations.

How fresh does a page need to be? Fresher helps, but not linearly. Pages under 30 days old were cited less often than pages one to three months old, and pages one to two years old performed nearly as well. Update meaningfully rather than churning dates.

Which pages should I fix first? The ones already retrieved but never cited. They have cleared the hardest stage and failed a cheaper one.

Sources

Ours

  • Finseo tracking data: citations per answer by engine, 731,000 answers in a 14 day window
  • Finseo tracking data: citation concentration by domain across 8.85 million cited sources

External

  • AirOps, The Fan-Out Effect, 2026: 16,851 queries, 50,553 responses, 353,799 pages
  • Ahrefs, Why ChatGPT Cites One Page Over Another, 2026: 1.4 million prompts
  • Ahrefs, We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved, 2026
  • Vishwakarma, Kumar and Jamidar, What Gets Cited: Competitive GEO in AI Answer Engines, Sprinklr, ACM SIGIR 2026: 252,000 trials across six systems
  • Kevin Indig, The science of how AI pays attention, Growth Memo 2026: 1.2 million answers, 18,012 verified citations
  • Yu et al., Structural Feature Engineering for Generative Engine Optimization (GEO-SFE), arXiv 2026
  • Aggarwal et al., GEO: Generative Engine Optimization, ACM KDD 2024: https://arxiv.org/abs/2311.09735
  • Liu and Xu, FeatGEO, Nanjing University of Information Science and Technology, 2026
  • Kurt Fischman, Does Schema Markup Predict AI Citation?, SSRN preprint 2026
  • Zhang, He and Yao, geo-citation-lab, 2026: 602 prompts, 21,143 citations, 18,151 pages
  • Sarkhedi, How Large Language Models Surface Personal Brands, 2026: 187 queries, 83 professionals

Disclosure: the tracking figures in this article come from Finseo, an AI visibility platform. External figures are attributed to their original source.