Getting cited by AI is not about clean formatting: the 4 gatekeepers 252,000 trials revealed
What separates content cited by AI answer engines is not clean formatting but topical relevance, recency, list position, and evidence. A study published on May 25, 2026 compared 18 content factors across six models including GPT-5.2, Claude, and Gemini 2.5 over 252,000 trials, and found that a recent timestamp, topical relevance, and the top list position are the gatekeepers that decide citation, with odds ratios above 10,000. Formatting changes, by contrast, had a negligible effect. ASAP summarizes the result from the primary source.
Change one factor, 252,000 times: what they watched
The researchers took the same information, changed exactly one factor, and paired the two versions to see which one got cited. Run this way, 18 factors were tested across six models over 252,000 trials. The gaps between factors then came into sharp focus. Some factors governed citation across all six models, while factors like structure and information layout left almost no trace in any of them.
The four gatekeepers that drop citation odds to zero when failed
The first gate of citation is four factors, and failing any one drops the citation chance to nearly zero. The figures published on May 25, 2026 are as follows.
| Gatekeeper factor | Effect (odds ratio) | Meaning |
|---|---|---|
| Topical relevance | Above 10,000 | Exact match to the query topic |
| Recent timestamp | Above 10,000 | 2026 vs 2019 |
| Top list position | Above 10,000 | Position 1 vs 2 in candidates |
| Price information | Above 10,000 | Stated in commerce queries |
All four factors exceeded an odds ratio of 10,000. An odds ratio of 10,000 is not a bonus level. It is closer to a threshold that separates passing from failing.
After the gate, the contest moves to tone and evidence
Tone and evidence are the secondary factors that decide citation in GPT-5.2 and the other models once a page clears the gatekeepers. The secondary factors that were significant in four or more models are as follows.
| Secondary factor | Effect (odds ratio) |
|---|---|
| Confident vs hedged tone | 2.7–754 |
| Specifications included | 8.6–243 |
| Evidence-backed claims | Above 2.1 |
| Query terms included | 6–40 |
Confident, declarative sentences were cited up to 754 times more than hedged, vague ones. Claims backed by evidence and sentences carrying the query's key terms were consistently favored too.
Where the conventional wisdom flipped: formatting and star ratings had no power
Formatting and social proof are far weaker drivers than common belief assumes, significant in only two or three of the six models including Claude and Gemini 2.5. Format changes such as structure and information layout barely moved citation. Social proof like ratings and reviews and the strength of the value proposition were significant in only two or three of the six models, breaking the assumption of universal persuasive power. The intuition that "clean formatting earns citations" was not supported by the data.
Why this result changes the game
These numbers matter because they suggest much of the AEO effort so far may have been aimed at the wrong target. Polishing a table of contents and prettifying tables feels like progress, but on an odds-ratio scale it rounds to noise next to the four gatekeepers. Topical relevance and recency, by contrast, are low-visibility work that often gets pushed down the list. In other words, the study nails down that citation is not a contest of looking good but of whether you answered the query exactly, recently, and in first position. The 754-times figure for confident tone points the same way: a sentence that hedges and leaves room is essentially surrendering its slot in the citation.
Three things a practitioner should change today
Ported to real content work, the priorities become clear. First, state the year and date in the body so recency reads as a signal. Second, put the exact key terms a reader would query directly into your sentences. Third, strip out hedged, noncommittal phrasing and write declaratively. The rules ASAP applies to every post — recency anchored to 2026, declarative BLUF, numeric and proper-noun evidence, query-term mirroring — point the same way as these 252,000 trials. One caveat, though: price information is a gatekeeper specific to commerce queries and does not carry over to informational content. And note the limit — the study measured citation probability, not the accuracy or trustworthiness of what gets cited once it is.
Source: Rahul Vishwakarma et al., "What Gets Cited: Competitive GEO in AI Answer Engines" (arXiv 2605.25517, 2026-05-25; six models, 4,320 scenarios, 252,000 trials).

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr