AI Citation Research: Source Order’s Impact Lessens, Highlighting Need for Content Quality
AI Citation Test Finds Source Order Matters Less Than Raw Data Suggests
A controlled study published September 14 on arXiv found that swapping source positions in AI search had far smaller citation effects than raw position data implied, while structured content rewrites shifted how citation credit was distributed.
The preprint by researchers Sriram Selvam and Anneswa Ghosh challenges a widely held assumption in AI search optimization: that where your content appears in a search result largely determines whether it gets cited. The findings carry significant implications for marketers, SEO professionals, and content strategists who are building strategies around AI citation visibility.
What the Researchers Actually Tested
The study prompted a GPT-5.4 search agent to answer 130 common questions through independent web searches. Researchers recorded every message and search result from the 129 questions the agent addressed. From those transcripts, they identified pairs of pages that appeared in the same search results and both supported the same fact — creating conditions where either page could fairly receive citation credit.
That process yielded 113 valid pairs. A blinded human review later confirmed 103 of them as genuine matches.
Each saved conversation was then replayed four ways. Researchers placed one page above or below the other and presented its text either as plain paragraphs or rewritten with headings, lists, or a table. Only the final answer was regenerated. Nearly all rewrites were produced by Grok 4.3, with GPT-5.4 used as a fallback for one pair.
The researchers noted an important limitation: because the wording differed between the two text versions, the test compared two distinct rewrites rather than isolating formatting changes alone. This distinction matters significantly when interpreting the results — a point that is easy to overlook in headline summaries of the research.
To understand why this type of controlled testing is both necessary and difficult, it helps to consider what artificial intelligence systems actually are and how they process information — including the layered mechanisms that determine which sources an AI agent retrieves, ranks, and ultimately references in a generated answer.
The Gap Between Raw Numbers and Controlled Results
In the raw data, pages appearing in the first position of a search call were cited 85.1% of the time, compared to 42.8% for pages in fifth position — a difference of 42.3 percentage points. That figure looks dramatic until the methodology is examined more closely.
The study emphasises that search providers typically place more relevant pages at the top, meaning the raw gap reflects both position and page quality simultaneously. When researchers actually swapped the same page to a higher position within its pair, the probability of it being cited increased by just 7.9 percentage points — a result that was not considered statistically significant after accounting for multiple tests.
A separate testing set of 56 pairs, where only order was switched, produced an estimated effect of exactly 0.0 percentage points, with a 95% confidence interval ranging from -5.4 to +5.4.
"This is an attribution-sensitivity warning, not an optimization tactic," the authors wrote in the paper's discussion section.
Why This Gap Matters for Practitioners
The distance between a 42-point raw gap and a statistically negligible controlled effect illustrates a problem that extends well beyond this single study. Correlation-based reporting — common in vendor white papers and industry analyses — can make position look like a lever when it may simply be a proxy for underlying content quality. Acting on that misreading could redirect content investment toward optimising placement rather than improving substance.
An overview of how artificial intelligence is being applied across business functions provides useful context here — particularly in understanding how AI-powered search and retrieval tools are being embedded into commercial workflows where citation accuracy carries real consequences.
Structured Rewrites, Model Randomness, and What the Data Can Actually Support
Structured Rewrites Shifted Credit Without Clearly Increasing Citations
Pages rewritten with headings and lists received an average of 0.50 more citation markers per answer compared to plain paragraph versions. The 95% confidence interval ranged from 0.20 to 0.84 additional markers. Given that the median answer in the test contained 29 citation markers across six documents, that shift represents a meaningful redistribution.
Critically, the total number of citations per answer did not rise. Citations on the competing page barely changed either. The authors interpreted this as structured formatting concentrating credit onto the reformatted page rather than generating new citation opportunities overall.
The primary pre-planned test — whether structured text increased the likelihood of a page being cited at all — showed a 4.5 percentage point increase. However, the 95% confidence interval ran from -1.4 to +10.4, and the authors acknowledged the study could only reliably detect effects of approximately 8.5 percentage points or more, making this result inconclusive.
A stricter formatting comparison, adjusting layout to one sentence per list row while keeping every word identical, boosted citation rates across all 113 pairs. When researchers repeated that test with a subset, the effect reversed — underscoring how unstable single-run conclusions can be.
Model Randomness Undermines Single-Run Conclusions
When 120 responses were retested using identical inputs, the decision to cite or not changed in 15% of cases — roughly one in seven. The researchers estimate that approximately 45% of the variation observed in a single test run stems from model randomness rather than any variable being tested.
This finding aligns with data reported by SparkToro in January, which found that ChatGPT and Google's AI Overviews each produced the same brand list less than 1% of the time when given identical prompts repeatedly.
The authors recommend that citation tests be run multiple times and that researchers report consistency across those runs rather than drawing conclusions from a single output. This recommendation has practical weight: a one-time test that appears to confirm a formatting win may simply have captured a favourable random output.
An examination of the risks and challenges AI presents for businesses reinforces why this kind of measurement instability deserves serious attention — particularly for organisations making content or budgetary decisions based on AI-generated outputs they may only be sampling once.
Limitations the Study Acknowledges Directly
An Ahrefs report from May found that pages cited by AI were approximately three times more likely to include JSON-LD schema. However, as this study illustrates, correlation in vendor reports does not confirm that changing the variable drives the outcome.
The study cannot confirm whether reformatting a live page would boost citations on the open web, since rewrites only applied to text already retrieved during the test — excluding crawling, retrieval, and ranking processes that operate before any formatting is ever evaluated.
Researchers re-ran the saved searches using Grok 4.3 and found that structured rewrites pointed in the same direction. However, fewer than half of Grok's first replies followed the correct citation format, limiting direct comparison.
The authors call for additional research testing each scenario multiple times, examining different search providers and models, and measuring both citation frequency and whether a page is cited at all. Independent replication across models and retrieval environments will be necessary before any of these findings can be treated as reliable guides to content strategy.
For further reading on the methodology behind controlled AI evaluation studies, the arXiv preprint repository hosts the full paper alongside related work in information retrieval and language model behaviour.
Three Ways This Research Can Be Applied
- Before attributing citation gains or losses to a single change, rerun your tests multiple times and compare consistency across outputs rather than relying on one result.
- Treat structured formatting as a potential credit-concentration tool rather than a guaranteed citation driver, and measure citation marker counts alongside simple cited-or-not tracking.
- When evaluating vendor reports or correlation-based findings about AI citations, ask whether the proposed variable was ever tested in isolation using a controlled swap — if not, the insight may not be actionable.