ChatGPT’s Search Index: Unlicensed Sites Measure Up Against Paid Content Partners

6

ChatGPT's Search Index Includes Small Sites Without Licensing Deals, New Data Shows

New research reveals that OpenAI's in-house search index serves unlicensed publishers the same way it serves paid content partners — raising fresh questions about the real value of costly content deals.

French SEO consultancy Resoneo analysed 1,249 ChatGPT answers captured in July and found no meaningful difference in how the platform's proprietary index treated licensed versus unlicensed websites. The finding carries significant implications for publishers who have signed — or are considering signing — content agreements with OpenAI, and for smaller sites wondering whether they have any shot at visibility inside AI-generated answers.

For publishers and SEO professionals, this research reframes what licensing deals with OpenAI actually deliver — and what they don't.


What Resoneo Found Inside ChatGPT's Search Pipeline

ChatGPT's server stream tagged each web result with the name of the pipeline that fetched it. One of four possible values was "labrador" — OpenAI's own in-house search index. When Resoneo examined pages served through that pipeline, a licensing deal made no measurable difference. The format, length, and freshness of results were identical for both partners and non-partners.

Resoneo describes labrador as an index supplemented with press feeds and open science archives. Because OpenAI can access it directly without paying a third party, it operates independently from commercial scraping arrangements.

How Account Type Shapes Which Index You're Served From

For free-account users, the labrador index dominated. Questions with settled answers, local business queries, and product searches appeared through that index almost every time. News results were split more evenly between labrador and content scraped from Google.

The picture shifted significantly for paid accounts in thinking mode. Of the 16,407 search results Resoneo recorded from those sessions, approximately 75% came from scraped Google results while the in-house index accounted for roughly 24%.

This distinction is worth examining carefully. The pipeline serving your content depends not just on what you publish, but on which version of ChatGPT your reader is using — a variable entirely outside your control. For businesses and publishers trying to understand their AI visibility, this makes single-account testing an unreliable method. Understanding how to increase traffic to your website through search requires accounting for an increasingly fragmented landscape where AI interfaces draw from multiple distinct sources simultaneously.

How an Earlier Misreading Was Corrected

The Resoneo data also supports a public correction made by SEO consultant Suganthan Mohanadasan in July. In June, Mohanadasan had examined ChatGPT's network traffic and described the labrador index as what appeared to be a licensed tier. He pointed to prominent domains including Reuters, The Guardian, the Wall Street Journal, and Wikipedia as evidence.

On July 14, Mohanadasan reversed that conclusion. A reader in Italy, using a free account, sent him captures showing that every publisher citation passed through the same pipeline — including small Italian outlets with no OpenAI deal. After re-running his own tests, Mohanadasan acknowledged in his summary table that he had "over-reached" with the tier claim. He noted that licensing deals are genuine but that his earlier reading reflected a single account's behaviour rather than the full picture.

The two investigations differed meaningfully in scale. Resoneo's dataset spans free and paid accounts, multiple countries, and logged-out sessions, with identical prompts replayed across different account types. Mohanadasan's counts came from one account, which he describes as directional, though his correction also drew on captures from two additional readers. Resoneo credits Mohanadasan's work as the foundation for their own research.

Around July 21, OpenAI stopped tagging search results with the name of the fetching system — the precise tag both investigations had been reading to identify pipeline origins. The timing of that change, so close to when this research became widely discussed, has not gone unnoticed in the SEO community.


What ChatGPT Actually Stores About Your Page

Beyond pipeline behaviour, Resoneo examined what OpenAI's index actually holds about the pages it serves. The consultancy reviewed 534 pages that ChatGPT cited and compared each with the stored snippets.

The 200-Character Snippet Window

Of the 463 pages that included an H1 heading, 387 snippets contained it — a rate of 83.6%. The snippet cuts off just after 200 characters and typically draws from the beginning of page content rather than the meta description. The Google-scrape pipeline still captures the meta description roughly one in three times.

With a median H1 length of 51 characters, approximately 150 characters of page content remain in the snippet window. This is a narrow space, and several template elements can quietly consume it before your actual content appears:

  • A section kicker appears before the H1 on 29% of pages and uses an average of 18 characters
  • A publication date appears on 11% of pages and consumes around 25 characters
  • Alt text from the first image appears on 9% of pages and can take up to 50 characters

One in every seven pages in the sample had no H1 markup at all. In those cases, Resoneo notes that the snippet begins with whatever subheading the page template provides — which may bear little relation to the page's actual subject matter.

Resoneo did not test whether optimising these elements increases the likelihood of being cited, but the structural implications are clear: template bloat above the fold directly reduces the content ChatGPT stores about your page.

Why This Matters Beyond Traditional SEO

The snippet behaviour described here is meaningfully different from how Google constructs its search result previews. Google's system draws flexibly from meta descriptions, body content, and structured data depending on the query. ChatGPT's labrador index appears to store a fixed character window from a fixed position — making page architecture a more deterministic factor than it has historically been in traditional search optimisation.

For businesses investing in content as a growth channel, this has practical consequences. Those exploring strategies to grow organic traffic to their website should now consider how their page templates perform not just for search engine crawlers, but for AI indexing systems that read content in a fundamentally different way.


What This Means for Publishers and SEO Professionals

Sites without an OpenAI content deal are already present in the index that handles the majority of free-account ChatGPT results. Whether a licensing deal increases citation frequency is a separate question that neither investigation examined. Resoneo focused on how pages are stored and served rather than how often they are selected.

That distinction matters. Publishers sign content agreements with OpenAI for multiple reasons, and appearing in free-user results looks like a weaker justification given these findings. Resoneo points out that partner articles reach OpenAI through a feed rather than a crawl, meaning a deal may change how content arrives — even if it does not change how it is ultimately served.

The Broader Implications for AI-Era Content Strategy

The research arrives at a moment when the relationship between AI platforms and content publishers is actively being renegotiated across the industry. OpenAI's licensing arrangements have been positioned publicly as a mechanism for compensating publishers whose work trains and informs AI systems. If those deals do not materially change how content is surfaced to users, the commercial logic underpinning them becomes harder to defend — both for publishers evaluating whether to sign and for those already inside existing agreements.

For smaller publishers and independent sites, the findings carry a different kind of significance. If the labrador index already includes unlicensed content and serves it identically to licensed material, then the primary barrier to AI visibility may be technical and structural rather than commercial. That shifts the strategic focus from negotiation to optimisation.

The growth of AI tools in business has created pressure on publishers of all sizes to understand how these systems interact with their content — and this research provides some of the clearest structural evidence to date about how one major platform's index actually operates.

As of publication, OpenAI's crawler documentation does not detail the labrador index or specify what its publisher agreements include. For further context on how AI companies approach web indexing and content licensing, the Reuters Institute Digital News Report provides relevant industry-level analysis of the evolving relationship between news publishers and AI platforms.

How to Act on This Information

  • Audit your page template to identify elements that print above the first paragraph and consume space in the 200-character snippet window, since that is what ChatGPT stores about your page
  • Verify H1 presence and length across key pages, as 83.6% of cited pages had their H1 included in ChatGPT's stored snippet — making it a high-priority optimisation signal
  • Test ChatGPT visibility across multiple account types before drawing conclusions, since free and paid thinking-mode sessions draw from different pipelines with significantly different index distributions
  • Review image alt text and publication date placement in your page template, as these elements can silently consume a meaningful portion of the available snippet window before your content appears
You might also like