AI Authorship: Pew Research Study Finds 10% of Webpages Exhibit Signs of AI-Generated Content

2

1 In 10 Webpages Now Shows Signs Of AI Authorship, Pew Research Finds

A landmark Pew Research Center study reveals that roughly 10% of English-language webpages sampled in July 2026 carry signs of AI authorship or editing — a figure that jumps to 35% among pages published after ChatGPT's launch.

The findings add hard data to a debate that has quietly reshaped how the internet is written, read, and ranked. For SEO professionals, content marketers, and publishers, the implications are immediate: the commercial web is changing faster than most industry benchmarks have captured — and the tools people use every day may already be driving that shift.


What Pew's Data Actually Found

Pew Research Center's Data Labs team published their analysis on August 20, 2026. The team ran nearly half a million webpages through an AI detector and flagged approximately 10% as showing signs of AI authorship or AI-assisted editing.

That overall number tells only part of the story. When researchers narrowed their sample to pages carrying a publication date after ChatGPT's November 2022 launch, the detection rate rose sharply to 35%. Before ChatGPT's release, all four major domain types — .com, .org, .edu, and .gov — sat at or below 1%.

The divergence since then has been striking:

  • .com domains registered an AI-authorship detection rate of 9.35% over a six-month average
  • .org domains came in at 4.59%
  • .edu domains held close to pre-ChatGPT baselines at 1.03%
  • .gov domains registered just 0.76%

Commercial domains have climbed fastest while academic and government sites have barely moved — a gap that carries significant implications for anyone working in digital publishing or SEO.

Pew was careful to note a key limitation. Their detection threshold captures AI-assisted editing alongside fully AI-generated content. A page that a human writer drafted and then polished using a built-in AI tool lands in the same flagged category as one produced entirely by an AI model from the first word. Understanding what artificial intelligence actually is and how it works provides useful context for interpreting what these detection tools are actually measuring — and where their boundaries lie.

A Note on Detection Reliability

It is worth emphasising that AI detection tools are probabilistic, not definitive. Pew's methodology acknowledges that flagged pages may include content where a human writer used light AI assistance — a grammar suggestion, a rephrased sentence — rather than wholesale generation. This distinction is technically invisible to current detectors, which means the 10% figure almost certainly captures a spectrum of human-AI collaboration rather than a clean count of fully machine-written pages.


The Linguistic Tells Spreading Across the Web

Beyond raw detection scores, Pew examined specific writing markers that have grown more common in pages published after ChatGPT's release. The findings reveal a measurable drift in how online text reads.

Patterns That Have Shifted Significantly

Em dash usage nearly doubled — climbing from 5.79 uses per 10,000 words in early 2023 to 11.19 by early 2026. Oxford comma frequency rose 63% over the same period. Words commonly associated with AI writing outputs — including "delve," "interplay," and "testament" — more than doubled in frequency across the sampled pages.

The negative parallelism construction — the rhetorical structure often written as "it's not X, it's Y" — increased from 0.87 to 2.36 uses per 10,000 words. While still relatively rare, its growth follows the same upward trajectory as the other markers.

Pew was explicit that none of these traits can definitively identify a single document as AI-generated. Human writers use all of them. The claim is about statistical rates across large bodies of text — a distinction that matters enormously when these scores are applied to individual pages in isolation.

Why These Markers Matter to Content Teams

For publishers and editorial teams, these linguistic shifts offer a practical internal audit framework. If your content output is trending toward higher em dash density, an uptick in words like "delve" or "testament," and more frequent use of negative parallelism constructions, that may signal heavier AI involvement than your editorial standards intend — regardless of whether a detector flags it. The role of a skilled professional digital content writer becomes more important, not less, as these markers become easier for readers and algorithms to recognise.


How Other Estimates Compare — and Why the Numbers Differ

Pew's 10% overall figure is not the only estimate in circulation, and it is not the highest. Two other major analyses published within the past year land at significantly different numbers.

Graphite, an SEO firm, estimated that 49.9% of newly published English-language articles were primarily AI-generated in the first quarter of 2026. A preprint from researchers at Imperial College London, the Internet Archive, and Stanford University found that by mid-2025, approximately 35% of newly published websites were detected as AI-generated or AI-assisted. That study has not yet been peer-reviewed.

Why the Gap Exists

The difference between these figures and Pew's is largely explained by methodology and sample composition:

  • Pew scored every English-language page in its collection, including forums, homepages, and category pages — content types where AI assistance is far less common
  • Graphite restricted its sample to pages with article schema markup of at least 100 words that its classifier identified specifically as articles or listicles
  • As Pew's own data suggests, articles are where AI authorship concentrates — forums and category pages are not

All three estimates relied on Pangram as a detection tool in some form. Graphite also incorporated Copyleaks and GPTZero. The choice of detector and the composition of the sample drive most of the variation in reported figures.

This methodological variance is not a flaw unique to AI content research. Any large-scale web analysis carries sampling decisions that shape outcomes significantly. Readers interpreting any headline figure — 10%, 35%, or 49.9% — should treat sample scope as the first question to ask. You can review Pew Research Center's Data Labs directly to examine their methodology in full.


What This Means for the Web Going Forward

The concentration of AI-generated or AI-assisted text on commercial domains is the most consequential finding for anyone working in digital marketing or SEO. The .com detection rate has risen in every reading since ChatGPT launched — precisely where the majority of SEO work is targeted and where competitive content pressure is highest.

The Invisible Integration Problem

AI editing is now a native feature embedded in Google Docs and Microsoft Word, making passive AI assistance nearly invisible to both writers and readers. None of the three studies examined here distinguish between text entirely produced by AI and text written by a human and refined by an AI tool. That distinction may matter less over time as the boundary between the two continues to dissolve.

This raises a broader strategic question for businesses investing in content at scale. The risks and challenges of artificial intelligence in business extend beyond automation and job displacement — they include reputational and quality risks that emerge when editorial oversight is reduced in favour of volume. Pew's threshold does not distinguish full AI generation from AI-assisted editing, meaning human review remains the clearest way to separate volume from value.

Whether a page is accurate, useful, and worth a reader's time is still determined by reading it — a point none of the studies dispute.

How to Apply These Findings Practically

  • Publishers and content teams can audit their output against the linguistic markers Pew identified — rising em dash frequency, "delve"-style vocabulary, and negative parallelism — as a practical internal quality signal before publication.
  • SEO professionals should factor sample methodology into how they interpret AI-content estimates, since a figure built on article pages will look markedly different from one built on the full crawlable web.
  • Businesses investing in content at scale should note that editorial standards and consistent human review remain the most reliable way to maintain quality as AI assistance becomes embedded across standard writing tools.
You might also like