AI
Pew: Over 35% of Recent Webpages Carry an AI Fingerprint
Pew Research scanned 500,000 pages: over 35% of English webpages published since ChatGPT's launch show signs of AI authorship. What the number includes, what it hides, and what it means for publishers.

500,000 pages scanned, one uncomfortable ratio. Pew Research Center has produced the first large-scale measurement of how much of the web is now machine-written, and the headline number lands hard: more than 35 percent of English-language webpages published since ChatGPT's launch in November 2022 show signs of AI authorship.
How the measurement worked
Pew pulled pages from the Common Crawl archive covering 2020 through 2026, scanned roughly 500,000 English-language pages, then ran a random sample of 10,000 pages collected in July 2026 through Pangram, an AI-text detection tool. One definitional detail matters more than any other: "signs of AI authorship" covers pages that were written or substantially edited by AI. The 35 percent is a blend of fully generated text and human drafts that a model reworked, so the share of pages written entirely by machines is lower, and unknown.
Detection itself is on firmer ground than it used to be. A separate study published the same week found that post-training pushes language models into a narrow, converging style, a kind of mode collapse, which keeps their output statistically identifiable even as the models improve. Detectors still make mistakes on individual pages; at the scale of thousands, the signal is real.
The feedback loop nobody ordered
Two audiences consume the web: people and models. Readers are getting visibly tired, and this week's community mood made that plain when a manifesto titled "Don't paste the AI, please" collected over a thousand points on Hacker News, aimed at raw model output being dumped into forums, emails, and pull requests. The second audience is the more troubling one. Search engines' AI summaries and chatbots draw from these same pages, which means models increasingly read, summarize, and cite each other's output. The measurement also stops at English; no comparable figure exists for other languages, though there is little reason to expect them to look better.
Where this leaves anyone publishing content
Our reading: the price of generic text has effectively hit zero, because anyone can generate it in seconds, and a third of the web already has. What still commands attention is content carrying something a model cannot invent, your own data, your operating experience, your test results. Using AI to draft is a workflow choice; publishing unedited output into a sea of identical output is a strategy for invisibility. In a market where 35 percent of pages share the same statistical fingerprint, being detectably human is turning into a measurable advantage.
Sources: Pew Research Center, TechCrunch, The Decoder

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.