A Third of the Web Is Now AI-Written. That's Not the Number You Should Worry About

Pew Research says over a third of web pages published since ChatGPT's launch show signs of AI authorship. The alarming part isn't the number — it's who gets crowded out along the way, and whose content trains the next generation of models.

Black technology professional working on a laptop while a digital display shows that one third of the web is AI-written, illustrating the growth of AI-generated content online.
AI-generated content is rapidly reshaping the web, but the bigger concern may not be how much content AI writes—it is how AI-written information affects search quality, trust, and the future of the internet.

The Number Everyone's Reacting to Is the Wrong One

Pew Research Center published a study this week with a startling headline figure: among web pages published since ChatGPT's November 2022 launch, over one-third (35%) now show significant signs of AI authorship. Researchers ran an open-weight detection model built by Pangram against roughly 490,000 English-language pages pulled from the Common Crawl archive, looking for the statistical fingerprints of machine-generated text doubled use of em dashes, a 63% jump in Oxford commas, more than double the usage of words like "delve," "pivotal," and "testament," and a near-tripling of "it's not X, it's Y" constructions.

It's a genuinely striking number, and it's already being treated as a verdict on the "death of the human internet." I think that's the wrong takeaway. The 35% figure describes volume. It says nothing about which corners of the web that volume is concentrated in, and that distribution is the part actually worth worrying about.

The Domain Breakdown Tells the Real Story

Buried in the same study is a far more useful number: AI-authorship rates vary enormously by domain type. On .com domains, roughly 10% of pages show AI-authorship signals. On .org domains, it's 4.6%. On .edu and .gov domains the parts of the web with institutional editorial standards, actual bylines, and reputational stakes it's about 1%.

Read plainly, that's not "the internet is being replaced by AI." It's "the parts of the internet built to be cheap and disposable content-farm SEO pages, affiliate listicles, low-stakes commercial copy are the parts absorbing most of the AI-generated volume." That's a less apocalyptic story than the 35% headline suggests, but it's also a more specific and more solvable problem, if publishers and platforms are honest about where it's actually concentrated.

Detection Is a Losing Game, and Everyone Building on It Knows It

Pangram's detector, like every AI-text classifier before it, works by spotting statistical regularities and statistical regularities are exactly the kind of thing that erodes the moment people start writing (or prompting) around them. Pew's own researchers flagged that these tools can misclassify content in both directions: human writers who happen to write in a clean, structured style get flagged, while AI text run through a light human edit slips past. The study's authors called the findings "likely at least directionally correct," which is an honest way of saying: treat this as a weather forecast, not a fact.

That matters because plenty of platforms — search engines, ad networks, academic institutions — are already building policy on top of detection scores like these. A methodology with real, acknowledged error margins is a shaky foundation for decisions with real consequences, like demonetizing a publisher or flagging a student's work.

The Quieter Risk: Whose Content Trains What Comes Next

Here's the part I think deserves more attention than it's getting. Large language models are trained, in part, on scraped web text and each new generation of models is increasingly training on a web that the previous generation helped write. Researchers have a name for what happens when a model trains repeatedly on its own outputs, or on data that's saturated with them: model collapse, a gradual narrowing of the outputs a system is even capable of producing.

That risk isn't evenly distributed. It concentrates wherever original, human-authored content was already scarce relative to demand which describes vast stretches of the web outside a handful of major languages and markets. If the input pool for a category or a language starts skewing synthetic, everything trained on it downstream inherits that drift.

What It Means for You

If you publish anything on the open web — a blog, product documentation, marketing pages — the practical lesson from the domain breakdown isn't "avoid AI tools." It's that editorial standards, real bylines, and institutional accountability are exactly what's currently correlating with lower AI-authorship rates, and that correlation is likely to become a trust signal readers and platforms increasingly reward, not just an artifact of who happened to adopt AI writing tools first. If you're building anything that consumes web-scraped data — a search product, a RAG pipeline, a fine-tuning dataset — this is also a good week to actually check your provenance and dedup practices rather than assume Common Crawl-style sources are as diverse as they were three years ago. And if your organization is using an AI-detection tool to make consequential decisions about people's content, Pew's own methodology notes are worth reading before you trust a single score.

Get the next one by email