On August 20, Pew Research Center published the biggest public measurement yet of the machine-written web: 490,000 pages, five and a half years, one detection model run over all of it. A third of everything published since ChatGPT launched shows significant signs of AI authorship. Put that next to what actually ranks and gets cited, and the story stops being about volume. The volume is real. The visibility never showed up.
Half the new web is written by AI. It wins 14% of Google.
Noah Davis, Zoe Harris, and Lucas Brown · Aug 21, 2026 · 13 min read
- ai-content
- statistics

TL;DR
- Pew's August 2026 study of 490,000 pages finds 10% of all sampled web pages, and over a third of pages published since ChatGPT launched, show significant signs of AI authorship.
- Every marker Pew tracked grew between January 2023 and July 2026: em dash frequency almost doubled, AI-typical vocabulary more than doubled, and negative parallelism nearly tripled.
- Supply and visibility have split. AI now writes roughly half of new online articles, yet 86% of articles ranking in Google and 82% of ChatGPT and Perplexity citations are human-written.
- Only 7% of number-one Google results are AI-generated, and AI-written articles cluster after position ten, so raw AI publishing competes for slots that barely exist.
- The pipeline that works is AI-drafted and human-owned: a person edits out the tells, adds a number or a position no model can generate, and cites the primary source under every figure.
What Pew actually measured, and how
Pew's Data Labs team pulled 490,000 English-language web pages from the Common Crawl archive, spanning January 2021 to July 2026, and ran them through the open Pangram detection model. Two numbers lead the report. 10% of all sampled pages show significant signs of AI authorship, and among pages published after ChatGPT's November 2022 launch the figure passes one in three. The full report is public, and TechCrunch covered it the same day.
Pew is careful about the limits, and so are we. Detection models misclassify individual documents. A single em dash proves nothing about a single page. The claim holds at the level of half a million pages, where the statistical fingerprints of model prose separate cleanly from human writing. Read it as a directionally correct census, and treat any one page's score as a guess.
The tells doubled, and you already know every one of them
The part of the report nobody else led with is the trend data on specific writing patterns. Between January 2023 and July 2026, across the whole sample of web text, em dashes went from 5.79 to 11.19 per 10,000 words. Oxford commas rose 63%, from 34.04 to 55.51. AI-typical vocabulary, the delve-and-showcase word list, more than doubled from 11.94 to 26.02. And negative parallelism, the "this isn't X, it's Y" construction, nearly tripled from 0.87 to 2.36.
Treat that trend data as a public, quantified checklist of what makes bought content read machine-made. Your readers pattern-match on these without knowing their names. So do the reviewers at any publication you pitch, and, increasingly, the classifiers that decide what surfaces in AI search. A content operation that ships drafts carrying these fingerprints is publishing its own production process.

Show the data behind this graphHide the data behind this graph
| Marker | Jan 2023 (per 10,000 words) | Jul 2026 (per 10,000 words) | Growth |
|---|---|---|---|
| Negative parallelism | 0.87 | 2.36 | +171% |
| AI-typical vocabulary | 11.94 | 26.02 | +118% |
| Em dashes | 5.79 | 11.19 | +93% |
| Oxford commas | 34.04 | 55.51 | +63% |
Where the machine-written web lives
The AI share is not evenly spread. Commercial .com domains show significant AI-authorship signs at roughly 10%, about ten times the rate of .edu and .gov pages, which sit near 1%. Nonprofit .org domains land at 4.6%.
That distribution is a trust map. The domains with editorial gatekeeping barely moved, and the domains where publishing is cheap and unsupervised absorbed almost all of the flood. Every buyer of content now publishes into the crowded end of that map. The scarce signals, a verifiable number, a source someone can open, a position someone will defend, are worth more there than they were in 2022.
AI authorship by domain type, July 2026
Share of pages showing significant signs of AI authorship in Pew's July 2026 sample.
- of .com pages show significant AI-authorship signsPew Research Center (2026)
- 10%
- of .org pagesPew Research Center (2026)
- 4.6%
- of .edu pagesPew Research Center (2026)
- 1%
- of .gov pagesPew Research Center (2026)
- 1%
Supply hit half. Visibility stopped at 14%.
Pew counts pages. Graphite, an SEO research firm that runs its own detection studies, counts articles, and its numbers frame the same event from the supply side: AI-generated articles hit 35.9% of new articles within twelve months of ChatGPT's launch, crossed 50.9% in late 2025, and have sat flat near half for five straight quarters. The growth stopped. Whatever economics were pushing raw AI publishing found their ceiling more than a year ago.
Why the ceiling? Graphite's 2025 companion study of 31,493 keywords answers it with the demand side. In Google's results, 86% of ranking articles are human-written. In ChatGPT and Perplexity citations, 82% are. Only 7% of number-one Google results are AI-generated, and the AI-written articles that do appear cluster after position ten, where nobody clicks. The machines writing half the web are, for the most part, writing pages that no search engine and no answer engine ever shows anyone.
A 2026 SEO study from Rankability that scored 487 competitive commercial results found the same shape: 83% of top results read as human-written. Three separate measurements with different methodologies land on the same shape. The flood is real, and the flood is losing.
Why the gap exists
Google's published position is that it rewards quality regardless of how content is produced, and penalizes scaled content whose primary purpose is manipulating rankings. Believe the second half of that sentence more than the first. What the ranking systems reward has converged on exactly what raw model output lacks: information that exists nowhere else, backed by a claim with a source under it.
Raw AI drafts fail that test in a detectable way. Pew's tell data shows the fingerprints are getting more common, which means the average unedited draft is getting easier to classify. And the same statistical signals a research lab can count at scale, a ranking system can weight at scale. Nobody needs a courtroom-grade detector to downrank a page. They need a probability, and Pew just published how legible that probability is.
If you use this well, the flood is your pricing power
Half the web conceding the quality bar means the bar is cheaper to clear than it has ever been. A company that publishes twenty pages a year with real numbers, named sources and an argument someone owns is now competing against a field that is one-third statistical filler. The .edu and .gov numbers show what scarcity looks like: the trusted end of the web is barely 1% machine-written, and trust concentrates where the flood is not.
AI still belongs in that operation. The drafting cost collapse is real and you should take it. What the data rules out is the version where the model's output ships unread. Read Graphite's own caveat closely: it did not evaluate AI-assisted content with heavy human editing, and it names that unmeasured route as possibly an effective strategy. So nobody has measured edited drafts failing the way raw ones do, and the mechanism points the other way, because what a detector counts is the raw output's fingerprints and an editor's whole job is to replace them with judgment. Your team's judgment is the thing the measurement rewards.
This is the pipeline we sell, so weigh our bias accordingly. When we build content production systems for a studio or a marketing team, the model does the assembly and a person does the arguing: the workflow drafts, then routes every piece through an editor who owns the claims before anything ships. We wire that gate in as part of a custom AI agent build, because the gate is the product. The machine gets you to a reviewable draft in minutes. It never gets to publish.
If you use it badly, you join the indistinguishable third
Publish raw model output at volume and you are competing for the 14% of ranking slots that AI content actually wins, mostly below position ten, with pages that carry fingerprints growing more legible every quarter. Rankability's study documents the terminal version: a site running 100% AI-detected content that was deindexed after a Google update, and recovered only after a human rewrite.
The cost is not only rankings. A third of the post-ChatGPT web reading as machine-made has trained your buyers to bounce on the pattern. The em dash cadence and the paragraph that summarizes itself now work as a label, and the label says nobody here checked this. Once a prospect files your blog under that label, no individual page gets a second read.
And the near-term risk compounds, because AI search is repricing visibility right now. The same week Pew published, Axios reported that Reddit's share of ChatGPT Search citations collapsed by 86% in about a week as OpenAI reweighted its sources. Citation flows in answer engines move that fast, with no notice. Content that earns citations on trust signals survives reweighting. Content that got lifted by one engine's temporary source mix does not.
What to change in your content operation this quarter
Run the audit Pew just handed you. Count the tells in your last ten published pieces per 10,000 words and compare against the averages above; if your blog tracks the 2026 curve, your production process is visible to anyone who looks. Mapping that gap, across content and every other function, is half of what our 10X audit exists to do. Then fix the pipeline rather than the prose. Every AI-drafted piece gets a named human owner who edits the argument, not the commas. Give each figure a primary source a reader can open, and make each piece carry one thing a model cannot generate, whether that is your own data or a worked example from your own operation.
Measure the demand side while you are at it. Track which pages earn citations in ChatGPT, Perplexity and AI Overviews, and treat a citation as the new ranking. It is the same discipline as running evals on a production agent: measure the output, or you are guessing. The ROI numbers on AI investments follow measurement discipline everywhere else we have looked, and content is no different.
If you would rather buy the pipeline than build it, that is the trade an AI content creation agency exists for, and the same one our own AI automation agency sells, ours included: a fixed-price build that wires drafting, editing gates and source checks into one system your team runs. Whoever builds it, hold them to the same bar this post is held to. Ask where the human edit happens, and walk if the answer is nowhere.

Show the data behind this diagramHide the data behind this diagram
- A model drafts the piece, then one gate decides the outcome: does a person edit it, own the claim and remove the tells?
- No: the piece ships as part of the indistinguishable third of the post-ChatGPT web, competing for the 14% of Google rankings and 18% of AI citations that AI-detected content wins.
- Yes: the editor adds a number or a position a model cannot generate and cites the primary source under every figure, so the piece carries the signals of the side that holds 86% of rankings and 82% of citations.
The questions worth asking about the AI-written web
How much of the internet is written by AI in 2026?+
Pew Research Center's August 2026 study puts it at 10% of all sampled web pages and over one-third of pages published since ChatGPT launched in November 2022, based on 490,000 pages run through the Pangram detection model. Graphite's article-level studies find AI-generated articles plateaued near 50% of new online articles since early 2025. The two numbers differ because the units differ: Pew counts all pages across five years, Graphite counts newly published articles.
Does Google penalize AI-generated content?+
Not for being AI-generated. Google's published guidance targets scaled content abuse, low-quality content produced at volume to manipulate rankings, whatever wrote it. The measured outcome still runs against raw AI output: 86% of ranking articles and 93% of number-one results read as human-written in Graphite's 31,493-keyword study. Quality investment after generation, not the tool, is what separates the content that ranks.
How can you tell whether content was written by AI?+
At the scale of a single page, unreliably: every detector misclassifies individual documents, and Pew says so plainly. Across a body of text the fingerprints are countable. Em dash frequency, Oxford commas, AI-typical vocabulary like delve, and negative parallelism all rose between 63% and 171% across the web from January 2023 to July 2026. If your published content tracks those curves, readers and classifiers can see it.
Do ChatGPT and Perplexity cite AI-written content?+
Mostly not. Graphite measured 82% of the articles cited by both engines as human-written. Citation mixes also shift abruptly: Axios reported Reddit's share of ChatGPT Search citations fell 86% inside a month during August 2026 when OpenAI reweighted sources. Earning citations on verifiable data and clear sourcing is more durable than optimizing for any engine's current source list.
Should we stop using AI to produce content?+
No. What the studies measured losing is raw machine output: that is the category detection classifies as AI-generated, and it wins 14% of rankings. Heavily human-edited AI drafts sit outside both studies' measurements, and Graphite explicitly calls that unmeasured route a possibly effective strategy. The drafting cost drops either way. The difference is that a person edits the argument, adds original substance and sources the claims before anything ships. What the data argues against is publishing model output nobody read. That content now competes for the thin end of a measured 86/14 split.
read next
Keep going
Want the drafting speed without joining the third?
We build content pipelines where the model drafts and your team owns every claim that ships. Fixed price, running in weeks, measured against the numbers in this post.
A starter build runs $1,500 to $2,500 fixed. If raw volume is what you want, we are the wrong shop.

Written by
Noah Davis · AI Research Writer
I research emerging AI developments and write in-depth articles that give readers the context behind them.
Hiking & nature photography

Written by
Zoe Harris · Newsletter Writer
I write newsletters that keep readers current on AI news and tools, with practical advice they can use.
Painting & illustration

Written by
Lucas Brown · AI Explainer Writer
I turn technical AI topics into explainers that show readers how the pieces fit together.
Playing guitar




