About 10% of webpages show signs of AI authorship: what Pew found in 490,000 pages
Pew Research Center reported on August 20, 2026 that roughly 10% of a random sample of 10,000 webpages taken in July 2026 showed significant signs of AI authorship, and that the share rose above one-third once the sample was narrowed to pages published after ChatGPT launched in November 2022. The study pulled about 490,000 English-language webpage texts from the Common Crawl web archive spanning January 2021 through July 2026 and classified them with Open Pangram, a machine learning detection model. On .com domains the share climbed from 1.09% in January 2021 to 9.35% in January 2026, more than an eightfold increase in five years. ASAP summarizes the Pew Research Center Data Labs report as the primary source.
9.35% on .com against 0.76% on .gov, a twelvefold gap
The domain breakdown Pew published for January 2026 is the sharpest split in the study: commercial .com pages came in at 9.35%, .org at 4.59%, .edu at 1.03%, and .gov at 0.76%. The distance between the commercial and government domains is more than twelvefold, and roughly ninefold against the education domain.
That gap cannot be explained by detector performance, since the same model classified all four domain types the same way. What remains is publishing incentive. On .com, producing pages in volume, quickly and cheaply, converts directly into revenue; on .gov and .edu, page count is not a performance metric at all.
The curve bent in 2023
The six .com data points do not describe a steady climb so much as a curve with a hinge. The share sat at 1.09% in January 2021 and 1.08% in January 2022, essentially flat across two years, then moved to 1.63% in January 2023, 3.76% in January 2024, 5.48% in January 2025, and 9.35% in January 2026.
The event sitting between January 2022 and January 2023 is the November 2022 release of ChatGPT. The flatness of the two pre-ChatGPT years matters as much as the rise. It suggests the detector carries a floor of roughly one percent, misreading a small share of human writing as machine-authored, and it lets everything above that floor be read as signal stacked on top of a known baseline.
Em dashes and Oxford commas, the stylistic fingerprints detection picked up
Separately from the classifications, Pew tracked how stylistic markers in web text changed. Comparing January 2023 with January 2026, em dash frequency doubled, Oxford commas rose 63%, AI-typical vocabulary more than doubled, and negative parallelism, the "not X, but Y" construction, nearly tripled.
Pew did not use these markers as the basis for its classifications, and attached an explicit caveat instead: em dashes or Oxford commas on their own do not mean a piece of writing was produced with AI, because humans use them too. The stylistic figures are presented as a trend observed across a large sample, not as evidence about any document.
10% and one-third are not the same claim
The number most likely to be quoted from this study is 10%, but the number carrying the weight is one-third. The difference lies entirely in the denominator. The 10% figure divides by the full July 2026 random sample of 10,000 pages, and the web still holds an enormous stock of pages made when ChatGPT did not exist. As the report itself notes, many pages in these random samples could not have been written by AI in the first place.
That makes the summary "10% of the internet is AI-written" actively misleading about what is being published now. Restricting to pages published after ChatGPT, and finding more than a third, is the figure that points at current production. Mixing stock and new inflow always tilts the number toward optimism. The opposite caution applies too: Pew states that the detector misclassifies human documents as AI and the reverse, so using this result to judge any individual page is not what it was built for. The figure measures the slope of the web, not the guilt of a document.
A marker loses its power the moment it becomes known
Pew's finding that em dash frequency doubled between January 2023 and January 2026 is a measurement that contains the seed of its own reversal. Once a stylistic feature is widely known as an AI tell, systems generating text are tuned to strip it and human writers start avoiding it as well. The measured object reacts to being measured, and markers like these lose discriminating power over time.
This is already visible in practice. Editorial rules banning particular punctuation and particular connective phrases because they "read like AI" are taking hold across organizations, and AI-generated text that follows those rules moves closer to human writing on exactly these detection signals. Two pressures follow. Detection has to move from surface style toward deeper structure, and style norms get flattened in the name of avoiding an AI tell. That is the main reason to doubt that Pew's stylistic numbers would replicate the same way three years from now.
When the document an answer engine cites was itself written by AI
The problem this study poses for search and answer engines is not the ratio but the loop. With a third of newly published pages showing AI signatures, answer engines build responses from that same web and attach citations to it, which means the probability that a cited page was machine-written rises in step.
A distinction matters here. Being written with AI is not the opposite of being good. The problem is that unverified bulk output presents the same surface as verified work produced in small quantities. That is also what the gap between 0.76% on .gov and 9.35% on .com implies: the side producing primary material and the side repackaging it for distribution are separating at the domain level.
For anyone producing content, the practical guidance is not to hide AI use but to leave things in a document that repackaging cannot produce. Primary source verification, measurements taken directly, a visible checking process, and evidence that cuts against the argument are the first elements dropped by bulk production pipelines, which is precisely why they function as a dividing line.
The Korean-language web is not in this sample
Pew limited the analysis to about 490,000 English-language webpages, and Korean-language pages are not part of the sample. The AI authorship rate on the Korean web cannot be read off this study, and copying the English figures across has no basis.
There are grounds for guessing at direction, though. What drove the rise here was .com, and the Korean web carries a heavy share of commercially motivated content with a long-established practice of high-volume publishing aimed at search traffic. The detector side points the other way. Training data for models like Open Pangram skews heavily toward English, so applying the same method to Korean would require validating the reliability of the classifications first. Running this study on the Korean web starts with rebuilding the detector, not with transplanting the numbers.
What the study does not answer
Pew wrote three limitations into the report itself, and the possibility of Open Pangram misclassification is the first. Detection models are imperfect and sometimes label human-written documents as AI-influenced and the reverse. Markers such as em dashes and Oxford commas are not evidence of AI on their own. And the analysis covers English-language pages only.
One further question falls outside the design. Whether a page bearing AI signatures is a page people actually read, or one that ranks in search results, cannot be determined from a Common Crawl sample. The set of crawled pages and the set of pages humans reach are not the same, and bulk-produced pages in particular tend to exist without being reached. What accumulates on the web and what people read are separate questions, and this report answered the first one.
Source: Pew Research Center Data Labs, "How much of the internet is written with AI?" (August 20, 2026), summarized by ASAP.

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr