01 Triggering Event
In June 2025, Cloudflare disclosed that AI bots now account for over half of all global internet traffic. The same month, the company announced layoffs of 20%, roughly 1,000+ employees, and CEO Matthew Prince published a WSJ op-ed on how to replace those roles with AI. Most recently, Prince appeared on The Verge's Decoder podcast and laid out Cloudflare's new positioning: a paid-passage, per-request billing, domain allowlist control panel between site owners and AI crawlers.
How I choose which Cloudflare employees to replace with AI
(This is the title of Prince's WSJ op-ed; I'm using it as a pull-quote callout.)
Worth hedging: this 51% figure is Cloudflare's own disclosure, which I haven't cross-checked against third-party traffic monitoring (Similarweb, Cisco Internet Report). But Cloudflare's vantage point covers a substantial share of global CDN requests, and the credibility of its self-published data is higher than the typical SaaS company — the order of magnitude should hold.
02 What This Really Means
On the surface, this is another Cloudflare product announcement. In reality, Prince is betting on a structural thesis: over the next five years, as AI companies continuously scrape training data from the public web and deploy agents to operate web forms and pages, this channel will require someone to settle and clear transactions. Cloudflare has positioned itself on this channel and believes it doesn't need to build anything new.
This is the second act of Stratechery's aggregation theory: Google in the 2000s aggregated all web links, became the hub of traffic distribution, and monetized from that position. Now Prince wants Cloudflare to aggregate AI's access rights to the web, transforming site owners into content suppliers who can refuse or set prices, and AI labs into payers, with Cloudflare taking a cut. That's what Prince was really saying on Decoder when he repeatedly emphasized the three layers of "owners can block, allow, or pay" — he's not selling anti-scraping tools, he's selling pricing power.
What I'm uncertain about is a key difference in the analogy: when Google aggregated, site owners were generally welcoming of being indexed (because it brought traffic). In the AI era, site owners are more uniformly hostile to crawlers, feeling they're being taken for free. That hostility looks like Cloudflare's obstacle, but it's actually the source of its bargaining power — the more site owners want to block AI, the more they need Cloudflare as an intermediary.
03 Historical Analogy
The closest parallel is the music industry vs. Apple iTunes / Spotify around 2010. Musicians originally posted songs for free on MySpace and LimeWire for anyone to rip, and before Apple and Spotify built the pipeline where listeners paid monthly subscriptions with per-play royalties, creators didn't see a cent. What Cloudflare is doing now is essentially a Spotify for AI on web content — turning every crawl into a measurable, billable, refusable event.
The second parallel is the 2013-2017 snippet tax dispute between Google and European publishers. Germany and Spain led with legislation requiring search engines to pay for snippets; Google briefly shut down Google News Spain and ultimately compromised. The difference is that the clearing entity then was government legislation, while this time it's a private infrastructure provider like Cloudflare, commercializing what had been a political question (whether mandatory payment violates net neutrality).
I'm not sure whether the analogy is overreaching — the music industry involves C-end consumer payments, AI data procurement is B2B, the bargaining models may be entirely different; snippet tax was government-driven, Cloudflare is private rent-seeking, with different legal foundations. But structurally, it's the same thing: the wholesale price of content is being forcibly reset from zero.
04 What This Means for AI Builders
Three things worth adjusting this month.
First, the TCO of training data needs to be recalculated. If Cloudflare's pay-per-crawl model works, the unit cost of any model relying on large-scale web scraping (including third-party fine-tuning datasets and retrieval pipeline refreshes) will rise. When AI builders choose data sources, they need to factor in the possibility of being charged by platforms like Cloudflare within their 12-month TCO estimates, rather than making decisions on the current cost structure of "free scraping."
Second, the distribution model for agent products needs to change. Currently, agents like Cursor and Anthropic Computer Use access websites through the default behavior of "quietly GET-ing." In the future, this will most likely move to a flow with "token-based authentication + metering budgets + domain allowlists." This means agent frameworks need to quickly support two primitives — "site identity + metering budgets" — otherwise they'll be blocked by anti-scraping mechanisms when going to production. OpenAI is already pushing A2A and Apps SDK, partly because of this pressure.
Third, compliance risk is rising for self-hosted scraping pipelines (self-built proxy pools + headless browser farms). Cloudflare itself is doing anti-scraping, and AI labs are doing counter-anti-scraping — this is an endless arms race that small teams can't afford. I'm not at a major AI lab, so I don't know whether they're already budgeting for this internally, but I'd guess OpenAI and Anthropic already have dedicated teams managing it.
05 Counter-Arguments / Risks
I may be wrong in underestimating AI labs' ability to bypass intermediaries.
OpenAI has already signed direct licensing deals with News Corp, Axel Springer, Reddit, Associated Press, and others; Anthropic is taking a similar path. Once leading AI labs use a combination of "exclusive deals + direct scraping of their licensed content" to bypass Cloudflare, what Cloudflare can really take a cut on is mid-tail and long-tail sites — and the model training value of long-tail data may be far below the pricing Cloudflare wants. This is the fundamental flaw of the Spotify model: the top 1% of content delivers 90% of the value, and long-tail royalties can't sustain a platform.
The second risk is Cloudflare's own organizational capability. I have doubts about whether it can simultaneously run four business lines — selling CDN, selling Zero Trust, selling Workers, and selling AI bot settlement. Moving from the technical layer to a content rights intermediary is moving from an engineering problem to a legal and business problem, and Cloudflare has no historical evidence of strength in the latter — look at its legal disputes with Verizon and end users; it hasn't handled them gracefully.
The third is that political risk is far higher than technical risk. Once Cloudflare becomes a toll booth on the AI data highway, both antitrust (Cloudflare is already large enough) and net neutrality fronts will come calling. Nilay directly asked Prince on the podcast, "Is this a good thing?", and Prince didn't answer directly. What I heard was evasion, not that he hadn't thought about it. I may be underestimating Prince's political skill — someone who's survived the FCC, net neutrality debates, SOPA/PIPA for ten years is unlikely not to have considered this move, but the risk doesn't disappear because of that.