The bots are winning.
That’s the clearest takeaway from the first section of TollBit’s most recent State of the Bots report. TollBit builds payment rails between publishers and AI crawlers, and the report includes commentary from Jonathan Roberts, Chief Innovation Officer at People Inc., who lays out the company’s approach to AI bots scraping content across its media properties: block unauthorized bots aggressively, allow access to legitimate crawlers, usually under some kind of licensing agreement.
It sounds good on paper. Then Roberts admits the part that matters: the bad actors have gotten too good to fully stop. Bad bots masquerade as legitimate ones like Google’s crawler, rotate IP addresses the moment they get blocked, and in some cases route traffic through networks of home devices to look like ordinary human visitors. People Inc. has real resources and deep expertise here, and it still can’t block everything.
That’s worth taking in if you run content, audience, or comms. If a company with People Inc.’s scale and know-how can’t fully lock down its content, what chance does a newsroom or brand with less resources have? Which means the fight over bot access can’t be the place where publishers define their future in the AI era. Build a strategy around winning that fight and you’ve already lost.
I should say at this point that Roberts reached to me directly after this column first ran in Fast Company. He pointed out that the big platforms, Google, Anthropic, OpenAI, aren’t the stealth bots I’m describing here, since they declare their crawlers, which is a fair point. Still, third-party scraping is big business, even if those bots are a small portion of the total. A little goes a long way.
But the smarter question isn’t which bots to block It’s what happens once they’re inside. Publishers shouldn’t throw the doors open to every crawler, but the more useful work is figuring out how content shows up for the end user, and what levers exist there to protect value, reward good behavior, and actually build a business.
The training fight is basically dead
This next part matters for anyone still treating training and retrieval as the same threat. They’re not, and conflating them leads to the wrong strategy. The AI-media fight has largely moved past training, not because it was ever fine for AI companies to train on crawled content, but because the thing that actually threatens media business models now is retrieval: people using AI as their discovery layer. Retrieval needs accurate, current information. Training doesn’t work that way.
Training large language models is something very few companies actually do, mostly because it’s expensive; training runs can run into the hundreds of millions or billions of dollars, and the goal is a model that predicts better, not one that answers informational queries reliably. That gap showed up constantly in the early days of consumer AI. Ask a chatbot something like the details of the Enlightenment and you’d get an answer that was roughly right in shape but shaky on details like names and dates. For accurate information, the AI has to retrieve it in real time. That’s a fundamentally different kind of bot request, with different stakes, because it’s serving one user’s query rather than absorbing content into a training pile.
None of this means training bots get a pass. Absent a deal, publishers should block them. But they should also drop any expectation of getting paid for training data. Few publishers have the scale to make their archive valuable enough to AI labs, and licensing deals have already shifted focus from training toward retrieval. Rob Kelly of the Media and the Machine newsletter has tracked this: of 94 publicly announced AI licensing deals, only about four in ten still include training rights. The market isn’t buying content to build smarter models anymore. It’s licensing content to deliver better answers, and that’s a distinction every publisher and comms team should have memorized by now.
Getting cited doesn’t pay the bills
A recent Digiday story looked at how hard it is for brands to connect AI visibility to actual business results, and the same problem applies to publishers, just with sharper edges. There’s some value in being the authoritative source an AI cites when it answers a question, but that value isn’t, by itself, something you can monetize.
That may be beginning to change. Right now, though, being cited is often where the relationship with the reader ends, not where it starts. Study after study backs this up: Pew Research found that clicks on links inside a Google AI summary land around 1% of visits, compared with 15% on a results page with no AI answer at all.
What do 1,000 journalists and PR pros know about AI that you don't? They took AI Quick Start, a 1-hour live class from The Media Copilot. 94% satisfaction. Find out how to work smarter with AI in just 60 minutes. Get 20% off with the code AIPRO: https://mediacopilot.ai/
The reader still gets value even when the publisher doesn’t get paid, and every emerging business model—pay-per-crawl, pay-per-use, licensing, ads served to bots—is really just an attempt to put a price on that gap. None of those models has become an industry standard yet, largely because none of them are easy to enforce. There are simply too many paths for content to leak into the ecosystem: stealthy bots crawling where they shouldn’t, AI companies buying data from gray-market scrapers, or plain old republishing and repackaging.
So flip the question. Instead of trying to stop every scrape, what if the leverage point was whether the content shows up in the answer at all?

Police the answer, not the crawl
Picture how this would work in an ideal world: A person asks a question, the answer engine goes and finds the best content to use, and provides a link back. We already know attribution works reasonably well inside retrieval systems: Every major AI engine gives citations today, and ProRata’s entire business model depends on accurately showing which sources contributed to an answer and how much.
Now add one more checkpoint. After the engine identifies its sources, it runs a second check: does it actually have legitimate access to each one, whether through licensing or some other business arrangement? Fail that check, and the content can’t be used. Company-level licensing makes this easy to verify. Pay-per-use or pay-per-crawl models could build average spend into their fees, or let the user allocate a budget for it.
Courts already run on a version of this logic: evidence obtained improperly can’t be used, even if it’s true. Process matters. Online content currently runs on the opposite assumption: that if it was reachable, it must be free for the taking. As bots keep outpacing anyone’s ability to block them, the answer layer becomes the obvious place to draw a line. Most scrapers, such as Common Crawl, Parallel, Diffbot, don’t operate major consumer-facing AI engines. The places people actually get their answers are a short list of large companies, which means publishers can’t chase every scraper, but they can write enforceable rules for the handful of surfaces that actually reach a human.
Follow that logic and a pricing model falls out of it almost naturally: every retrieval is its own transaction. If ChatGPT pulls information from your site twice in a day for two different users, those are two separate proxies for two separate people, not one bulk request from OpenAI. ChatGPT can’t buy a single subscription and call it covered for everyone. It’s Sam Altman’s micropayments idea, just arriving through a side door, and it’s the direction Cloudflare is already moving toward.
Thirty Napsters, Still No Spotify
None of this happens because platforms decide to be good citizens. The companies that would need to run these checks are the same ones benefiting from skipping them, and publishers have almost no leverage to force the issue. The one party that does is government, which is part of why some form of regulation now feels close to inevitable. New York’s Stealth Crawler Prohibition Act, which would force bots to identify themselves, is a reasonable start, and a bipartisan federal Stealth Bots Bill is already working its way through Congress. Still, bot identification if a fairly low bar, and the fact that it took legislation to clear that bar says plenty about where things stand.
Roberts’ framing is dead on: There are now about 30 “Napsters of content,” and still no Spotify. The reason is straightforward. Nobody has built the mechanism that makes misusing someone’s work actually cost something. That check doesn’t belong at the crawler, where publishers keep losing ground. It belongs at the answer, where a small number of companies decide what billions of people see. Publishers have spent three years guarding the door. The more useful fight is on the other side of that door, in the answer itself. It’s time to start treating it that way.
A version of this column appears in Fast Company.







