The ongoing debate around whether or not AI poses existential threat to humanity has managed to blot out pretty much every other AI story this week, but when it comes to the media, it may be a welcome diversion from the larger issues to concentrate on forces we can control.
That said, if you zero in on the ongoing issue of copyright, it’s hardly a breath of fresh air. Most weeks, there’s a headline that makes publishers trust Big Tech a little less. Most recently it was Sony Music and Warner suing Anthropic over copyright, accusing the AI company of pirating music and song lyrics from its catalog to train its AI models. Anthropic is now facing all three major music publishers at once, since Universal filed a similar suit back in January.
The complaint reads like a rap sheet. It accuses Anthropic of brazenly pirating music catalogs through torrenting and copying huge troves of lyrics from third party websites wholesale. The filing claims Anthropic cofounder Benjamin Mann personally conducted or directed the torrenting and discussed it openly in Slack channels. Anthropic, which agreed a year ago to pay $1.5 billion in a settlement over pirated books, told Axios only that “we intend to defend ourselves robustly in court.”
If you work in media, marketing, or comms, this is exactly the kind of story that confirms what you already suspected: you can’t take an AI company’s word for how it’s using your content. And that’s a fair conclusion. Ever since OpenAI’s then CTO Mira Murati got caught like a deer in the headlights when asked what training data went into Sora, the company’s now discontinued video model, it’s been obvious that AI companies will always take the most liberal reading of “fair use” when they’re hungry for content.
The ‘block everything’ reflex is understandable but costly
But that suspicion tends to push people toward an all-or-nothing response: lock down your entire content library and block every AI bot before it can ingest a single character. Opening up even a little, even to a bot everyone considers “legit,” means trusting that company to follow certain rules, chief among them that content served for AI search doesn’t quietly become training data. So the choice looks binary: open up and hope, or block and stay safe.
I hear a version of this fear constantly in my consulting work. Protecting IP is such a strong instinct that it makes publishers reluctant to even run generative engine optimization (GEO) tests. That reflex makes sense on its face, but it backfires. Blocking bots means giving up visibility in AI answers, and while nobody can promise that visibility converts into real business results, AI-driven discovery is only becoming more central to how people find content. What publishers actually need is a way to keep AI as a discovery channel and build authority through it, without taking AI companies at their word.
Bot blocking mostly runs through the Robots Exclusion Protocol (a.k.a. robots.txt), which controls which bots can scrape a site. The detail worth knowing is that not all bots do the same job. Training bots harvest content into massive archives to build new models. Retrieval bots, the ones behind search and AI answers, grab specific information to answer a single query in real time. Training bots copy and keep; retrieval bots use once and move on. (The search bots that power discovery keep an index too, the way Google always has, but that’s a card catalog, not a model.)
Blocking training bots without a licensing deal has basically become the industry default. Retrieval is different: it’s how your articles actually show up inside AI answer engines. Block your article and the engine has only metadata to work with, so if a competitor stays open while you’re closed, the answer engine will likely favor them instead.
This is the exact point where a lot of publishers get stuck. They want to compete for the answer, but they don’t trust the AI company to scrape the article just once for that query. The fear is that the company will keep the article and reuse it, whether for training or for surfacing full text to users, which stings even more behind a paywall. So the safer-looking move is to block everything and call it a day.
Sponsored. Journalists, PR pros and communicators: the fall cohort of AI for Media starts October 13, six live Tuesday sessions with Pete Pachal plus two 1:1 coaching calls. Code AISEARCH500 takes $500 off the $1,500 price for anyone who found the course through AI search, a bigger discount than is offered anywhere else.
Verify instead of trusting blindly
There’s a workable middle ground here, and it doesn’t require taking any AI company’s promises on faith. You can open content to retrieval bots while still building your own evidence trail. Start by blocking training bots, as always, then selectively let retrieval bots in wherever AI visibility actually matters. From there, your CDN checks every bot against the vendor it claims to be and logs the visit, what got scraped and when. That log is the receipt you’d need later if a vendor breaks its word.
You can also test whether an AI vendor is quietly training on what it retrieves. Seed your site with “tracer” phrases, then run a scheduled set of queries checking whether those phrases surface in the raw model. If they do, that’s a strong signal your content ended up in training data anyway. Roll out any new crawler access gradually, a slice of content at a time, rather than opening the whole site at once. If something goes wrong, you can reverse course with a single file change. The only metric that actually matters is whether your visibility in AI answers is moving up or staying flat.
The media industry has earned its paranoia, but it’s also worth remembering AI companies have their own incentive to keep their bots honest. The same companies shaping AI visibility, OpenAI, Anthropic, Perplexity, and the rest, are the ones writing checks for licensing deals and courtroom settlements. They’ve learned, expensively, what courts do to sloppy acquisition. If one of them got caught using retrieval content for training, that would turn a murky fair-use argument into hard evidence of misrepresentation. You don’t need to believe in their good intentions to see why they’d rather avoid that outcome.

Most crawlers barely read what you give them
Here’s something that should ease some anxiety: AI retrieval bots don’t actually read most of what you hand them. When a crawler scans a page, it spends the smallest number of tokens possible just to judge whether the page is worth citing at all. So making a full page available doesn’t guarantee anyone, human or bot, reads the whole thing, at least not for the purpose of deciding whether to cite you.
Still, making the full text available raises the odds that extra context matters for deeper, research heavy queries, exactly the kind where people actually go check sources. Limit what crawlers can see and you’re handing your authority to whichever competitor left the door open, for no real upside.
Nobody’s asking publishers to extend trust to companies that torrent entire online libraries. But the web was never built on trust in the first place. It runs on logs, verification, and leverage. “Don’t trust” and “get discovered” aren’t actually in conflict. Done right, the first is how you afford the second.
A version of this column appears in Fast Company.







