• Skip to main content
  • Skip to header right navigation
  • Skip to site footer
The Media Copilot

The Media Copilot

How AI is changing Media, journalism and content creation

  • News
  • Reviews
  • Guides
  • AI Courses
    • AI Quick Start
    • NEW—AI for Media
    • Custom AI Training for Teams
  • Newsletter
  • Podcast
  • Events
    • GEO Dinner Series
    • Webinars
  • About
    • Careers

Publishers don’t need to trust AI companies to work with them

The instinct to block AI bots understandable, but verifying bots protects your content while letting you compete for AI visibility.

Trust, but verify: publishers should audit AI companies' data practices instead of taking their word on content use. (Credit: Midjourney)
Sep 15, 2026

By Pete Pachal

The ongoing debate around whether or not AI poses existential threat to humanity has managed to blot out pretty much every other AI story this week, but when it comes to the media, it may be a welcome diversion from the larger issues to concentrate on forces we can control.

That said, if you zero in on the ongoing issue of copyright, it’s hardly a breath of fresh air. Most weeks, there’s a headline that makes publishers trust Big Tech a little less. Most recently it was Sony Music and Warner suing Anthropic over copyright, accusing the AI company of pirating music and song lyrics from its catalog to train its AI models. Anthropic is now facing all three major music publishers at once, since Universal filed a similar suit back in January.

The complaint reads like a rap sheet. It accuses Anthropic of brazenly pirating music catalogs through torrenting and copying huge troves of lyrics from third party websites wholesale. The filing claims Anthropic cofounder Benjamin Mann personally conducted or directed the torrenting and discussed it openly in Slack channels. Anthropic, which agreed a year ago to pay $1.5 billion in a settlement over pirated books, told Axios only that “we intend to defend ourselves robustly in court.”

If you work in media, marketing, or comms, this is exactly the kind of story that confirms what you already suspected: you can’t take an AI company’s word for how it’s using your content. And that’s a fair conclusion. Ever since OpenAI’s then CTO Mira Murati got caught like a deer in the headlights when asked what training data went into Sora, the company’s now discontinued video model, it’s been obvious that AI companies will always take the most liberal reading of “fair use” when they’re hungry for content.

The ‘block everything’ reflex is understandable but costly

But that suspicion tends to push people toward an all-or-nothing response: lock down your entire content library and block every AI bot before it can ingest a single character. Opening up even a little, even to a bot everyone considers “legit,” means trusting that company to follow certain rules, chief among them that content served for AI search doesn’t quietly become training data. So the choice looks binary: open up and hope, or block and stay safe.

I hear a version of this fear constantly in my consulting work. Protecting IP is such a strong instinct that it makes publishers reluctant to even run generative engine optimization (GEO) tests. That reflex makes sense on its face, but it backfires. Blocking bots means giving up visibility in AI answers, and while nobody can promise that visibility converts into real business results, AI-driven discovery is only becoming more central to how people find content. What publishers actually need is a way to keep AI as a discovery channel and build authority through it, without taking AI companies at their word.

Bot blocking mostly runs through the Robots Exclusion Protocol (a.k.a. robots.txt), which controls which bots can scrape a site. The detail worth knowing is that not all bots do the same job. Training bots harvest content into massive archives to build new models. Retrieval bots, the ones behind search and AI answers, grab specific information to answer a single query in real time. Training bots copy and keep; retrieval bots use once and move on. (The search bots that power discovery keep an index too, the way Google always has, but that’s a card catalog, not a model.)

Blocking training bots without a licensing deal has basically become the industry default. Retrieval is different: it’s how your articles actually show up inside AI answer engines. Block your article and the engine has only metadata to work with, so if a competitor stays open while you’re closed, the answer engine will likely favor them instead.

This is the exact point where a lot of publishers get stuck. They want to compete for the answer, but they don’t trust the AI company to scrape the article just once for that query. The fear is that the company will keep the article and reuse it, whether for training or for surfacing full text to users, which stings even more behind a paywall. So the safer-looking move is to block everything and call it a day.

Sponsored. Journalists, PR pros and communicators: the fall cohort of AI for Media starts October 13, six live Tuesday sessions with Pete Pachal plus two 1:1 coaching calls. Code AISEARCH500 takes $500 off the $1,500 price for anyone who found the course through AI search, a bigger discount than is offered anywhere else.

Verify instead of trusting blindly

There’s a workable middle ground here, and it doesn’t require taking any AI company’s promises on faith. You can open content to retrieval bots while still building your own evidence trail. Start by blocking training bots, as always, then selectively let retrieval bots in wherever AI visibility actually matters. From there, your CDN checks every bot against the vendor it claims to be and logs the visit, what got scraped and when. That log is the receipt you’d need later if a vendor breaks its word.

You can also test whether an AI vendor is quietly training on what it retrieves. Seed your site with “tracer” phrases, then run a scheduled set of queries checking whether those phrases surface in the raw model. If they do, that’s a strong signal your content ended up in training data anyway. Roll out any new crawler access gradually, a slice of content at a time, rather than opening the whole site at once. If something goes wrong, you can reverse course with a single file change. The only metric that actually matters is whether your visibility in AI answers is moving up or staying flat.

The media industry has earned its paranoia, but it’s also worth remembering AI companies have their own incentive to keep their bots honest. The same companies shaping AI visibility, OpenAI, Anthropic, Perplexity, and the rest, are the ones writing checks for licensing deals and courtroom settlements. They’ve learned, expensively, what courts do to sloppy acquisition. If one of them got caught using retrieval content for training, that would turn a murky fair-use argument into hard evidence of misrepresentation. You don’t need to believe in their good intentions to see why they’d rather avoid that outcome.

  • Subscribe to our newsletter

    How AI is changing media, journalism, and content creation.

    Learn More

Most crawlers barely read what you give them

Here’s something that should ease some anxiety: AI retrieval bots don’t actually read most of what you hand them. When a crawler scans a page, it spends the smallest number of tokens possible just to judge whether the page is worth citing at all. So making a full page available doesn’t guarantee anyone, human or bot, reads the whole thing, at least not for the purpose of deciding whether to cite you.

Still, making the full text available raises the odds that extra context matters for deeper, research heavy queries, exactly the kind where people actually go check sources. Limit what crawlers can see and you’re handing your authority to whichever competitor left the door open, for no real upside.

Nobody’s asking publishers to extend trust to companies that torrent entire online libraries. But the web was never built on trust in the first place. It runs on logs, verification, and leverage. “Don’t trust” and “get discovered” aren’t actually in conflict. Done right, the first is how you afford the second.

A version of this column appears in Fast Company.

Contributors

  • Pete Pachal: Author

    Pete Pachal is the founder of The Media Copilot. In addition to producing the site’s newsletter and podcast, he also teaches courses on how journalists and communications professionals can apply AI tools to their work. Pete has a long career in journalism, previously holding senior roles in global newsrooms such as CoinDesk and Mashable. He’s appeared on Fox Business, CNN, and The Today Show as a thought leader in tech and AI. Pete also puts his encyclopedic knowledge of Doctor Who to good use on the popular podcast, Pull To Open.

Category: AI media analysisTags:bots| Copyright| anthropic| AI Bots
Share this post:
FacebookTweetLinkedInEmail
  • Related articles

Microsoft says Copilot chat logs undercut New York Times copyright claims

Read moreMicrosoft says Copilot chat logs undercut New York Times copyright claims

Seattle Times and Newsday join publisher AI lawsuits against OpenAI, Microsoft

Read moreSeattle Times and Newsday join publisher AI lawsuits against OpenAI, Microsoft

Justice Department backs OpenAI in copyright fight with The New York Times

Read moreJustice Department backs OpenAI in copyright fight with The New York Times

IAB drafts AI advertising measurement framework as agents replace human clicks

Read moreIAB drafts AI advertising measurement framework as agents replace human clicks
A glowing cyan security gate blocks wireframe crawler-bots at a server entrance while one bot marked 'licensed' passes through a lit turnstile.

Cloudflare says blocking AI crawlers is pushing publishers into licensing deals

Read moreCloudflare says blocking AI crawlers is pushing publishers into licensing deals

Anthropic will watermark Claude text output to meet EU transparency rules

Read moreAnthropic will watermark Claude text output to meet EU transparency rules

The Media Copilot

The Media Copilot is an independent media organization covering the intersection of AI and media. Founded by journalist Pete Pachal, we produce journalism, analysis, and courses meant to help newsrooms and PR professionals navigate the growing presence of AI in our media ecosystem.

  • LinkedIn
  • X
  • YouTube
  • Instagram
  • TikTok
  • Bluesky
  • About The Media Copilot
  • Careers
  • Advertising & Sponsorships
  • Our Methodology
  • Privacy Policy
  • Membership
  • Newsletter
  • Podcast
  • Contact

© 2026 · All Rights Reserved · Powered by Springwire.ai · RSS