• Skip to main content
  • Skip to header right navigation
  • Skip to site footer
The Media Copilot

The Media Copilot

How AI is changing Media, journalism and content creation

  • News
  • Reviews
  • Guides
  • AI Courses
    • AI Quick Start
    • NEW—AI for Media
    • Custom AI Training for Teams
  • Newsletter
  • Podcast
  • Events
    • GEO Dinner Series
    • Webinars
  • About

The bots publishers should be letting through the door

Not every bot in your logs is an enemy. The ones publishers keep shutting out are often the ones building authority in AI answers.

Ornate iron gate at the entrance of a digital newspaper archive with small robot crawlers approaching
Blocking every AI crawler feels safe, but it hands authority to whoever is willing to be read. (Credit: Google Gemini)
Jul 21, 2026

By Pete Pachal

Most media companies are now pretty aggressive when it comes to blocking AI bots from their content. Reports say most major publications almost universally block crawlers from major AI companies like OpenAI and Anthropic, and many small to midsize publishers mirror that, configuring their robots exclusion protocol settings, or robots.txt, to keep them out.

It’s a defensible choice. Sometimes it’s the right one. But it also crudely compresses a complex issue into a binary: Block or don’t block. And the reasoning is often just as straightforward, pointing to the fact that most AI platforms don’t give back significant human traffic. Meanwhile, the crawl volume has become punishing enough that infrastructure cost is a factor. From an ROI standpoint, it’s an easy call.

There are exceptions. Google is the big one, because presence in AI Overviews are is currently welded into Search itself, so slamming the door on Googlebot isn’t really an option for most media companies (although the balance of that equation is starting to shift). And if you’ve cut a licensing deal with a specific platform, its crawler obviously gets a pass.

The bigger problem with a pure block-or-allow framing is that it treats direct traffic and direct revenue as the only things worth measuring. That’s too narrow a window to really get a full picture of the opportunity AI presents. An AI-forward strategy has to be more thorough: sort the bots by function, understand how each one connects your work to an audience, treat that audience as real even when it’s not a human clicking through, and shape how your content actually appears once it lands inside someone else’s answer.

Sort the bots before you swing the axe

Different bots are doing different jobs, and lumping them together is where most publisher policies go wrong. There are several types, but broadly three matter most to publishers: training bots, search bots, and retrieval bots. At least those are the ones that operate inside the “legitimate” ecosystem—declared, documented, and mostly respectful of robots.txt. Set the unauthorized scrapers aside for a minute.

The one worth a closer look is the retrieval bot. Retrieval bots typically have “-user” as part of their name, and their job is narrow: fetch a specific page in the moment a person asks an AI chatbot a question. Their activity is modest and predictable, usually hitting only a small number of pages, and they can often be rate-limited or served cached content. Cloudflare, in fact, has made it easier to distinguish between uses like search indexing, real-time AI input, and AI training, and its AI Crawl Control tools now include options to allow, charge, or block specific AI crawlers.

Letting retrieval bots through, and possibly rate-limiting a small population of search bots, can materially change how often your work shows up in AI summaries.

The reflexive publisher answer is “so what?”—you can’t cash a citation at the bank. Which is true, but it misses the strategic point: If you’re not in the answer, someone else is. Every citation your competitor gets is a small deposit into their authority on that topic, and audiences follow the citations that keep showing up. The scarce referral traffic that does exist tends to route to whoever’s already trusted by the model. Winning in AI means understanding that authority is the prize, not traffic.

Signal value, don’t donate it

Chasing authority doesn’t mean surrendering everything to get it. Your most valuable assets—whether they be content, community, or experiences—need to remain yours, not given freely to AI systems. Strategic content is where the discipline shows: careful choices about what an AI can read versus what a reader still has to come to you to see. Metadata, snippets, access controls, and explicit instructions for LLMs are the levers. Used well, they can teach a crawler that something valuable exists on your site without handing over the substance.

Take an agriculture trade publisher with a heavy, data-rich report on how pesticides affect soybean demand. A snippet might describe exactly what kind of data is within and why the data is relevant to certain types of research while not revealing the data or conclusions. The report page itself can spell out what’s open, what’s restricted, and how a qualified user gets to the full document.

The end state is asymmetric on purpose: the AI knows the report is authoritative, but anyone who wants to actually read it has to come to the publisher and clear some kind of gate—a subscription, an email, a partner login. The goal isn’t to hide entirely from AI answer engines. It’s to make them aware of the publisher’s value without letting them reproduce that value.

  • Subscribe to our newsletter

    How AI is changing media, journalism, and content creation.

    Learn More

Your archive is your alpha

There’s an interesting framing here. Palantir cofounder Alex Karp said in a widely shared CNBC interview, where he advised enterprise AI customers to stop using models from the major AI labs, since it was effectively giving them their “alpha”—the company data that gives them an edge over competitors. Media companies, unfortunately, didn’t get to make that choice cleanly. Most of what publishers produce is public, and much of it was already ingested into the first wave of foundation models.

Those same lab-built models are now competitors, not just infrastructure. Answer engines keep users on the platform instead of routing them to the outlet that produced the reporting. This is, of course, the foundational idea behind the many lawsuits and licensing deals between the AI companies and the media.

But the market has moved past the original training-data fight. To give users the best and most current answer, the model needs live information at the moment of the query, and that shifts the value from training rights to retrieval rights. Rob Kelly has a useful analysis of this shift, showing that training rights are no longer a given in publicly announced AI licensing deals. In his database, only about 4 in 10 public 2026 deals include training rights, a sign that the market is moving from “buy content to build better models” toward “license content to deliver better answers.”

Most publishers aren’t going to land their own deal with OpenAI, Anthropic, or the rest. That’s the honest baseline. It doesn’t mean there’s no way to extract value from your content in an AI marketplace—it means the first step is making your content machine-readable in the first place, which is a separate discipline from bot blocking.

The number of places publishers can market their content to AI experiences is growing. Factiva, for example, includes content from thousands of suppliers, all licensed and Microsoft describes its new Publisher Content Marketplace as a way to support licensed access to premium content while preserving publisher control, independence, and sustainable revenue. The most ambitious move is to build your own agentic layer that legitimate customers and partners can hit via MCP (model context protocol).

That requires real engineering, often with a partner, but the payoff is control over the experience. Rather than letting some third-party crawler scrape and interpret your material, you’ve already done the interpretation, and any authorized external bot gets the version you’ve pre-processed — a relay, not a raw pull. That’s the whole idea behind AI media projects like Reuters’ new MCP server, which allows customers to search, retrieve, and use the Reuters content they subscribe to inside AI workflows.

I’ve made the case before that publisher-built agents and AI-ready archives change the shape of the business: once you’ve done the hard work of formatting, ingesting, and processing your archive for AI, you start to look less like a content supplier and more like a tool vendor. Reuters clearly recognizes that publishers who don’t build some kind of controlled retrieval layer over their archives are letting third-party crawlers set the rules. Eventually, that will lead to negotiating from weakness.

Now a word about the bad actors. Unauthorized scrapers hoover up huge parts of what publishers produce and then sell that data in various black and gray markets. Certainly, this is where blocking should be encouraged, and publishers need both reliable tools for doing so as well as broader ecosystem support, such as what Cloudflare has done to better identify bots and potentially monetize their activity. As I wrote earlier this year, unauthorized AI crawling is rampant, and publishers need more than wishful thinking and a robots.txt file to deal with it.

Turn defense into leverage

Blocking, on its own, is a defensive posture, and defense doesn’t win the AI cycle. The answer is not to throw the doors open or nail them shut. It is to build a gate with rules. Give retrieval bots enough to make your authority visible. Keep training crawlers and unauthorized scrapers away from the material they have no business ingesting. Use snippets, metadata, paywalls, and rate limits to separate discovery from access. Then do the harder work behind the gate: clean the archive, structure it, make it legible to machines, and expose it through channels you actually control.

The giant AI licensing deal is not coming for most publishers. That’s fine—it was never a real business plan. The better play is to make your work legible to AI systems without making it free, so when the market does come looking for trusted answers, you’re not begging to be included. You’re already the gate.

A version of the column appears in Fast Company.

Contributors

  • Pete Pachal: Author

    Pete Pachal is the founder of The Media Copilot. In addition to producing the site’s newsletter and podcast, he also teaches courses on how journalists and communications professionals can apply AI tools to their work. Pete has a long career in journalism, previously holding senior roles in global newsrooms such as CoinDesk and Mashable. He’s appeared on Fox Business, CNN, and The Today Show as a thought leader in tech and AI. Pete also puts his encyclopedic knowledge of Doctor Who to good use on the popular podcast, Pull To Open.

Category: AI media analysisTags:bot blocking| agents
Share this post:
FacebookTweetLinkedInEmail

What do 1,000 journalists and PR pros know about AI that you don't? They took AI Quick Start, a 1-hour live class from The Media Copilot. 94% satisfaction. Find out how to work smarter with AI in just 60 minutes. Get 20% off with the code AIPRO: https://mediacopilot.ai/

  • Related articles

Illustration of a woman at a control panel managing AI company toggles for OpenAI, Anthropic, Google, and Microsoft

Creators get new say over AI scraping through Cloudflare–beehiiv partnership 

Read moreCreators get new say over AI scraping through Cloudflare–beehiiv partnership 
Illustration of friendly robots passing through a glowing gate toward menacing red-eyed robots

Reuters and Time flip the script on AI bots with blocking whitelists

Read moreReuters and Time flip the script on AI bots with blocking whitelists
AI content scraping

Inside the AI scraping economy nobody wants to talk about

Read moreInside the AI scraping economy nobody wants to talk about
Photo of a focused professional in an office setting

What an agentic newsroom will look like

Read moreWhat an agentic newsroom will look like
Smartphone displaying the Claude Mythos logo on a keyboard

UK and US financial regulators hold emergency meetings over Anthropic’s Claude Mythos

Read moreUK and US financial regulators hold emergency meetings over Anthropic’s Claude Mythos
Human hand on keyboard with ghostly AI agent hands working on floating task panels — illustrating Microsoft Copilot agentic workflows

Microsoft turns Microsoft 365 Copilot into a broader agentic work platform

Read moreMicrosoft turns Microsoft 365 Copilot into a broader agentic work platform

The Media Copilot

The Media Copilot is an independent media organization covering the intersection of AI and media. Founded by journalist Pete Pachal, we produce journalism, analysis, and courses meant to help newsrooms and PR professionals navigate the growing presence of AI in our media ecosystem.

  • LinkedIn
  • X
  • YouTube
  • Instagram
  • TikTok
  • Bluesky
  • About The Media Copilot
  • Advertising & Sponsorships
  • Our Methodology
  • Privacy Policy
  • Membership
  • Newsletter
  • Podcast
  • Contact

© 2026 · All Rights Reserved · Powered by Springwire.ai · RSS