• Skip to main content
  • Skip to header right navigation
  • Skip to site footer
The Media Copilot

The Media Copilot

How AI is changing Media, journalism and content creation

  • News
  • Reviews
  • Guides
  • AI Courses
    • AI Quick Start
    • NEW—AI for Media
    • Custom AI Training for Teams
  • Newsletter
  • Podcast
  • Events
    • GEO Dinner Series
    • Webinars
  • About

Publishers Turn to AI ‘Honeypots’ to Fight Content Scraping

As AI companies continue collecting web data for training, publishers are testing digital traps designed to waste crawlers’ time and make large-scale scraping more expensive.

Some publishers are testing AI “honeypots” that lure web crawlers into endless mazes of worthless content, raising the cost of scraping journalism without permission or payment. (Credit: ChatGPT)
Jul 21, 2026

By Romy Abu-Fadel

As AI companies continue scraping the web for training data, some publishers are experimenting with a new tactic that doesn’t only block unwanted bots, but tries to waste their time and money.

Known as LLM honeypotting, Digiday describes the approach as a form of deception to lure AI crawlers into consuming plausible-looking but ultimately worthless content, aimed at making large-scale data collecting so computationally expensive that it becomes less economically viable. 

Who’s doing this? A small number of publishers and e-commerce companies are looking for alternatives to traditional bot-blocking. And while the technique remains early and experimental, it’s one way for media companies to protect their content amid a fight to create standards for how AI companies track, value and compensate journalism. 

A cybersecurity tactic adapted for AI

Cyberhoneypotting is a long-standing cybersecurity strategy. The technique is to build a decoy target for attackers—or LLM bots, in this case—that can lure bad actors away and even gather intelligence on their capabilities and methods. In this case, the goal is to make scraping content without compensation more costly than it’s worth. 

Simon Wistow, co-founder of content delivery network provider Fastly, describes the philosophy as one of “Chang[ing] the economics of attacking.” If abusing a system becomes significantly more expensive than the value gained, the entire model will become unsustainable, he argues. 

Applied to AI crawlers, the strategy is equipped against all automated visitors, regardless of whether they are operated by major AI companies or smaller third-party scraping firms. 

Publishers can implement the tactic in several ways. They can introduce subtle delays, difficult (for computers) problems to solve before admittance, or endless mazes of contents and files filled with meaningless AI-generated text. Some honeypotting techniques go even further and try to inject bad data into the bots’ training datasets, “poisoning” the AI results.

The tactic’s adoption

Large e-commerce brands are already testing the technology successfully, said Wistow. News publishers are also showing increased interest, although he declined to identify specific customers. 

Even so, skepticism towards the strategy remains. 

Frederick Jahn, co-founder of AI company Centennal, argues sophisticated scrapers can often detect or avoid honeypots altogether. 

“I think it’s a good concept, but more on a marketing level, and like a gimmick,” Jahn said. He argues that publishers would be better served by creating real barriers to stealth crawlers, who are often not shown maze pages. 

Supporters of honeypotting maintain that, even if most scrapers adapt, increasing operational costs across thousands or millions of requests could make smaller scraping businesses financially unsustainable. 

“If they could burn through that 10 million funding in one crawl then suddenly those businesses aren’t viable and suddenly the whole market collapses, and that’s kind of what you’re going for,” said Wistow.

  • Subscribe to our newsletter

    How AI is changing media, journalism, and content creation.

    Learn More

Costs and limitations for publishers

The strategy isn’t free. Generating and serving millions of fake pages is more expensive than simply blocking unwanted traffic. And larger publishers with more resources are better able to plan and implement the strategy.

Wistow said the approach is unlikely to become widespread, in part because of the consequences of filling the internet with yet more intentionally deceptive content. 

“Hallucinations happen even with good data, just because of the way LLMs work,” said Wistow. “This is about changing the economics for the people abusing your site, not running some giant disinformation campaign.”

Contributors

  • Romy Abu-Fadel: Author

    Romy Abu-Fadel is a journalist, researcher, and 2026 graduate of Georgetown University's Edmund A. Walsh School of Foreign Service. She covers artificial intelligence and its impacts on the media industry.

  • Christopher Allbritton: Editor

    Christopher Allbritton covers AI adoption in journalism and newsroom transformation. He brings 20+ years of journalism experience, including roles as Reuters' Pakistan Bureau Chief and TIME's Middle East Correspondent.

Category: NewsTags:privacy| publishers| webscraping| cybersecurity
Share this post:
FacebookTweetLinkedInEmail

What do 1,000 journalists and PR pros know about AI that you don't? They took AI Quick Start, a 1-hour live class from The Media Copilot. 94% satisfaction. Find out how to work smarter with AI in just 60 minutes. Get 20% off with the code AIPRO: https://mediacopilot.ai/

  • Related articles

A journalist sits at a cluttered London newsroom desk staring at a monitor displaying a sharply declining analytics traffic graph under fluorescent office lighting

Google search traffic to drop by half by Q3 2027 for UK publishers

Read moreGoogle search traffic to drop by half by Q3 2027 for UK publishers
Overhead view of a nighttime digital newsroom with journalists at monitors and a wall display showing a rapidly rising crawl counter beside a flatlined referral traffic graph

Meta now drives most AI agent traffic while sending publishers few visitors

Read moreMeta now drives most AI agent traffic while sending publishers few visitors
Stack of dog-eared Fiction Feast magazines on a cluttered writing desk beside a handwritten manuscript, a lamp casting warm light over an empty chair

Bauer’s Take a Break drops freelance writers as AI drafts fiction stories

Read moreBauer’s Take a Break drops freelance writers as AI drafts fiction stories
A fictional byline photo dissolves into pixels on a glowing screen, surrounded by Alabama small-town newspaper printouts while a hand holds a phone confirming the papers are active

AI fake news network invents the collapse of 47 local Alabama newspapers

Read moreAI fake news network invents the collapse of 47 local Alabama newspapers
Cloudflare bouncer protecting club from bots

Cloudflare’s new plan could change how AI pays publishers

Read moreCloudflare’s new plan could change how AI pays publishers
Exterior of Cloudflare's corporate headquarters

Cloudflare will block AI training crawlers by default on ad-supported sites

Read moreCloudflare will block AI training crawlers by default on ad-supported sites

The Media Copilot

The Media Copilot is an independent media organization covering the intersection of AI and media. Founded by journalist Pete Pachal, we produce journalism, analysis, and courses meant to help newsrooms and PR professionals navigate the growing presence of AI in our media ecosystem.

  • LinkedIn
  • X
  • YouTube
  • Instagram
  • TikTok
  • Bluesky
  • About The Media Copilot
  • Advertising & Sponsorships
  • Our Methodology
  • Privacy Policy
  • Membership
  • Newsletter
  • Podcast
  • Contact

© 2026 · All Rights Reserved · Powered by Springwire.ai · RSS