By
As AI companies continue scraping the web for training data, some publishers are experimenting with a new tactic that doesn’t only block unwanted bots, but tries to waste their time and money.
Known as LLM honeypotting, Digiday describes the approach as a form of deception to lure AI crawlers into consuming plausible-looking but ultimately worthless content, aimed at making large-scale data collecting so computationally expensive that it becomes less economically viable.
Who’s doing this? A small number of publishers and e-commerce companies are looking for alternatives to traditional bot-blocking. And while the technique remains early and experimental, it’s one way for media companies to protect their content amid a fight to create standards for how AI companies track, value and compensate journalism.
A cybersecurity tactic adapted for AI
Cyberhoneypotting is a long-standing cybersecurity strategy. The technique is to build a decoy target for attackers—or LLM bots, in this case—that can lure bad actors away and even gather intelligence on their capabilities and methods. In this case, the goal is to make scraping content without compensation more costly than it’s worth.
Simon Wistow, co-founder of content delivery network provider Fastly, describes the philosophy as one of “Chang[ing] the economics of attacking.” If abusing a system becomes significantly more expensive than the value gained, the entire model will become unsustainable, he argues.
Applied to AI crawlers, the strategy is equipped against all automated visitors, regardless of whether they are operated by major AI companies or smaller third-party scraping firms.
Publishers can implement the tactic in several ways. They can introduce subtle delays, difficult (for computers) problems to solve before admittance, or endless mazes of contents and files filled with meaningless AI-generated text. Some honeypotting techniques go even further and try to inject bad data into the bots’ training datasets, “poisoning” the AI results.
The tactic’s adoption
Large e-commerce brands are already testing the technology successfully, said Wistow. News publishers are also showing increased interest, although he declined to identify specific customers.
Even so, skepticism towards the strategy remains.
Frederick Jahn, co-founder of AI company Centennal, argues sophisticated scrapers can often detect or avoid honeypots altogether.
“I think it’s a good concept, but more on a marketing level, and like a gimmick,” Jahn said. He argues that publishers would be better served by creating real barriers to stealth crawlers, who are often not shown maze pages.
Supporters of honeypotting maintain that, even if most scrapers adapt, increasing operational costs across thousands or millions of requests could make smaller scraping businesses financially unsustainable.
“If they could burn through that 10 million funding in one crawl then suddenly those businesses aren’t viable and suddenly the whole market collapses, and that’s kind of what you’re going for,” said Wistow.

Costs and limitations for publishers
The strategy isn’t free. Generating and serving millions of fake pages is more expensive than simply blocking unwanted traffic. And larger publishers with more resources are better able to plan and implement the strategy.
Wistow said the approach is unlikely to become widespread, in part because of the consequences of filling the internet with yet more intentionally deceptive content.
“Hallucinations happen even with good data, just because of the way LLMs work,” said Wistow. “This is about changing the economics for the people abusing your site, not running some giant disinformation campaign.”







