• Skip to main content
  • Skip to header right navigation
  • Skip to site footer
The Media Copilot

The Media Copilot

How AI is changing Media, journalism and content creation

  • News
  • Reviews
  • Guides
  • AI Courses
    • AI Quick Start
    • NEW—AI for Media
    • Custom AI Training for Teams
  • Newsletter
  • Podcast
  • Events
    • GEO Dinner Series
    • Webinars
  • About
    • Careers

Publishers finally have AI buyers but no say in the price

A real market for AI inference is forming, with payment rails to match. What publishers still lack is leverage.

Illustration of a newsstand vendor watching robot buyers hold up blank price tags while a conveyor belt carries newspapers away
Data brokers have built an estimated $1 billion business reselling scraped content, and very little of that money reaches the newsrooms that produced it. (Credit: Google Gemini)
Oct 6, 2026

By Pete Pachal

If you work in the media, you’ve probably heard some version of the origin story the AI labs like to tell. In the early days of generative AI, they were so heads-down building large language models that nobody thought much about the raw material. After all, big data sets like Common Crawl already existed, and trawling the open web was how search engines had worked for decades. Chatbots were arguably a variation on the same idea. Who could’ve guessed that sourcing information this way would get so controversial?

Everybody, as it turns out. In a brief made public Sept. 17 in The New York Times’s lawsuit against OpenAI and Microsoft, internal documents show people at both companies saying things that suggest they understood exactly what they were taking, and how it would look.

Microsoft’s Brent Hecht, a director of applied science, warned in a memo in early 2023 that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft,” and called it “the largest theft of labor in human history.” When a researcher described getting around the Times paywall, OpenAI President Greg Brockman replied, “ah nice.” And Nick Turley, OpenAI’s head of ChatGPT, called chatbots an “existential threat” to publishers.

How much any of this sways the case is anyone’s guess, but the documents confirm something the whole industry already knows: Content has value, even if it’s scraped “for free” from the internet. That value can vary with the type of content, how unique it is, and what the customer of that content wants to do with it. But nobody on either side of the courtroom thinks the number is zero.

That includes the government, by the way. The Department of Justice took the unusual step of filing a statement of interest in the Times case, addressing whether training AI models on publishers’ content counts as fair use, which would effectively give AI companies a pass. The DOJ said it does, shocking absolutely no one. The administration has been quite clear that it sees any concession on the copyright question as helping the Chinese win the AI race, since they won’t abide by the same rules. President Donald Trump’s own summary: “China’s not doing it.”

The money is in answers, not training

Here’s the thing, though: All of that is about training models. None of it settles AI search or, more broadly, inference. In fact, when talking specifically about what AI systems produce, the DOJ conceded that “an output reconstructing and disseminating an original copyrighted work may not be transformative,” referring to one of the pillar conditions of the fair-use doctrine. If you wanted a one-line description of AI search, a machine for summarizing current reporting, that comes pretty close.

The unsealed documents also go straight at another fair-use pillar, harm to the market for the original work. Turley wrote that OpenAI’s products “are largely substitutive, period,” and Microsoft CEO Satya Nadella testified that using chatbots “has substituted” for visiting the original sources.

The Times filed its lawsuit in December 2023, when the copyright conversation was largely about training. The latest back-and-forth shows the case is far from over, and that uncertainty explains why a market for training data has been so slow to form. Nobody wants to build a framework to pay for something that a court might declare free a few months from now. Training is where the lawsuits are. Inference is where the money is.

Inference is the other side of the coin: AI answers about what’s happening right now. The better an AI system handles those, the more valuable it is, so you’d expect AI companies to be lining up to pay for the quality content that makes their products better.

That mostly hasn’t happened. Despite a few licensing deals here and there, the AI industry has been mostly uninterested in building a marketplace or payment technology where they can buy content in real time for a fair price. Brian Morrissey at the Rebooting suggests this is because most content isn’t unique enough. Even if you have the world’s best enchilada recipe, the AI only needs one. Commoditization means a race to the bottom, and in this case the bottom sits pretty close to zero.

The middlemen got paid first

That theory holds up if publisher deals are all you look at. Zoom out and the picture changes. I’ve written about the class of data brokers, essentially AI “middlemen,” who scrape the internet at scale and then resell the data. Going by names like Exa, Parallel, and Tavily, their customers aren’t just AI companies, but a whole cadre of enterprise businesses, including ad agencies, investment firms, even other publishers. One estimate, cited in Matthew Scott Goldstein’s widely circulated report on the scraper economy, pegged the market at about $1 billion.

So a market for inference exists. The money just mostly goes around the media, and the problem’s getting worse. The bot-protection company DataDome put out a report that showed “bad” bot traffic grew 124% in a year, more than nine times faster than human traffic, meaning more bots are visiting sites and scraping data even when they’re told not to. Scraping alone jumped 185%.

Sponsored. Journalists, PR pros and communicators: the fall cohort of AI for Media starts October 13, six live Tuesday sessions with Pete Pachal plus two 1:1 coaching calls. Code AISEARCH500 takes $500 off the $1,500 price for anyone who found the course through AI search, a bigger discount than is offered anywhere else.

One detail in the report deserves a second look. Meta, which just released its personal AI agent, Muse, is responsible for 46.3% of all AI bot traffic DataDome tracked in the first half of the year, well ahead of OpenAI. Meta also has deals with several publishers, including USA Today, CNN, Fox News, People Inc., and others. So the biggest data harvester in AI is also a paying customer. It just picks the moments when it pays.

Steering inference money back toward publishers will take work, but there are early signs of it. Parallel has introduced a way to pay publishers for their contributions to agent tasks, with The Atlantic and Fortune among its first partners. Companies like Cloudflare, TollBit, and ProRata have all built payment rails for “good” bots to pay for what they take, and Cloudflare just shifted to paying publishers when their content shapes an AI answer, not just when it’s fetched.

Even Google is opening up payments to some publications when their content contributes significantly to answers in AI Overviews, AI Mode, and Gemini. It’s very early, but the idea that AI inference could work something like the YouTube Partner Program no longer sounds far-fetched.

  • Subscribe to our newsletter

    How AI is changing media, journalism, and content creation.

    Learn More

The buyer still names the price

So the market’s there, and so are the pipes for payment. Publishers have everything they need to get paid as data suppliers to AI systems, except an answer to one question: Who sets the price? The publisher, the exchange, or the buyer?

For now, the buyer does. If you’re looking to buy content, there are more than a dozen data brokers willing to sell it to you for the lowest possible price, and the exchanges are nascent, with wildly inconsistent pricing. Even Google’s program pays whatever Google decides; publishers in the pilot described the math to Digiday as “quite black box,” and one called the offers “lowball.”

Jonathan Roberts, People Inc.’s chief innovation officer, described the situation this summer as having 30 Napsters for content, but no Spotify. Today, you could argue there are a few baby Spotifys. But a reliable exchange alone didn’t get artists paid. In the case of the music industry, rights holders also had enough collective weight to insist on it. Publishers have the first part forming, but almost none of the second.

With the government clearly sitting this one out, that leverage has to come from publishers themselves. One route is acting collectively on licensing. SPUR, a coalition of publishers building standards for how AI systems license and track their content (the Associated Press recently joined), might be the beginnings of this.

The other route is controlling access, through better protections and agentic tools of their own. That would take a major shift, since DataDome’s report also shows 65.3% of the more than 21,000 popular websites it tested didn’t stop a single one of its test bots.

AI systems have gotten a lot better in recent months, but an answer can only be as good as the information behind it. Even Microsoft’s Nadella knows this, saying in the court documents, “anything that is paywalled should be licensed.” The market just hasn’t been told what that license should cost.

Setting the terms means showing up with leverage, whether that comes from banding together or from locking the doors. Until publishers have it, buyers will keep assuming the content is theirs for the taking. Right now, they’re not wrong.

A version of this column appears in Fast Company.

Contributors

  • Pete Pachal: Author

    Pete Pachal is the founder of The Media Copilot. In addition to producing the site’s newsletter and podcast, he also teaches courses on how journalists and communications professionals can apply AI tools to their work. Pete has a long career in journalism, previously holding senior roles in global newsrooms such as CoinDesk and Mashable. He’s appeared on Fox Business, CNN, and The Today Show as a thought leader in tech and AI. Pete also puts his encyclopedic knowledge of Doctor Who to good use on the popular podcast, Pull To Open.

Category: AI media analysisTags:AI summaries| Copyright| AI licensing
Share this post:
FacebookTweetLinkedInEmail
  • Related articles

Judge dismisses Chegg, Penske antitrust suits over Google AI Overviews

Read moreJudge dismisses Chegg, Penske antitrust suits over Google AI Overviews

Appeals court upholds Thomson Reuters AI copyright win

Read moreAppeals court upholds Thomson Reuters AI copyright win
Rob Kelly on The Media Copilot podcast: Who gets paid for AI licensing?

What 99 AI licensing deals reveal about the future of publishing

Read moreWhat 99 AI licensing deals reveal about the future of publishing

Axios says AI is breaking the link economy

Read moreAxios says AI is breaking the link economy

Unsealed OpenAI emails bolster NYT copyright case

Read moreUnsealed OpenAI emails bolster NYT copyright case

Microsoft researcher called AI training the ‘largest theft of labor in human history’

Read moreMicrosoft researcher called AI training the ‘largest theft of labor in human history’

The Media Copilot

The Media Copilot is an independent media organization covering the intersection of AI and media. Founded by journalist Pete Pachal, we produce journalism, analysis, and courses meant to help newsrooms and PR professionals navigate the growing presence of AI in our media ecosystem.

  • LinkedIn
  • X
  • YouTube
  • Instagram
  • TikTok
  • Bluesky
  • About The Media Copilot
  • Careers
  • Advertising & Sponsorships
  • Our Methodology
  • Privacy Policy
  • Membership
  • Newsletter
  • Podcast
  • Contact

© 2026 · All Rights Reserved · Powered by Springwire.ai · RSS