bots Archives - The Media Copilot https://mediacopilot.ai/tag/bots/ How AI is changing Media, journalism and content creation Wed, 05 Aug 2026 01:16:57 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.2 https://mediacopilot.ai/wp-content/uploads/2024/08/cropped-cropped-Media-Copilot-favicon-60x60.jpeg bots Archives - The Media Copilot https://mediacopilot.ai/tag/bots/ 32 32 Judge allows Reddit’s data scraping lawsuit against Perplexity to procee https://mediacopilot.ai/reddit-perplexity-data-scraping-lawsuit/ Mon, 03 Aug 2026 13:50:55 +0000 https://mediacopilot.ai/?p=9530 A Manhattan judge advanced most of Reddit's copyright and conspiracy claims against Perplexity and three data-scraping firms over AI training data.

The post Judge allows Reddit’s data scraping lawsuit against Perplexity to procee appeared first on The Media Copilot.

]]>

A Manhattan federal judge on Friday refused to dismiss the core of Reddit’s lawsuit against Perplexity AI, allowing the platform to pursue claims that the AI search startup illegally bypassed technical protections to scrape Reddit content without paying for a license.

The ruling doesn’t determine whether Perplexity broke the law. But it does keep alive a legal theory that could prove as important as copyright in the battle over AI training data: that circumventing a platform’s access controls can itself create liability.

U.S. District Judge Paul Engelmayer ruled that Reddit can continue pursuing claims that Perplexity and three data scraping companies unlawfully bypassed protections designed to limit automated access to Reddit content, according to Reuters. He also found that Reddit has standing to sue over the alleged misuse of posts created by its users.

The defendants include Lithuania-based Oxylabs, Russia-based AWMProxy and Texas-based SerpApi. Reddit alleges the companies extracted its content from billions of search results without permission and that Perplexity worked with at least one of them to obtain the data. Perplexity, which does not license Reddit’s content, denies the allegations.

While Engelmayer dismissed several secondary claims, he allowed Reddit’s central allegations—including conspiracy and unlawful circumvention—to move forward.

Perplexity argues Reddit is attempting to control access to publicly available webpages it doesn’t own, using security measures it didn’t create and on behalf of users who never authorized the lawsuit. The company says it will continue defending what it calls the open internet.

Reddit argues the issue isn’t whether its content is publicly visible but whether companies deliberately bypassed technical barriers to collect it at scale without permission or compensation.

SerpApi attorney Jeff Homrig echoed that view, arguing the company accesses public search results rather than Reddit itself and that publicly available information doesn’t become proprietary simply because a platform later decides to charge for access.

That distinction could have broad consequences across the AI industry.

Reddit has transformed its user-generated content into a licensing business, striking deals with Google and OpenAI while describing itself in court filings as the most frequently cited source in AI-generated answers. If companies can obtain the same data by scraping around access controls, the value of those licensing agreements is diminished.

The lawsuit is one of a growing number testing how AI companies acquire training data. Authors, record labels, news publishers and other content owners have all sued AI developers over the use of copyrighted material. Reddit is also pursuing a separate lawsuit against Anthropic in California state court over similar scraping allegations.

For publishers, platforms and AI companies, Friday’s ruling is significant because it points beyond copyright. A favorable ruling on Reddit’s circumvention claims could give content owners another legal lever to protect their data—and strengthen their hand in negotiating AI licensing agreements.

The case is still in its early stages. But by allowing the core scraping and circumvention claims to proceed, the Southern District of New York has signaled that the next major battleground over AI training data may not be copyright alone. It may be whether AI companies can legally get around the digital gates content owners have built.

The post Judge allows Reddit’s data scraping lawsuit against Perplexity to procee appeared first on The Media Copilot.

]]>
Time starts building ads for AI agents as bot traffic overtakes humans https://mediacopilot.ai/time-ads-ai-agents-markdown/ Fri, 31 Jul 2026 14:01:47 +0000 https://mediacopilot.ai/?p=9487 Time is placing brand FAQs inside markdown pages aimed at AI crawlers, with Ally Bank and the Project Management Institute among its first buyers.

The post Time starts building ads for AI agents as bot traffic overtakes humans appeared first on The Media Copilot.

]]>

Ally Bank and the Project Management Institute have become two of the first brands to buy an ad meant to be read by a machine, not a person. Time began serving ads to AI agents this month, formatting them as sponsored FAQs stuffed with brand messaging and dropping them into stripped-down copies of its pages, according to Digiday.

The move follows Time’s decision last month to convert all its webpages into markdown, text-only versions that strip out design and images, making them easier for AI systems to crawl and process. The publisher’s bet is that greater accessibility will boost its visibility. By placing ads within those markdown files, Time also hopes to monetize the growing volume of AI bot traffic while that strategy plays out.

To build the ads, Time is working with an AI ad tech platform called Mobian, which converts the pages and generates the agent ads from a brand brief. The output gets turned into a PDF for humans to approve, much like a standard branded content deal. Mobian then feeds the same FAQ questions to AI search engines and tracks visibility, favorability and accuracy over time.

Mobian co-founder and CEO Jonah Goodhart said the shift reflects a growing reality: publishers and brands increasingly need to optimize for AI systems as much as human audiences.

“Maybe it’s more important to influence the agent than even the human, because with a human you influence one person. When you influence ChatGPT, you’re influencing potentially all of ChatGPT,” he told Digiday.

Goodhart said roughly 15% of brands now run their own markdown pages for AI crawlers, a figure he expects to grow as companies adapt to what he describes as a two-track internet.

Time COO Mark Howard declined to disclose traffic figures but pointed to TollBit data showing the publisher receives more AI crawler requests than most of the roughly 7,000 sites in the company’s network. During major events such as the Time100 franchise, bot activity surges so dramatically that AI crawlers outnumber human visitors on most days. The trend mirrors Cloudflare’s finding that automated bots now account for more than half of all web traffic.

Time is positioning those AI visits as a new source of advertising revenue, selling one
“agent ad” per markdown page. The ads are part of a broader generative engine optimization offering that reflects publishers’ growing focus on AI discovery over traditional search traffic.

The approach comes with uncertainty. No major AI company has explained how its models handle ads embedded in markdown files—or whether they recognize them as ads at all. Rob Derow, a managing director at BCG X, told Digiday the lack of standards is the biggest risk. If AI companies ultimately treat the practice like cloaking—showing crawlers content different from what humans see—the pages could be devalued, much as Google penalized similar SEO tactics.

To reduce that risk, Time labels each placement as sponsored content and identifies the advertiser, despite no current requirement to do so. “We don’t know yet because this is brand new, and we believe that we are paving the first path forward here,” COO Mark Howard told Digiday.

The post Time starts building ads for AI agents as bot traffic overtakes humans appeared first on The Media Copilot.

]]>
Meta now drives most AI agent traffic while sending publishers few visitors https://mediacopilot.ai/meta-ai-agent-traffic-datadome-q2-2026/ Fri, 17 Jul 2026 13:12:18 +0000 https://mediacopilot.ai/?p=9082 Overhead view of a nighttime digital newsroom with journalists at monitors and a wall display showing a rapidly rising crawl counter beside a flatlined referral traffic graphDataDome logged 17.7 billion AI agent requests in Q2 2026, with Meta's crawlers generating most of the volume and almost no referral traffic back to sites.

The post Meta now drives most AI agent traffic while sending publishers few visitors appeared first on The Media Copilot.

]]>

Meta’s crawlers hit websites 9.1 billion times in the second quarter of 2026 and sent almost nobody back in return. That single figure, pulled from DataDome’s Q2 2026 AI Traffic Report, captures the widening split between the agents that consume publisher infrastructure and the ones that actually deliver readers.

DataDome’s network processed 17.7 billion AI agent requests between April and June, a 45% jump from Q1’s 12.2 billion. The report draws on 5 trillion signals analyzed daily across more than 400 enterprises. Since January, the network has logged over 30 billion AI agent requests in total, and the monthly curve kept climbing: 4.77 billion in April, 6.29 billion in May, 6.60 billion in June.

Meta drove most of that growth. Its two crawlers do different jobs. Meta-ExternalAgent reads websites to train AI models without sending traffic or compensation back to publishers. Meta-WebIndexer works more like Google’s crawler, indexing pages so Meta AI can answer real-time queries. In Q2, ExternalAgent grew 74% to 5.3 billion requests and WebIndexer grew 163% to 3.75 billion. In June, WebIndexer passed ExternalAgent in monthly volume for the first time, a sign Meta is investing in the answering side of AI as much as the training side.

Crawl volume and referral value are pulling apart. ChatGPT-User, the top agent in Q1, fetched pages 6% less often in Q2. Yet OpenAI’s chatbot still commands 80% to 88% of all AI-driven referral traffic and grew referrals 17% quarter over quarter. Among the rest, Claude referrals more than doubled to 876,000, Perplexity grew 37%, and Grok collapsed 74% to just 24,000 visits.

The distinction matters because publishers have spent the past two years arguing that AI companies use their content without returning referral traffic or other value in exchange. That tension runs through the AI scraping economy, and DataDome’s numbers put figures on it.

The report also flags a new signal worth watching: Model Context Protocol traffic. MCP, the connective layer between AI agents and external tools that Anthropic introduced in late 2024, went from negligible volume to peaks near 500,000 requests a day. Most requests come from AI agents taking inventory through calls such as initialize, tools/list and prompts/list. Rather than reading content, those requests reveal what an agent intends to do before it takes action.

For newsrooms and publishers, the practical message is that bot-or-not detection no longer cuts it. Meta-ExternalAgent, Meta-WebIndexer, ChatGPT-User and a chat session all demand different responses. DataDome recommends agent-level classification, MCP monitoring, and identity validation through standards like Web Bot Auth rather than trusting user-agent strings, which are easily spoofed. Any allowlist granting automatic access based on a trusted agent name is exposed.

Already, 54% of DataDome customers have adopted agent trust policies. The firm frames that as a leading indicator. As Q3 data arrives, the open question is whether Meta’s tilt toward real-time indexing holds, and whether publishers can tell the difference between an agent burning their bandwidth and one bringing them an audience.

The post Meta now drives most AI agent traffic while sending publishers few visitors appeared first on The Media Copilot.

]]>
Reuters and Time flip the script on AI bots with blocking whitelists https://mediacopilot.ai/reuters-time-block-ai-bots-whitelist/ Thu, 11 Jun 2026 01:05:41 +0000 https://mediacopilot.ai/?p=8345 Illustration of friendly robots passing through a glowing gate toward menacing red-eyed robotsTwo major publishers are blocking all AI bots by default and only letting approved crawlers through.

The post Reuters and Time flip the script on AI bots with blocking whitelists appeared first on The Media Copilot.

]]>

Reuters and Time are blocking all AI bots by default and only letting approved crawlers through—a whitelist approach that more publishers are adopting as the volume of unauthorized scraping grows.

As Digiday reports, both publishers moved to block AI bots last month, joining People Inc. and The Atlantic, which adopted similar strategies earlier this year and late last year respectively. The goal is simple: content costs money to produce, and AI companies have been taking it without paying.

“We saw that there was an imbalance between the value that publishers like Reuters provide and the value that Reuters receives in kind, and so instead we went from a default allow-all to a default disallow all,” said Josh London, head of Reuters Professional, which oversees the direct-to-consumer and direct-to-professional businesses. Reuters has since signed AI licensing agreements with Microsoft and Meta, according to the report.

The publishers aren’t relying on any single tool. Reuters uses robots.txt files, a method that is voluntary and non-binding, and one that many AI bots simply ignore. The approach is meant to create friction and signal that access requires negotiation. “If you want this, let’s have a conversation and then we can allow you to access,” said Alphonse Hardel, head of agency at Reuters, who leads the content licensing business.

Time allows roughly 70 bots on its site, ranging from AI lab crawlers and social platforms to its own operational systems. The volume of bot traffic has become significant enough that Time sees it as leverage for a future AI visibility product it’s developing for brand clients.

The economics are also shifting. Blocking bots cuts server costs: Hardel said the expense of the bot-blocking vendor can be nearly offset by the reduction in non-human traffic. At People Inc., the shift from a block list to an allow list meant going from blocking roughly 2,100 user agents to over 30,000, said Lindsay Van Kirk, the company’s SVP of innovation, speaking at an IAB Tech Lab event in May.

“Adding two full seconds of latency to the majority of scrapers when you implement a block-all-bots approach is a really good thing, even if they have to go through,” Van Kirk said. “Every scraper who has to pay a home proxy network in order to get access to the content is margin that you are taking out of their business.”

The IAB Tech Lab has published guidance on bot management, and the SPUR Coalition—a publisher group formed earlier this year with major news organizations—announced significant new membership as it works to create technical standards for AI licensing and content protection.

For Reuters, the change hasn’t reduced site traffic. After monitoring bot activity over an extended period, the company had enough data to identify which bots it could block without hurting revenue. The publisher maintains a public robots.txt file that lists approved bots, a benchmark that also supports enforcement discussions, said Phil Andraos, general manager of Reuters Digital.

“It’s not a set it and forget it approach,” London said. “The value of content is something that we ignore at our own peril, especially as AI scales.”

The post Reuters and Time flip the script on AI bots with blocking whitelists appeared first on The Media Copilot.

]]>
Cloudflare CEO: Bots have overtaken human traffic online https://mediacopilot.ai/bots-passed-human-traffic-online-cloudflare-ceo/ Fri, 05 Jun 2026 11:39:40 +0000 https://mediacopilot.ai/?p=8234 Aerial illustration of a busy highway interchange at night with AI-tagged colored carsFor the first time, bots account for more web traffic than humans, according to Cloudflare data.

The post Cloudflare CEO: Bots have overtaken human traffic online appeared first on The Media Copilot.

]]>

For the first time in the internet’s history, bots account for more web traffic than humans.

Cloudflare CEO Matthew Prince announced the milestone this week, according to Tom’s Hardware, noting that automated traffic has now eclipsed human-generated requests online, months ahead of even his own projections.

“Welp, that happened faster than I predicted,” Prince wrote on X. “Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots have now passed human traffic online for the first time in the Internet’s history.”

According to Cloudflare’s Radar data, bots represented roughly 57% of all HTTP requests as of late April 2026, with humans accounting for the remaining 43%. Bot traffic has held between 53% and 60% in the weeks since. Prince said the actual crossover occurred in the last few months, though the data is messy enough that pinning down an exact date is difficult.

The shift underscores how quickly AI agents have transformed web traffic patterns. Before the generative AI era, bot traffic sat at around 20% of all web activity, with Google’s web crawler serving as the largest single source. Now, AI agents performing tasks on behalf of users are generating requests at a scale that dwarfs human browsing behavior.

Prince illustrated the contrast at SXSW earlier this year: “If a human were doing a task—let’s say you were shopping for a digital camera—you might go to five websites. Your agent or the bot that’s doing that will often go to 1,000 times the number of sites that an actual human would visit. So it might go to 5,000 sites. And that’s real traffic, and that’s real load, which everyone is having to deal with and take into account.”

The reaction to Prince’s announcement was swift. Tech billionaire Elon Musk replied with a single “Wow” to the post.

The full picture is more nuanced. While bots now dominate HTML request traffic—reading pages, scraping content, indexing sites—humans still account for roughly 65% of total web activity when the metric expands to include app usage, video streaming, maps, and social media scrolling. Bots have overtaken humans in the specific act of navigating and reading the web, but not in the broader measure of people actually using the internet.

Cloudflare, which handles approximately one-fifth of all global web traffic, has been tracking the trend closely. The company’s 2026 Threat Intelligence Report also found that bots now account for 94% of all login attempts across its network, meaning only 6% of login attempts come from actual humans.

The crossing point Prince initially forecast for 2027 arrived in 2026. What once required a two-year runway happened in a matter of months.

The post Cloudflare CEO: Bots have overtaken human traffic online appeared first on The Media Copilot.

]]>
Alliance for Audited Media opens ethical AI certification to publishers https://mediacopilot.ai/aam-ethical-ai-certification-all-publishers/ Thu, 23 Apr 2026 14:06:59 +0000 https://mediacopilot.ai/?p=6112 Abstract blue illustration of a checkmark inside concentric circles, representing an ethical AI certification sealThe move signals a push to bring industry-wide accountability to AI adoption in newsrooms.

The post Alliance for Audited Media opens ethical AI certification to publishers appeared first on The Media Copilot.

]]>

The Alliance for Audited Media announced Wednesday it is making its Ethical AI Certification available to all publishers as part of AAM membership. The certification provides a structured framework for developing, implementing, and demonstrating responsible AI governance.

Publishers are navigating rapid AI adoption, emerging regulatory proposals, and rising audience expectations for disclosure and oversight. Recent research from the Local Media Association and Trusting News found that nearly 99 percent of news audiences expect human involvement when AI is used. The finding underscores a challenge for publishers integrating AI tools: audiences notice and adjust their trust based on how disclosure is handled.

AAM developed the certification with publisher input. It evaluates companies across eight key areas including transparency, governance, bias and fairness, and privacy. The program was recently updated to incorporate elements of the IAB’s AI Transparency and Disclosure Framework to strengthen industry alignment, according to the announcement.

Richard Murphy, AAM’s CEO, president and managing director, said trust remains a critical factor in how AI is implemented and used. “By expanding access to our certification, we’re helping publishers demonstrate and communicate responsible AI use to their subscribers, advertisers and partners,” he said.

During the certification process, companies receive feedback on their AI governance policies and oversight mechanisms. Publishers who complete the program can display AAM’s Ethical AI Certification seal on their websites and media kits, and receive a listing on AAM’s Assurance List.

Publishers can find more information at auditedmedia.com.

The post Alliance for Audited Media opens ethical AI certification to publishers appeared first on The Media Copilot.

]]>