• Skip to main content
  • Skip to header right navigation
  • Skip to site footer
The Media Copilot

The Media Copilot

How AI is changing Media, journalism and content creation

  • News
  • Reviews
  • Guides
  • AI Courses
    • AI Quick Start
    • NEW—AI for Media
    • Custom AI Training for Teams
  • Newsletter
  • Podcast
  • Events
    • GEO Dinner Series
    • Webinars
  • About

Judge allows Reddit’s data scraping lawsuit against Perplexity to procee

A Manhattan judge advanced most of Reddit’s copyright and conspiracy claims against Perplexity and three data-scraping firms over AI training data.

(Credit: ChatGPT)
Aug 3, 2026

By The Copilot

A Manhattan federal judge on Friday refused to dismiss the core of Reddit’s lawsuit against Perplexity AI, allowing the platform to pursue claims that the AI search startup illegally bypassed technical protections to scrape Reddit content without paying for a license.

The ruling doesn’t determine whether Perplexity broke the law. But it does keep alive a legal theory that could prove as important as copyright in the battle over AI training data: that circumventing a platform’s access controls can itself create liability.

U.S. District Judge Paul Engelmayer ruled that Reddit can continue pursuing claims that Perplexity and three data scraping companies unlawfully bypassed protections designed to limit automated access to Reddit content, according to Reuters. He also found that Reddit has standing to sue over the alleged misuse of posts created by its users.

The defendants include Lithuania-based Oxylabs, Russia-based AWMProxy and Texas-based SerpApi. Reddit alleges the companies extracted its content from billions of search results without permission and that Perplexity worked with at least one of them to obtain the data. Perplexity, which does not license Reddit’s content, denies the allegations.

While Engelmayer dismissed several secondary claims, he allowed Reddit’s central allegations—including conspiracy and unlawful circumvention—to move forward.

Perplexity argues Reddit is attempting to control access to publicly available webpages it doesn’t own, using security measures it didn’t create and on behalf of users who never authorized the lawsuit. The company says it will continue defending what it calls the open internet.

Reddit argues the issue isn’t whether its content is publicly visible but whether companies deliberately bypassed technical barriers to collect it at scale without permission or compensation.

SerpApi attorney Jeff Homrig echoed that view, arguing the company accesses public search results rather than Reddit itself and that publicly available information doesn’t become proprietary simply because a platform later decides to charge for access.

That distinction could have broad consequences across the AI industry.

Reddit has transformed its user-generated content into a licensing business, striking deals with Google and OpenAI while describing itself in court filings as the most frequently cited source in AI-generated answers. If companies can obtain the same data by scraping around access controls, the value of those licensing agreements is diminished.

The lawsuit is one of a growing number testing how AI companies acquire training data. Authors, record labels, news publishers and other content owners have all sued AI developers over the use of copyrighted material. Reddit is also pursuing a separate lawsuit against Anthropic in California state court over similar scraping allegations.

For publishers, platforms and AI companies, Friday’s ruling is significant because it points beyond copyright. A favorable ruling on Reddit’s circumvention claims could give content owners another legal lever to protect their data—and strengthen their hand in negotiating AI licensing agreements.

The case is still in its early stages. But by allowing the core scraping and circumvention claims to proceed, the Southern District of New York has signaled that the next major battleground over AI training data may not be copyright alone. It may be whether AI companies can legally get around the digital gates content owners have built.

Posts co-authored by The Copilot are drafted with AI and then carefully edited by Media Copilot editors. Our AI-assisted process allows us to bring more valuable content to our readers while preserving accuracy and quality.

Contributors

  • The Copilot: Author

    I'm a generative AI writer for The Media Copilot. I help author posts, and with the help of human editors, play a growing role in the site's content strategy.

  • Romy Abu-Fadel: Editor

    Romy Abu-Fadel is a journalist, researcher, and 2026 graduate of Georgetown University's Edmund A. Walsh School of Foreign Service. She covers artificial intelligence and its impacts on the media industry.

Category: NewsTags:Perplexity| licensing| bots| Copyright| Lawsuit
Share this post:
FacebookTweetLinkedInEmail

What do 1,000 journalists and PR pros know about AI that you don't? They took AI Quick Start, a 1-hour live class from The Media Copilot. 94% satisfaction. Find out how to work smarter with AI in just 60 minutes. Get 20% off with the code AIPRO: https://mediacopilot.ai/

  • Related articles

Time starts building ads for AI agents as bot traffic overtakes humans

Read moreTime starts building ads for AI agents as bot traffic overtakes humans

ChatGPT blocks author-style requests as publishing safeguards tighten

Read moreChatGPT blocks author-style requests as publishing safeguards tighten

Meta licenses Newsmax journalism for AI search across Facebook and Instagram

Read moreMeta licenses Newsmax journalism for AI search across Facebook and Instagram

Muck Rack licenses MIT Technology Review for expanded AI news monitoring

Read moreMuck Rack licenses MIT Technology Review for expanded AI news monitoring

Citing trusted news brands increases confidence in AI responses, UK Ipsos survey finds

Read moreCiting trusted news brands increases confidence in AI responses, UK Ipsos survey finds
The Delhi High Court's colonial-era sandstone facade glows warm at dusk, with a security guard standing near the entrance gate.

Delhi High Court rules OpenAI’s training on ANI news content is fair dealing

Read moreDelhi High Court rules OpenAI’s training on ANI news content is fair dealing

The Media Copilot

The Media Copilot is an independent media organization covering the intersection of AI and media. Founded by journalist Pete Pachal, we produce journalism, analysis, and courses meant to help newsrooms and PR professionals navigate the growing presence of AI in our media ecosystem.

  • LinkedIn
  • X
  • YouTube
  • Instagram
  • TikTok
  • Bluesky
  • About The Media Copilot
  • Advertising & Sponsorships
  • Our Methodology
  • Privacy Policy
  • Membership
  • Newsletter
  • Podcast
  • Contact

© 2026 · All Rights Reserved · Powered by Springwire.ai · RSS