A Manhattan federal judge on Friday refused to dismiss the core of Reddit’s lawsuit against Perplexity AI, allowing the platform to pursue claims that the AI search startup illegally bypassed technical protections to scrape Reddit content without paying for a license.
The ruling doesn’t determine whether Perplexity broke the law. But it does keep alive a legal theory that could prove as important as copyright in the battle over AI training data: that circumventing a platform’s access controls can itself create liability.
U.S. District Judge Paul Engelmayer ruled that Reddit can continue pursuing claims that Perplexity and three data scraping companies unlawfully bypassed protections designed to limit automated access to Reddit content, according to Reuters. He also found that Reddit has standing to sue over the alleged misuse of posts created by its users.
The defendants include Lithuania-based Oxylabs, Russia-based AWMProxy and Texas-based SerpApi. Reddit alleges the companies extracted its content from billions of search results without permission and that Perplexity worked with at least one of them to obtain the data. Perplexity, which does not license Reddit’s content, denies the allegations.
While Engelmayer dismissed several secondary claims, he allowed Reddit’s central allegations—including conspiracy and unlawful circumvention—to move forward.
Perplexity argues Reddit is attempting to control access to publicly available webpages it doesn’t own, using security measures it didn’t create and on behalf of users who never authorized the lawsuit. The company says it will continue defending what it calls the open internet.
Reddit argues the issue isn’t whether its content is publicly visible but whether companies deliberately bypassed technical barriers to collect it at scale without permission or compensation.
SerpApi attorney Jeff Homrig echoed that view, arguing the company accesses public search results rather than Reddit itself and that publicly available information doesn’t become proprietary simply because a platform later decides to charge for access.
That distinction could have broad consequences across the AI industry.
Reddit has transformed its user-generated content into a licensing business, striking deals with Google and OpenAI while describing itself in court filings as the most frequently cited source in AI-generated answers. If companies can obtain the same data by scraping around access controls, the value of those licensing agreements is diminished.
The lawsuit is one of a growing number testing how AI companies acquire training data. Authors, record labels, news publishers and other content owners have all sued AI developers over the use of copyrighted material. Reddit is also pursuing a separate lawsuit against Anthropic in California state court over similar scraping allegations.
For publishers, platforms and AI companies, Friday’s ruling is significant because it points beyond copyright. A favorable ruling on Reddit’s circumvention claims could give content owners another legal lever to protect their data—and strengthen their hand in negotiating AI licensing agreements.
The case is still in its early stages. But by allowing the core scraping and circumvention claims to proceed, the Southern District of New York has signaled that the next major battleground over AI training data may not be copyright alone. It may be whether AI companies can legally get around the digital gates content owners have built.







