Reddit sues Perplexity for data mining: Key facts about the case

  • Reddit has filed a lawsuit in New York against Perplexity and three companies for allegedly extracting unauthorized data.
  • Perplexity denies the accusations and defends fair access to public knowledge; SerpApi and Oxylabs also reject the charges.
  • The platform already licenses content from Google and OpenAI; it submitted a pre-trial notice and cites a 40x increase in referrals to Reddit.
  • The case, which concerns the Lithuanian company Oxylabs, has touched Europe and rekindled the debate on scraping and rights within the EU regulatory framework.

Reddit sues Perplexity for data mining

The San Francisco-based social network has filed a federal lawsuit in New York against Perplexity AI and several firms linked to web data harvesting, alleging that they have obtained Reddit content without permission to feed AI-based tools.

According to the document, Perplexity would not have a license to use the platform's material, while Reddit has reached agreements with other technology companies such as Google and OpenAI; in addition, after a cease and desist request Submitted last year, the company claims that Reddit mentions in Perplexity's system increased forty-fold.

What is reported

Reddit claims that various scraping services would have circumvented anti-extraction measures from the platform and collected publications through Google search results, describing the practice as a “data laundering economy” on an industrial scale.

The lawsuit details that Perplexity would have used at least one of these providers to obtain Reddit content, instead of subscribe to a license with the platform itself, and that the extractors would have masked identities and locations to circumvent controls.

Who are those involved?

In addition to Perplexity, the litigation points to Oxylabs UAB (Lithuania), to the AWMProxy domain (which Reddit describes as linked to a former Russian botnet) and to the startup SerpApi (Texas), which places the case on a map that mixes actors from the United States and Europe.

The response of the defendant companies

Perplexity has stated that it has not yet been formally notified and that it will vigorously defend users' right to access freely and fairly to public knowledge, highlighting that its approach aims to provide accurate answers with AI in a responsible manner.

From SerpApi, a spokesperson has completely rejected the accusations and has advanced that the company will defend itself vigorously in court; Oxylabs, for its part, expressed surprise and disappointment, asserting that it had received no prior contact from Reddit and defending its collection of public data.

Regarding AWMProxy, the platform indicates that it has not been possible collect comments Of the entity.

Background and license agreements

This legal step adds to another front opened by Reddit: in June it filed a similar lawsuit against the AI ​​company Anthropic, a procedure that remains in progress after being transferred to a federal court.

Reddit emphasizes that its community, made up of thousands of subreddits and more than 100 million daily users, is a key source of internet conversations, which is why it has signed licenses with Google, OpenAI and other firms for model training.

On the stock market, after learning of the legal action, Reddit shares closed the session with a drop of more than 4% in New York, reflecting the market's sensitivity to data disputes in the AI ​​sector.

Implications for Europe and Spain

The presence of EU-based Oxylabs introduces a European angle to the controversy and puts the debate on the use of public data, scraping and the limits of copyright under EU law.

Beyond the US litigation, European players – including publishers, platforms and developers – remain closely watching how the balance is access to information publicly available with rights protection and terms of use, in a context marked by the Copyright Directive and the emerging regulatory framework for AI.

What Reddit is asking for and next steps

The company requests a financial compensation unspecified and an injunction preventing Perplexity from using Reddit data, pending a court ruling on whether any rights were violated and the scope of any injunctions.

The procedural times and the fit of the defenses remain to be defined, but everything points to this case will set a precedent in a field where public interest in information, intellectual property, and the training needs of AI systems collide.

The battle between platforms with large repositories of human conversation and artificial intelligence companies is intensifying: with licensing on one side and scraping allegations on the other, the dispute between Reddit and Perplexity illustrates the new board where the value, permissions and limits of online data are negotiated.

How to buy from ChatGPT-2
Related article:
Buying from ChatGPT: A Complete Guide to Leveraging AI in Your Online Shopping

Add as preferred source in Google