Skip to main content
AI System Evidence: What the Perplexity Discovery Order Means for Your ProgrameDiscovery & Legal Holds
5 min readFor Compliance Officers

AI System Evidence: What the Perplexity Discovery Order Means for Your Program

Your company uses an AI tool to draft responses, summarize documents, or route customer inquiries. A lawsuit lands. Opposing counsel asks for logs of what your team asked the system and what it returned. Do you know where that data lives, how long you keep it, or what it costs to produce?

A recent Southern District of New York order in a copyright dispute involving Perplexity AI just mapped the discovery obligations for retrieval-augmented generation systems. The court ordered production of the RAG database itself and six months of user activity logs. If your organization deploys similar tools, this case shows exactly where your exposure lives.

What Happened

Britannica and Merriam-Webster sued Perplexity AI, alleging the company's answer engine copied their copyrighted content into its retrieval index and reproduced it in user-facing responses. The plaintiffs requested snapshots of Perplexity's RAG database and extended user activity logs to prove infringement. Perplexity had already produced baseline data in a parallel case but resisted the expanded request, citing cost and burden.

The court granted partial relief: one additional RAG snapshot and six months of UAL data, with cost-sharing capped at $6,000 per month.

Timeline

  • 2024: Dow Jones and New York Post sued Perplexity over RAG architecture.
  • August 2025: District Judge Failla denied Perplexity's motion to dismiss in Dow Jones.
  • Britannica filing: Perplexity produced baseline data matching Dow Jones volumes without dispute.
  • May 8, 2026: Britannica requested ten additional months of data.
  • This order: Court granted six months of UAL data and one RAG snapshot, applying Rule 26(b)(1) and Rule 26(b)(2)(C).

Which Controls Failed or Were Missing

1. No documented retention schedule for system-generated logs
Perplexity's UAL database captured four data points per query: user input, retrieval instructions, model prompts, and output. The volume ran approximately 20 terabytes per day. Nothing in the record showed a Records Control Schedule tied to business need or regulatory requirement. The system kept everything, making it all discoverable.

2. Cost estimation without supporting documentation
Perplexity informally estimated hosting costs at $12,000 to $14,000 monthly, then filed declarations citing hundreds of thousands. The court noted an "orders of magnitude difference" and viewed the higher figure "with considerable skepticism." The gap likely reflected extraction costs from deep storage, but Perplexity never explained the breakdown on the record. The court capped relief at the lower figure Britannica had already offered.

3. No pre-litigation discovery protocol
Perplexity fought proportionality after the request landed. In a parallel case cited in the order, parties who negotiated a cost-sharing framework before motion practice avoided this fight entirely. Perplexity proposed cost-sharing only after Britannica requested the additional data, too late to control the scope or cost allocation.

4. Unclear data classification for AI-generated records
The RAG database and UAL logs became evidence of the system's core function. Your organization likely treats chatbot logs, model outputs, or retrieval indexes as ephemeral system data. This case shows they're business records the moment they document a transaction, decision, or customer interaction.

What the Relevant Standard Requires

Rule 26(b)(1) defines discoverable information as relevant to any party's claim or defense and proportional to the needs of the case. Proportionality weighs the importance of the issues, the amount in controversy, the parties' resources, and the burden of production.

Rule 26(b)(2)(C) allows a court to limit discovery if the burden outweighs its likely benefit, considering whether the information is obtainable from a more convenient source or whether the producing party had a reasonable opportunity to obtain it during discovery.

Rule 26(c)(1)(B) permits cost-shifting when good cause exists, typically under the Zubulake framework: accessibility of the data, likelihood of finding relevant information, availability from other sources, and the producing party's resources.

The court applied the registration-date rule from the Dow Jones precedent: the technical discovery window tracks when copyrights became effective, not when suit was filed. For Britannica, most registrations became effective between April and December 2025, setting the outer boundary for relevant data.

Lessons and Action Items for Your Team

Map your AI system's data flows now
Document what your RAG, chatbot, or summarization tool stores: user inputs, retrieval queries, model prompts, outputs, and metadata. Identify where each dataset lives, how long it persists, and whether it moves to cheaper storage over time. If you can't answer these questions before a legal hold lands, you can't estimate production costs or negotiate scope.

Write a retention schedule for AI-generated logs
Apply your Records Control Schedule to system-generated data. If your UAL logs document customer service interactions, they're business records under the same retention rule as call transcripts. If they're diagnostic data with no business value after 90 days, declare that and dispose on schedule. The failure here wasn't keeping the data; it was keeping it without a documented reason.

Document cost assumptions before you cite them
Perplexity's informal estimate became the ceiling when it couldn't explain why the formal number moved. If you tell opposing counsel a production will cost $10,000, and your vendor later quotes $80,000, put the reason in writing: extraction from tape, format conversion, hosting infrastructure, whatever changed. Courts assume the worst explanation when you don't provide one.

Negotiate discovery protocols before you fight
The order cites a parallel case where parties agreed to cost-sharing terms in advance and avoided motion practice. If your organization uses AI tools that generate high-volume logs, propose a sampling methodology, cost-sharing framework, or phased production schedule during Rule 26(f) conferencing. You'll get better terms negotiating than litigating.

Classify AI outputs as records when they document decisions
If your model drafts a contract clause, recommends a product, or routes a support ticket, the output is a business record the moment it influences a transaction. Your General Records Schedule likely covers "correspondence" or "transactional records" already. Extend those categories to AI-generated content that serves the same function, and apply the same retention and legal hold obligations.

The Perplexity order won't be the last time a court treats an AI system's database as the evidence itself. Prepare your program for that discovery now, before the subpoena arrives.

Rule 26(b)(1) text

You Might Also Like