Skip to main content
The state of ai impact assessment
Should You Filter Collaboration Data Before Collection?Audit & Assessment
5 min readFor eDiscovery Specialists

Should You Filter Collaboration Data Before Collection?

The question at hand

Your legal team receives a preservation notice. The scope includes Slack, Teams, email, and shared drives for twelve custodians over eighteen months. You're facing terabytes of potential data. Do you collect everything first and filter later, or apply filters upfront to reduce what you preserve?

This decision impacts your storage costs, review timelines, and exposure to spoliation claims. Legal teams are divided on this issue, with valid arguments on both sides based on different risk tolerances.

The tension arises from collaboration platforms. These systems generate massive volumes of content, much of it irrelevant to any given matter. Automated notifications, emoji reactions, duplicate files, and unrelated conversations all get swept up in broad collection scopes. This is what practitioners call noise: data that gets collected or reviewed even though it has little relevance to the legal matter at hand.

The case for filtering early

Proponents of early filtering argue that you can't afford to preserve everything. The data volumes are simply too large, and most of what you'd collect has no evidentiary value.

Consider what noise includes: duplicate files across multiple shared folders, outdated document versions, automated system notifications, conversations outside the relevant time period, and threads where custodians were copied but never engaged. Collecting all of this inflates your storage costs and buries relevant evidence under layers of irrelevant content.

Advanced filtering tools let you apply defensible criteria before preservation. You can exclude irrelevant file types (system logs, temporary files), deduplicate at the collection stage, and apply date ranges that match your matter's scope. Some platforms allow filtering by participant, thread activity, or content patterns.

The efficiency gains are significant. Reducing your dataset by 60% through defensible filters cuts storage costs, speeds up your Early Case Assessment timeline, and provides reviewers with a cleaner dataset. Your legal team can make informed decisions faster because they're not wading through noise.

The risk argument is also important. Every piece of data you collect becomes part of your preservation obligation. If you collect broadly and then fail to maintain that data properly, you've created spoliation exposure. Narrower collection means narrower obligations.

The case for collecting first

The opposing view is equally compelling: you don't know what you don't know. Filtering before collection creates irreversible decisions when you have the least information about your case.

The spoliation risk cuts both ways. If you apply filters that exclude potentially relevant data, you can't get it back. Courts don't look kindly on preservation decisions that prioritize cost savings over completeness. A document that seemed irrelevant at the outset might be crucial six months into discovery.

Collaboration platforms complicate this further. A Slack thread that looks like casual banter might contain the only contemporaneous discussion of a key decision. An email where your custodian was CC'd but didn't reply might still show notice of critical information. Automated notifications can establish timelines. You won't know any of this until you've reviewed the content.

There's also a practical reality: your filters are only as good as your understanding of the platform's data structures. If you don't know how Teams organizes threaded conversations or how Slack handles edited messages, your filters might exclude relevant content by accident. Collection tools capture metadata and relationships that become important later, even if the content itself seems minor.

The conservative approach is to collect broadly and filter during review. Modern review platforms can handle the volume, and you can apply more sophisticated filters once you understand your case better. Storage is cheaper than sanctions.

Where practitioners actually land

Most eDiscovery specialists use a hybrid approach that shifts based on case specifics.

For matters with clear scope and well-defined custodians, early filtering makes sense. If you're responding to an employment dispute with a three-month window and five participants, you can apply date ranges and participant filters with confidence. The risk of excluding relevant data is low, and the efficiency gains are substantial.

For complex litigation with evolving theories or unclear boundaries, practitioners collect more broadly. They'll still apply obvious filters (file types, system files, duplicates), but they preserve the full conversation threads and metadata. The filtering happens during Early Case Assessment when the legal team has enough context to make informed decisions.

The tool matters too. Platforms that allow you to store data centrally and apply filters retroactively give you more flexibility. You can collect once, then experiment with different filter criteria as your understanding of the case develops. This reduces the pressure to make perfect filtering decisions at the outset.

Our take

Filter what you can defend, but don't optimize yourself into spoliation.

The practical middle ground is to apply conservative filters at collection (date ranges that bracket the relevant period, obvious file-type exclusions, deduplication) and save the aggressive filtering for Early Case Assessment. At that stage, you have the data preserved, you understand the case theory, and you can make informed decisions about what actually matters.

Your Records Freeze procedures should document your filtering criteria and the rationale behind them. If you exclude certain platforms, file types, or date ranges, write down why. If you later need to defend your preservation decisions, contemporaneous documentation of your reasoning is your best protection.

The collaboration-platform problem isn't going away. Data volumes will keep growing, and noise will remain a challenge. The solution isn't to avoid filtering entirely or to filter so aggressively that you create gaps in your preservation. It's to build processes that let you make defensible decisions at each stage, with the flexibility to adjust as you learn more about your matter.

Start with broad preservation, apply conservative filters, and use Early Case Assessment to refine your scope. That approach balances efficiency with defensibility, and it gives your legal team room to make informed decisions without committing to irreversible choices before they understand the case.

eDiscovery

Promotional banner for the Pentest Readiness checklist download

You Might Also Like