Predictive Coding
Predictive coding is a computer-assisted approach used in electronic discovery to help identify which documents in a large collection are likely relevant to a legal matter. Rather than having people review every document, a small set of documents reviewed by experts is used to train software to rank or classify the remaining documents. Note that the same term is also used, unrelatedly, in neuroscience and signal processing, but that sense does not apply here.
In the e-discovery context, predictive coding refers to a form of technology-assisted review (TAR) in which machine-learning classifiers are trained on human coding decisions applied to a sample of documents, then used to predict relevance (or other coding categories) across a larger review population. It is commonly implemented through workflows such as continuous active learning (sometimes described as CAL or TAR 2.0), among other protocols, and is typically evaluated using measures such as recall and precision to demonstrate defensibility. The evidence packet supplied here does not contain e-discovery-specific sources; the available sources describe unrelated meanings of the term in neuroscience/cognitive science and in data compression, so the practitioner-level detail above (including any reference to seminal definitional work or case-law acceptance) cannot be substantiated from the provided evidence and should be verified against authoritative e-discovery sources before publication.
Why it matters
Predictive coding matters to information governance and legal teams because the volume of electronically stored information involved in litigation, investigations, and regulatory requests has grown well beyond what manual, document-by-document review can address in a timely or cost-effective way. By allowing a smaller set of expert coding decisions to guide the classification of a much larger population, predictive coding can materially reduce the review burden while providing a structured, measurable basis for demonstrating the thoroughness of an approach. This connects e-discovery practice to broader recordkeeping concerns, since the ability to locate, assess, and account for relevant records depends on how well an organization has captured, classified, and retained its information in the first place.
Defensibility is central to why the technique is used with care. Because predictive coding relies on statistical prediction rather than exhaustive human review, parties are often expected to explain and, where challenged, justify the process they followed. Metrics such as recall and precision are commonly used to characterize how completely relevant material was identified and how much of the retrieved set was actually relevant. Whether and how a particular predictive coding workflow is accepted can depend heavily on jurisdiction, court, sector, and the expectations of the parties involved, so requirements should not be treated as uniform.
The evidence packet supplied for this entry does not contain e-discovery-specific sources; it covers unrelated senses of the term in neuroscience and data compression. Accordingly, specific claims about foundational definitional work, particular protocols, or case-law acceptance in the e-discovery context could not be substantiated here and should be verified against authoritative e-discovery references before they are relied upon for practice.
Who it's relevant to
Inside Predictive Coding
Common questions
Answers to the questions practitioners most commonly ask about Predictive Coding.