Skip to main content
Category: E-Discovery and Legal Holds

Predictive Coding

Simply put

Predictive coding is a computer-assisted approach used in electronic discovery to help identify which documents in a large collection are likely relevant to a legal matter. Rather than having people review every document, a small set of documents reviewed by experts is used to train software to rank or classify the remaining documents. Note that the same term is also used, unrelatedly, in neuroscience and signal processing, but that sense does not apply here.

Formal definition

In the e-discovery context, predictive coding refers to a form of technology-assisted review (TAR) in which machine-learning classifiers are trained on human coding decisions applied to a sample of documents, then used to predict relevance (or other coding categories) across a larger review population. It is commonly implemented through workflows such as continuous active learning (sometimes described as CAL or TAR 2.0), among other protocols, and is typically evaluated using measures such as recall and precision to demonstrate defensibility. The evidence packet supplied here does not contain e-discovery-specific sources; the available sources describe unrelated meanings of the term in neuroscience/cognitive science and in data compression, so the practitioner-level detail above (including any reference to seminal definitional work or case-law acceptance) cannot be substantiated from the provided evidence and should be verified against authoritative e-discovery sources before publication.

Why it matters

Predictive coding matters to information governance and legal teams because the volume of electronically stored information involved in litigation, investigations, and regulatory requests has grown well beyond what manual, document-by-document review can address in a timely or cost-effective way. By allowing a smaller set of expert coding decisions to guide the classification of a much larger population, predictive coding can materially reduce the review burden while providing a structured, measurable basis for demonstrating the thoroughness of an approach. This connects e-discovery practice to broader recordkeeping concerns, since the ability to locate, assess, and account for relevant records depends on how well an organization has captured, classified, and retained its information in the first place.

Defensibility is central to why the technique is used with care. Because predictive coding relies on statistical prediction rather than exhaustive human review, parties are often expected to explain and, where challenged, justify the process they followed. Metrics such as recall and precision are commonly used to characterize how completely relevant material was identified and how much of the retrieved set was actually relevant. Whether and how a particular predictive coding workflow is accepted can depend heavily on jurisdiction, court, sector, and the expectations of the parties involved, so requirements should not be treated as uniform.

The evidence packet supplied for this entry does not contain e-discovery-specific sources; it covers unrelated senses of the term in neuroscience and data compression. Accordingly, specific claims about foundational definitional work, particular protocols, or case-law acceptance in the e-discovery context could not be substantiated here and should be verified against authoritative e-discovery references before they are relied upon for practice.

Who it's relevant to

Litigation and e-discovery teams
Legal professionals managing document review use predictive coding to reduce the volume requiring manual review and to prioritize likely-relevant material. They are typically responsible for selecting a workflow, overseeing training decisions, and being prepared to explain or defend the process if it is challenged.
Information governance officers
Governance leads are concerned with how predictive coding intersects with the wider accountability framework, including how well records were captured, classified, and retained, since the quality of those upstream practices affects the reliability of any review conducted over the resulting collection.
Records managers
Records managers may be drawn into e-discovery when identifying, preserving, and producing records subject to a matter. Sound classification and retention practices support more predictable review, though predictive coding itself operates on collected documents rather than replacing records controls.
Compliance and legal hold coordinators
Those responsible for legal holds and regulatory responses need to understand how predictive coding fits within preservation and production obligations, and that acceptance of such methods can vary by jurisdiction, court, and sector, so no single approach should be assumed to be universally accepted.

Inside Predictive Coding

Technology-Assisted Review (TAR)
The broader category of e-discovery workflows in which predictive coding operates. Predictive coding is the machine-learning component within TAR, used to classify documents as relevant or non-relevant to a matter. TAR encompasses the end-to-end process, including human review decisions that train and validate the system.
Supervised machine-learning model
The underlying mechanism by which predictive coding operates. Reviewers make relevance determinations on a set of documents, and the system generalizes from those coded examples to rank or classify the remaining collection. The model's performance depends on the quality and consistency of human coding decisions.
Seed set and training
The initial set of documents coded by subject-matter experts or reviewers to teach the model. Depending on the workflow, training may occur once or iteratively as additional documents are reviewed and fed back into the model.
Continuous Active Learning (CAL / TAR 2.0)
A widely discussed workflow variant in which the system continually re-ranks the collection and surfaces the documents most likely to be relevant for ongoing review, updating the model as reviewers proceed rather than relying on a single fixed training round. This is often contrasted with earlier one-time-training approaches sometimes labelled TAR 1.0.
Validation and quality measures
Metrics commonly used to assess model performance, such as recall (the proportion of relevant documents identified) and precision (the proportion of retrieved documents that are actually relevant). Sampling and statistical estimation are typically used to demonstrate the defensibility of the review to opposing parties or courts.
Defensibility and judicial acceptance
The context in which predictive coding is used to reduce the cost and time of large-scale document review while producing an outcome that can be defended. Commentary attributed to Grossman and Cormack is frequently cited in defining the term, and predictive coding has received acceptance in case law in various jurisdictions, though the extent and conditions of acceptance depend on the jurisdiction and the specific matter.

Common questions

Answers to the questions practitioners most commonly ask about Predictive Coding.

Is predictive coding the same as the predictive coding or predictive processing concept from neuroscience?
No. In the records management and e-discovery context, predictive coding refers to a technology-assisted review (TAR) method used to prioritise or classify documents for relevance during discovery. It is unrelated to the neuroscience theory of predictive processing, and that term should not be treated as an alias here. The overlap is only in the words, not the meaning.
Does predictive coding fully automate document review and remove the need for human reviewers?
No. Predictive coding relies on human input to function. Subject-matter experts or reviewers code a set of documents, and the system uses that coding to model relevance across a larger population. Human judgment remains central to training, quality control, and validation, so the technology typically supports rather than replaces reviewers.
How does a predictive coding workflow typically begin?
Workflows commonly begin with reviewers coding documents to train the system. Some approaches use a defined seed or training set reviewed by subject-matter experts, while continuous active learning (often described as CAL or TAR 2.0) instead feeds reviewer decisions back into the model on an ongoing basis, allowing the ranking of documents to update as review proceeds. The specific approach depends on the tool and the matter.
How is the effectiveness of a predictive coding process usually validated?
Effectiveness is typically assessed using measures such as recall and precision, often applied to a statistically drawn sample of the document population. Validation aims to demonstrate that the process identified a defensible proportion of relevant material. The acceptable thresholds and sampling methods depend on the matter, the jurisdiction, and any agreements between the parties or directions from the court.
What documentation is important for defensibility when using predictive coding?
Maintaining a clear record of the methodology is generally advisable, including how training or coding decisions were made, how the model was validated, and what quality-control steps were taken. Contemporaneous documentation supports the ability to explain and defend the process if challenged. Depending on the jurisdiction and the parties' agreements, some level of transparency about the workflow may be expected or negotiated.
Does the use of predictive coding require cooperation or disclosure between opposing parties?
Practice varies. In some jurisdictions and matters, courts and parties have addressed the transparency of the TAR process, including whether and how the methodology or aspects of the training are disclosed. Whether cooperation is required or advisable depends on the applicable procedural rules, case law, and any protocol agreed between the parties, so requirements should be confirmed for the relevant jurisdiction.

Common misconceptions

Predictive coding removes the need for human reviewers.
Predictive coding depends on human relevance decisions to train and validate the model. Human judgement, particularly from those familiar with the matter, typically drives coding decisions and quality control; the technology assists and scales that judgement rather than replacing it.
Predictive coding is the same thing as Technology-Assisted Review.
Predictive coding is more precisely the machine-learning classification component within TAR. TAR is the broader workflow, which includes training approaches such as one-time training and Continuous Active Learning, along with human review and validation steps. The terms are often used loosely but are not strictly synonymous.
Once a court has accepted predictive coding, it is automatically acceptable everywhere.
Judicial acceptance has developed in case law in various jurisdictions, but whether and how predictive coding may be used, and what validation is expected, depends on the jurisdiction, the court, and the circumstances of the particular matter. Acceptance in one context should not be assumed to apply universally.

Best practices

Document the workflow choice (for example, a single-training approach versus Continuous Active Learning) and the reasoning behind it, so the process can be explained and defended if challenged.
Use consistent, well-qualified reviewers or subject-matter experts to make the coding decisions that train the model, since model quality depends heavily on the consistency of human input.
Apply validation using recognized measures such as recall and precision, supported by appropriate sampling, to demonstrate the completeness and reliability of the review.
Engage with opposing parties and, where relevant, the court early about the intended use of predictive coding, as transparency about methodology often supports defensibility.
Confirm the expectations and any applicable requirements in the relevant jurisdiction before adopting predictive coding, since acceptance and validation standards vary by jurisdiction and matter.
Retain records of training rounds, coding decisions, model versions, and validation results to preserve an auditable account of how the review was conducted.