Skip to main content
Category: E-Discovery and Legal Holds

Technology-Assisted Review

Also known as: TAR, Predictive Coding
Simply put

Technology-Assisted Review (TAR) is a process in which computer software helps sort through large volumes of documents by learning from decisions made by expert human reviewers. Instead of people reading every document, the software classifies documents so that reviewers can focus on a smaller, more relevant set. It is commonly used during eDiscovery, the identification and review of documents in the context of legal matters.

Formal definition

Technology-Assisted Review (TAR), also referred to as predictive coding, is a process that combines human legal expertise with software algorithms to classify documents in large data sets, typically during eDiscovery. Expert reviewers provide input, often by coding a subset of documents, from which the software electronically classifies the remaining documents to identify those likely to be relevant. The general purpose is to narrow the population requiring human review to a smaller, more relevant set, thereby supporting more efficient and focused review. The specific workflows, protocols, and conditions under which TAR is applied depend on the matter, applicable practice guidelines, and jurisdiction; the term describes a review methodology within eDiscovery and should not be equated with records management classification or retention processes.

Why it matters

The volume of electronically stored information involved in litigation, regulatory inquiries, and internal investigations has grown to a scale where document-by-document human review is often impractical within the time and cost constraints of a matter. Technology-Assisted Review addresses this pressure by narrowing the population of documents requiring human attention to a smaller, more relevant set, which can support a more efficient and focused review. For organizations facing large-scale discovery obligations, the availability of a defensible, more scalable review methodology is a significant practical consideration.

TAR also matters because its acceptability is shaped by practice guidelines, protocols, and the expectations of courts and opposing parties, which vary by matter and jurisdiction. Bodies within the eDiscovery community have published best-practice guidance intended to help parties assess whether and under what conditions TAR should be applied. The methodology's credibility depends heavily on documented process, expert input, and transparency, rather than on the technology alone, so understanding TAR is relevant to anyone responsible for demonstrating that a review was reasonable and defensible.

It is important to keep TAR's scope clear. TAR is a review methodology used within eDiscovery to classify documents by likely relevance to a legal matter. It should not be equated with records management classification, retention scheduling, or disposition processes, which serve different purposes across the records lifecycle. Professionals should be careful not to assume that a tool or workflow validated for one context carries over to the other.

Who it's relevant to

Litigation and eDiscovery teams
Legal professionals managing document identification and review in litigation or regulatory matters use TAR to focus reviewer effort on a smaller, more relevant set of documents. They are typically responsible for selecting and documenting an appropriate protocol and demonstrating that the review process was reasonable, which depends on the matter and applicable practice guidelines.
Records and information governance professionals
Those responsible for information governance should understand where TAR fits and where it does not. TAR is a review methodology within eDiscovery for classifying documents by likely relevance to a matter; it is distinct from records management classification, retention, and disposition. Recognizing this boundary helps avoid conflating discovery review with lifecycle recordkeeping controls.
Compliance and investigations staff
Teams conducting internal investigations or responding to regulatory inquiries may encounter TAR as a means of managing large data sets efficiently. Its appropriateness and the conditions of its use depend on the matter, applicable guidance, and jurisdiction, so these staff benefit from understanding the role of expert human input and documented process in supporting a defensible review.

Inside TAR

Predictive Coding / Machine Learning Classification
The core analytical component in which a machine learning model is trained to distinguish relevant from non-relevant documents based on human coding decisions, typically applied within e-discovery or large-scale review contexts. The model learns from example documents and applies its classification across a broader collection.
Training / Seed Set
A subset of documents reviewed and coded by knowledgeable human reviewers to teach the model. Depending on the TAR workflow adopted, this may be a defined seed set (often associated with earlier-generation approaches) or a continuously updated training pool.
Continuous Active Learning (CAL) and Simple Passive Learning Workflows
Distinct approaches to how the model is trained. Continuous active learning iteratively selects documents for review to refine the model as review proceeds, while simple passive or one-time training workflows rely on a fixed training set. The choice of workflow depends on the tool, the matter, and organizational or legal expectations.
Human Review and Subject-Matter Input
The expert coding decisions and quality oversight that guide and validate the model. TAR typically supplements rather than replaces human judgment, and the reliability of outputs depends heavily on the consistency of human coding.
Validation and Quality Measures
Statistical and procedural checks used to assess the effectiveness of the review, which may include sampling and estimation of measures such as recall and precision. The specific metrics and acceptable thresholds often depend on the matter, jurisdiction, and any applicable court or regulatory expectations.
Defensibility and Documentation
The record of the methodology, decisions, and validation steps used, maintained to demonstrate that the process was reasonable and proportionate. Expectations regarding disclosure of TAR methodology to opposing parties or regulators vary by jurisdiction and forum.

Common questions

Answers to the questions practitioners most commonly ask about TAR.

Does Technology-Assisted Review replace human reviewers entirely?
No. Despite the name, TAR is designed to support and prioritize human review rather than eliminate it. The process typically depends on human subject-matter experts to make relevance determinations that train and validate the system, and human judgment remains central to quality control, defensibility, and handling of ambiguous or borderline material. TAR is more accurately understood as a way of focusing reviewer effort than as full automation of review.
Is TAR the same as keyword searching?
Not in the same sense. Keyword searching retrieves documents based on the presence or absence of specified terms, whereas TAR generally uses statistical or machine-learning techniques to infer relevance from patterns learned during training, extending beyond literal term matching. The two approaches are often used together, with keyword methods sometimes informing culling or sampling, but they rely on different underlying mechanisms and should not be treated as interchangeable.
How do you decide when TAR is appropriate for a given matter?
Suitability typically depends on factors such as the volume and nature of the material, the proportion of likely relevant documents, time and cost constraints, and the defensibility expectations of the relevant forum. Because acceptance and expectations vary by jurisdiction, sector, and the specific proceeding, organizations often assess proportionality and consult applicable procedural rules or guidance before adopting TAR.
What role do subject-matter experts play in a TAR workflow?
Subject-matter experts commonly provide the relevance determinations used to train or validate the system, and their consistency directly affects outcomes. Depending on the workflow adopted, experts may review seed or training sets, resolve conflicting coding decisions, and participate in validation. Clear, consistent criteria for what counts as relevant are generally important to reliable results.
How is the effectiveness of a TAR process typically measured and validated?
Effectiveness is often evaluated using sampling and statistical measures that estimate how completely relevant material has been identified and how much non-relevant material remains. Validation approaches depend on the workflow and the expectations of the relevant forum, and organizations frequently document their methodology, sampling decisions, and results to support defensibility. The specific metrics and thresholds considered acceptable can vary by jurisdiction and matter.
What documentation should be maintained to support the defensibility of a TAR process?
Defensibility generally benefits from a documented record of the methodology used, including decisions about training, coding criteria, validation and sampling steps, and quality-control measures. Maintaining such documentation as reliable evidence of the process can support later scrutiny. What is expected in practice depends on organizational policy and the requirements of the applicable jurisdiction or proceeding.

Common misconceptions

TAR fully automates review and removes the need for human involvement.
TAR is generally a technology-assisted process rather than a fully automated one. Human reviewers typically provide the coding decisions that train the model and the oversight that validates its output; the quality of results depends substantially on that human input.
TAR and records management are essentially the same discipline.
TAR is primarily a review and classification technique associated with e-discovery and large-scale document assessment, not a lifecycle recordkeeping practice. Records management concerns the control of records as evidence of activity across their lifecycle, and while TAR may operate over collections that include records, it addresses a narrower task of relevance or responsiveness classification.
A single TAR workflow or validation approach is universally accepted and applicable everywhere.
TAR encompasses more than one workflow, such as continuous active learning and simple passive learning, and the acceptability of a given approach, its validation measures, and any disclosure obligations depend on jurisdiction, forum, and the specifics of the matter.

Best practices

Ensure that human reviewers who code training documents are knowledgeable and apply consistent criteria, since the reliability of the model depends heavily on the quality and consistency of that coding.
Select a TAR workflow, such as continuous active learning or a passive training approach, that is appropriate to the matter, the tool's capabilities, and any applicable legal or organizational expectations.
Incorporate validation and quality measures, including sampling where appropriate, to assess the effectiveness of the review rather than relying on the model's outputs without verification.
Document the methodology, coding decisions, and validation steps to support defensibility, recognizing that expectations for disclosure of TAR methodology vary by jurisdiction and forum.
Confirm applicable jurisdictional and forum-specific requirements before relying on TAR, as acceptable approaches, validation thresholds, and disclosure obligations differ across regimes.
Treat TAR as a supplement to, rather than a replacement for, human judgment and subject-matter oversight throughout the review.