Skip to main content
Category: Access and Security

Sensitive Information Type

Also known as: SIT, Sensitive information type entity, Pattern-based classifier
Simply put

A Sensitive Information Type is a predefined rule that automatically recognizes particular kinds of confidential content, such as social security numbers, credit card numbers, or bank account details, based on the patterns those items typically follow. It is used within data governance platforms to help find and sort sensitive material so it can be handled appropriately. The term is closely associated with Microsoft Purview, though the general concept of pattern-based content detection is not unique to any one product.

Formal definition

A Sensitive Information Type (SIT) is a pattern-based classifier that detects specific categories of sensitive content by matching defined patterns, and it may combine supporting elements such as keywords, formats, and proximity or confidence criteria to identify candidate matches. In Microsoft Purview, SITs are provided as built-in entity definitions covering data types like social security, credit card, and bank account numbers, and can also be created and managed as custom SITs where organizational requirements demand it. SITs support automated detection and classification of content, but their scope is limited to identifying that information matches a pattern; they do not themselves determine records status, retention, disposition, or the authoritative recordkeeping properties (authenticity, reliability, integrity, usability) of the content in which the sensitive data appears. The specific entity definitions and management capabilities described here are documented in the context of Microsoft Purview and should not be assumed to apply identically to other platforms or jurisdictions.

Why it matters

Organizations increasingly hold large volumes of unstructured content across email, documents, and collaboration platforms, much of which may contain confidential material such as social security numbers, credit card numbers, or bank account details. Locating that material manually is often impractical at scale, which is why pattern-based detection has become a common component of data governance programs. Sensitive Information Types provide a mechanism to surface content that matches recognizable patterns, enabling it to be identified and sorted so that appropriate handling can follow.

From an information governance perspective, the value of a SIT lies in supporting downstream decisions rather than making them. Detecting that a document contains what appears to be a credit card number does not by itself establish whether that document is a record, what retention period applies, or how it should ultimately be disposed of or preserved. Those determinations depend on organizational policy and, in many jurisdictions, on statutory and regulatory requirements that vary by sector and location. A SIT is best understood as an input to classification and risk-management workflows, not a substitute for the recordkeeping judgments that govern authenticity, reliability, integrity, and usability of the underlying content.

Because pattern matching identifies candidates rather than confirmed facts, professionals should treat SIT results as indicative and subject to review. Patterns can produce both false positives and missed matches, and the confidence associated with a detection depends on how the SIT is configured. Relying on detection alone, without governance policy to interpret and act on the results, can create a false sense of assurance about where sensitive information resides and how it is being controlled.

Who it's relevant to

Information Governance Officers
Those responsible for accountability frameworks may use SITs as one input into locating and classifying sensitive content across systems. They should be clear that a SIT identifies that content matches a pattern but does not itself determine records status, retention, or disposition, which remain governed by policy and, in many jurisdictions, by statutory and regulatory requirements.
Records Managers
Records managers may find SIT detection useful for surfacing confidential material that requires controlled handling. They should recognize that detecting sensitive data does not establish the authoritative recordkeeping properties of a document, nor does it decide whether the content is a record, a copy, a draft, or transitory information.
Privacy and Data Protection Professionals
Practitioners concerned with personal and confidential data can use SITs to help identify where categories such as financial or identity-related information appear in an organization's content. Because privacy obligations depend on jurisdiction and sector, detection results should be interpreted within the applicable legal framework rather than treated as a compliance outcome in themselves.
Compliance Leads
Compliance teams may incorporate SITs into workflows for identifying content that carries regulatory or contractual sensitivity. They should account for the limitations of pattern matching, including the potential for false positives and missed matches, and pair automated detection with review and defined governance policy.
Platform and IT Administrators
Administrators configuring data governance platforms, particularly Microsoft Purview, work directly with built-in and custom SIT definitions, including creating, modifying, and removing custom types. They should be aware that capabilities documented for one platform do not necessarily transfer identically to other tools or environments.

Inside SIT

Category of protected content
A sensitive information type identifies a class of content that an organization treats as requiring heightened handling, such as personal identifiers, financial account details, health information, or credentials. The precise categories that qualify as sensitive depend on jurisdiction, sector, and organizational policy.
Detection or matching criteria
Sensitive information types are often defined by patterns, keywords, formats, or contextual rules used to recognize the content within records and other information. The reliability of such detection varies, and matches may include false positives or miss content depending on how the criteria are configured.
Handling and control implications
Classifying content as a sensitive information type typically triggers additional controls across the lifecycle, which may affect access restrictions, retention decisions, disposition, and security measures. These implications are governed by organizational policy and applicable legal or regulatory obligations.
Relationship to the record
A sensitive information type describes the nature of content that may reside within a record, a copy, a draft, or transitory information. Identifying sensitive content does not by itself confer record status; whether an item is an authoritative record depends on its authenticity, reliability, integrity, and usability as evidence of activity.

Common questions

Answers to the questions practitioners most commonly ask about SIT.

Is a sensitive information type the same thing as a record?
No. A sensitive information type is a classification describing the nature or category of information based on the harm that could result from its unauthorized disclosure, alteration, or loss. A record, by contrast, is content that serves as evidence of activity and typically must exhibit properties such as authenticity, reliability, integrity, and usability. Sensitivity is an attribute that may apply to a record, but it may equally apply to data, drafts, copies, or transitory information that are not records at all. Treating the two as interchangeable conflates a security or privacy characteristic with the evidential status of the content.
Does labeling something as a sensitive information type determine how long it must be kept?
Not directly. Sensitivity classification concerns the protection and handling that content requires, while retention concerns how long content is kept before disposition. These are distinct dimensions that often intersect but are not equivalent. In many organizations, retention is driven by the business, legal, and regulatory value of the record, whereas sensitivity is driven by the potential for harm. A single item may be highly sensitive yet short-lived, or non-sensitive yet subject to a lengthy statutory retention period. Requirements depend on jurisdiction, sector, and organizational policy, so the two should generally be managed as related but separate controls.
How can an organization decide which sensitive information types it needs to define?
Definitions typically flow from an assessment of the information the organization holds, the obligations that apply to it, and the harm that could arise from mishandling. Many organizations begin by reviewing applicable privacy, security, and regulatory obligations, which vary by jurisdiction and sector, and by consulting stakeholders across legal, compliance, security, and business functions. The resulting set of types is usually documented in policy and aligned with the organization's broader classification scheme. The scope of what qualifies as sensitive should be stated explicitly to avoid ambiguity at the point of handling.
Where should sensitive information type designations be applied in the records lifecycle?
Depending on organizational policy, sensitivity is often assigned as early as creation or capture, so that appropriate handling controls apply from the outset. It may also be reviewed at classification and revisited at points such as transfer or disposition, since the sensitivity of content can change over time. Applying the designation early tends to support consistent protection, but organizations should also allow for reassessment, because a type initially considered sensitive may cease to be so, or vice versa, as circumstances change.
How do sensitive information types relate to legal holds and disposition?
Sensitivity classification and disposition are separate controls that must be coordinated. When content is subject to a legal hold, its scheduled disposition is typically suspended regardless of its sensitivity, and hold requirements differ across jurisdictions and matters. Sensitivity may, however, influence how disposition is carried out, for example by calling for secure destruction methods or additional authorization before transfer. It is generally advisable to ensure that disposition processes account for both the sensitivity of the content and any applicable holds rather than treating either factor in isolation.
What roles are usually involved in maintaining sensitive information type definitions?
Maintenance is often shared across functions rather than owned by any single role. Records and information governance staff commonly steward the classification scheme, while privacy, security, legal, and compliance functions contribute expertise on obligations and risk, and business units provide context on the content itself. Because requirements depend on jurisdiction and sector and can change over time, many organizations establish periodic review so that definitions remain accurate. Clear accountability for updates and for communicating changes to those who handle the content is generally considered good practice.

Common misconceptions

Sensitive information types and records are the same thing, so anything flagged as sensitive is automatically a record.
A sensitive information type characterizes the content, not the evidential status of the item. Sensitive content may appear in records, copies, drafts, or transitory information alike, and its presence does not determine whether something qualifies as an authoritative record.
The set of categories that count as sensitive is universal and fixed across organizations and countries.
What is treated as sensitive typically depends on jurisdiction, sector, and organizational policy. Categories and their associated obligations vary, so a definition appropriate in one regime may not apply in another.
Automated detection of a sensitive information type is definitive and identifies all relevant content accurately.
Detection often relies on patterns, keywords, or contextual rules that can produce false positives or miss content depending on configuration. Results generally require human review and should not be treated as complete or infallible.

Best practices

Define sensitive information types with reference to your organization's policies and the legal and regulatory obligations of the relevant jurisdictions and sectors, rather than assuming a single universal standard.
Distinguish the identification of sensitive content from the determination of record status, and ensure classification workflows address both separately where relevant.
Treat automated detection results as indicative rather than definitive, and incorporate human review to manage false positives and missed content.
Document the detection criteria and handling rules associated with each sensitive information type so that controls across access, retention, and disposition are applied consistently and defensibly.
Review and update sensitive information type definitions periodically to reflect changes in applicable obligations, organizational risk posture, and the content landscape.
Apply appropriate controls consistently to sensitive content wherever it resides, including copies, drafts, and transitory information, not only to authoritative records.