Guide · PII redaction

PII redaction before an LLM

Detect and replace covered personal data before a prompt reaches the model, then test the route and review its record.

Updated

The short answer. Put PII detection in the request path before the LLM call. Define the values the detector must cover, replace each match with a stable placeholder, send the transformed prompt, and record the request. Test known matches, known limits, and every application route before approving use.

SUPERWISE® Sentinel is one deployed option for this path. It replaces emails, phone numbers, account numbers, and API keys with placeholders before a prompt leaves the PC, and logs every request. It does not redact names or passwords.

The five steps

  1. Set the boundary. Name the data that must stay out of the LLM and every route that can carry it.
  2. Detect covered values. Choose recognizers for the formats and languages your prompts contain.
  3. Replace before sending. Transform each match at the gateway before the provider receives the request.
  4. Test the route. Use synthetic examples for expected matches, expected misses, and ordinary text.
  5. Review the record. Confirm each covered request appears in the operational log.

1. Set the PII redaction boundary

Start with the company’s AI acceptable use policy. List the values staff must keep out of model prompts, the approved AI tools, the request routes, the owner, and the reporting path. Include pasted text, uploaded files, command output, connected apps, and application API calls.

Keep the language precise. Detection finds configured data patterns. Redaction or replacement changes the matched text. Anonymization is a wider goal that requires the organization to consider whether the remaining information can still identify a person. The AI glossary defines the gateway and guardrail terms used here.

2. Detect the values your prompts contain

PII detection starts with categories and recognizers. Test each category with synthetic values that match the formats your business uses. Include punctuation, spacing, country formats, labels, and surrounding prose in the test set.

  • Record the category and example format for every required match.
  • Record expected misses so reviewers understand the boundary.
  • Keep names, passwords, and context-dependent details in the test plan when the chosen tool leaves them unchanged.
  • Repeat tests when a recognizer, application route, or prompt format changes.

3. Replace matches before the LLM call

Place the control where every covered request crosses it. The detector returns the location and type of each match. The redaction step substitutes a placeholder, then the gateway forwards the transformed prompt to the model provider.

Stable, typed placeholders keep the sentence usable while showing what changed. A customer email can become <EMAIL>, and an account number can become <ACCOUNT_NUMBER>. Use synthetic data to confirm the provider-facing text contains the placeholder and excludes the original value.

4. Test detection, replacement, and logging

Run the same test through every browser, desktop tool, CLI, and application integration in scope.

  • Send one synthetic value for every covered category.
  • Confirm each original value is absent from the provider-facing prompt.
  • Confirm ordinary text remains readable after replacement.
  • Send synthetic examples for documented limits and record the unchanged result.
  • Confirm the request appears in the log and the owner can retrieve it.

5. Choose a managed gateway or an open-source framework

Data anonymization tools cover different jobs. A deployed gateway provides a request path and operational record. An open-source framework gives developers components to integrate into an application or data workflow.

Microsoft Presidio is an open-source approach for PII detection and anonymization. Its Analyzer identifies PII, and its Anonymizer can redact, replace, hash, or encrypt detected values (Presidio text anonymization). Presidio’s own documentation warns that automated detection does not guarantee every sensitive value will be found. Teams integrate, configure, operate, and test those components.

See the direct Sentinel and Microsoft Presidio comparison, or review the broader LLM gateway PII redaction solution.

Where Sentinel fits

Sentinel is the answer when the required path matches its defined coverage. It replaces emails, phone numbers, account numbers, and API keys with placeholders before a prompt leaves the PC, and logs every request. It does not redact names or passwords.

Read the PII redaction documentation, then test each required format and route with synthetic data. Keep names and passwords out of prompts through the policy, training, and access controls.

Frequently asked questions

What is PII redaction?

PII redaction detects personal identifiers in text and removes or replaces the matched values before another system receives the text.

What is PII detection?

PII detection is the step that identifies text spans or structured values that match defined personal-data categories. The result can then be reviewed, blocked, removed, or replaced.

How do you anonymize data before sending it to an LLM?

Define the prohibited data, detect it at a gateway before the model call, replace each matched value with a placeholder, test every covered route, and review the request log.

What is the difference between PII masking and redaction?

Masking obscures some or all of a value. Redaction removes or replaces the matched value. Teams should name the exact operation their tool performs and test the resulting text.

Sources

All sources read on 2026-10-07.