Pith. sign in

REVIEW 4 cited by

Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.10785 v1 pith:MZUWQKQJ submitted 2023-11-16 cs.CL cs.LG

classification cs.CLcs.LG
keywords textinformationsanitizationcoherencecontextualdatadomainslanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the context of information systems, text sanitization techniques are used to identify and remove sensitive data to comply with security and regulatory requirements. Even though many methods for privacy preservation have been proposed, most of them are focused on the detection of entities from specific domains (e.g., credit card numbers, social security numbers), lacking generality and requiring customization for each desirable domain. Moreover, removing words is, in general, a drastic measure, as it can degrade text coherence and contextual information. Less severe measures include substituting a word for a safe alternative, yet it can be challenging to automatically find meaningful substitutions. We present a zero-shot text sanitization technique that detects and substitutes potentially sensitive information using Large Language Models. Our evaluation shows that our method excels at protecting privacy while maintaining text coherence and contextual information, preserving data utility for downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Truthful Text Sanitization Guided by Inference Attacks

    cs.CL 2024-12 conditional novelty 7.0 of 10

    INTACT generates abstraction-sorted replacement candidates for sensitive spans and selects the most specific candidate that resists LLM-based inference attacks, achieving a strong privacy-utility trade-off on the Text...

  2. Public Data Assisted Differentially Private In-Context Learning

    cs.AI 2025-09 conditional novelty 4.0 of 10

    A private ICL algorithm that aggregates LLM responses with DPM clustering and uses public data representatives achieves near-non-private utility at epsilon=1.

  3. Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    A survey of LLM routing and hierarchical inference techniques that proposes an unvalidated unified evaluation metric called the Inference Efficiency Score.

  4. Image deidentification in the XNAT ecosystem: use cases and solutions

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A XNAT/DicomEdit pipeline scored 99.61% on the MIDI-B deidentification benchmark after test-set feedback, with remaining failures concentrated in addresses and burned-in pixels.

Pith tools