Pith. sign in

REVIEW 1 cited by

Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18371 v2 pith:H2GL4JYO submitted 2024-10-24 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords textaudioclapinferencedatadetectorinputmembership
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio can disclose PII, particularly when combined with related text data. Therefore, it is essential to develop tools to detect privacy leakage in Contrastive Language-Audio Pretraining(CLAP). Existing MIAs need audio as input, risking exposure of voiceprint and requiring costly shadow models. We first propose PRMID, a membership inference detector based probability ranking given by CLAP, which does not require training shadow models but still requires both audio and text of the individual as input. To address these limitations, we then propose USMID, a textual unimodal speaker-level membership inference detector, querying the target model using only text data. We randomly generate textual gibberish that are clearly not in training dataset. Then we extract feature vectors from these texts using the CLAP model and train a set of anomaly detectors on them. During inference, the feature vector of each test text is input into the anomaly detector to determine if the speaker is in the training set (anomalous) or not (normal). If available, USMID can further enhance detection by integrating real audio of the tested speaker. Extensive experiments on various CLAP model architectures and datasets demonstrate that USMID outperforms baseline methods using only text data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming

    cs.CR 2025-05 reject novelty 5.0 of 10

    A reward-driven, PPO-finetuned LLM pipeline claims to generate diverse, evasive webshell payloads with higher escape rates than prompt-engineering baselines.

Pith tools