Pith. sign in

REVIEW 2 cited by

Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01946 v3 pith:XNMYHUIP submitted 2024-06-04 cs.CR cs.CL

classification cs.CRcs.CL
keywords attacksbilevesignaturespoofingtextcontentbi-leveldetectability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text watermarks for large language models (LLMs) have been commonly used to identify the origins of machine-generated content, which is promising for assessing liability when combating deepfake or harmful content. While existing watermarking techniques typically prioritize robustness against removal attacks, unfortunately, they are vulnerable to spoofing attacks: malicious actors can subtly alter the meanings of LLM-generated responses or even forge harmful content, potentially misattributing blame to the LLM developer. To overcome this, we introduce a bi-level signature scheme, Bileve, which embeds fine-grained signature bits for integrity checks (mitigating spoofing attacks) as well as a coarse-grained signal to trace text sources when the signature is invalid (enhancing detectability) via a novel rank-based sampling strategy. Compared to conventional watermark detectors that only output binary results, Bileve can differentiate 5 scenarios during detection, reliably tracing text provenance and regulating LLMs. The experiments conducted on OPT-1.3B and LLaMA-7B demonstrate the effectiveness of Bileve in defeating spoofing attacks with enhanced detectability. Code is available at https://github.com/Tongzhou0101/Bileve-official.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HaloMark: A Spectral Threshold for Embedding-Vector Watermarking under C2PA

    cs.CR 2026-08 reject novelty 6.0 of 10

    HaloMark proposes publishing an LSH commitment in the C2PA sidecar to watermark embeddings, but it derives the secret key from public manifest data, making the key recoverable by any sidecar observer.

  2. SoK: Watermarking for AI-Generated Content

    cs.CR 2024-11 conditional novelty 3.0 of 10

    A systematization of knowledge on watermarking for AI-generated content, unifying definitions, threat models, evaluation methods, and representative schemes across modalities.

Pith tools