Pith. sign in

REVIEW 3 cited by

Secure Transformer Inference Protocol

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00025 v2 pith:SYGE7KBD submitted 2023-11-14 cs.CR cs.LG

classification cs.CRcs.LG
keywords securetransformerinferencesecuritytwo-partyefficiencymodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Security of model parameters and user data is critical for Transformer-based services, such as ChatGPT. While recent strides in secure two-party protocols have successfully addressed security concerns in serving Transformer models, their adoption is practically infeasible due to the prohibitive cryptographic overheads involved. Drawing insights from our hands-on experience in developing two real-world Transformer-based services, we identify the inherent efficiency bottleneck in the two-party assumption. To overcome this limitation, we propose a novel three-party threat model. Within this framework, we design a semi-symmetric permutation-based protection scheme and present STIP, the first secure Transformer inference protocol without any inference accuracy loss. Experiments on representative Transformer models in real systems show that STIP has practical security and outperforms state-of-the-art secure two-party protocols in efficiency by millions of times.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs

    cs.CR 2025-05 conditional novelty 8.0 of 10

    A sequential vocabulary-search attack decodes original prompts from unpermuted and permuted LLM hidden states, compromising PermLLM, STIP, and Centaur.

  2. Cascade: Token-Sharded Private LLM Inference

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Cascade performs LLM inference by sharding the token sequence across non-colluding nodes, claiming resistance to vocabulary-matching and learning-based reconstruction attacks while being orders of magnitude faster than SMPC.

  3. Scaling up FHE-based Privacy-Preserving ML: Higher Throughput, Longer Inputs for LLama-3-8B

    cs.CR 2026-01 conditional novelty 5.0 of 10

    A CKKS FHE system runs Llama-2-7B private inference on 4096-token prompts (only the last 128 encrypted) in 85s prefill and 33s/token on 8 RTX-4090 GPUs, with a mismatched abstract claiming faster Llama-3-8B numbers.

Pith tools