Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

A Survey on Private Transformer Inference

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Private transformer inference has two bottlenecks — large matrix multiplications and non-linear functions — and the surveyed systems differ mainly by cryptographic setup and which layer they optimize.

desk verdict A well-organized PTI survey that is currently too incomplete to use: the promised evaluation guidelines are missing and the comparison tables mislabel several systems. read the letter →

arxiv 2412.08145 v1 pith:DCMPTJEH submitted 2024-12-11 cs.CR cs.AI

classification cs.CRcs.AI
keywords privatetransformerinferencesecuremulti-partycomputationhomomorphicencryptionSoftmaxapproximationGeLUmatrixmultiplicationprivacy-preservingmachinelearningarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey aims to establish that the design space of private transformer inference from 2022 to 2024 can be organized around two cryptographic bottlenecks: large matrix multiplications in the linear layers and complex non-linear functions (Softmax, GeLU, LayerNorm) in the attention and feed-forward blocks. It reviews the surveyed systems by their setup (two-party, two-party with a trusted dealer, and three-party), by the cryptographic tool they employ, and by which transformer component they optimize, and it proposes evaluation guidelines that report communication volume, runtime, and accuracy loss together. The survey's comparisons show a trade-off: homomorphic-encryption-based systems achieve low communication and non-interactivity but high computation, while MPC-hybrid systems trade communication and trust assumptions for speed. If the survey is accurate, a practitioner can use its tables to choose a system based on threat model and network environment rather than on isolated latency claims.

What carries the argument

The carrying mechanism is a layer-wise decomposition of a transformer encoder into linear operations (matrix multiplications in attention and feed-forward layers) and non-linear operations (Softmax, GeLU, and LayerNorm), cross-classified by cryptographic setup (two-party, two-party with a trusted dealer, three-party). Within that grid, the load-bearing objects are: secret-sharing schemes with Beaver triples (precomputed shared randomness that turns secure multiplication into one communication round) or re-sharing (refreshing shares by exchanging noisy local results) for secure multiplication; homomorphic-encryption schemes (BFV, CKKS, RNS-CKKS) with ciphertext packing and SIMD operations for matrix multiplication; and approximation techniques for non-linear functions, including low-degree polynomials, Taylor/Maclaurin/Chebyshev/Fourier series, and look-up tables. These mechanisms let the survey compare systems along the same axes: each system's reported communication volume, runtime, and accuracy loss are tied to which mechanisms it uses and which layer it optimizes.

What would settle it

A reader can settle the central claim by cross-checking every row of Tables 3, 4, 8, 9, 10, and 12 against the cited papers' own reported numbers and reference lists. The draft already shows two misattributions (BOLT appears with marker [26] in Table 3 and [43] in Table 4; Curl appears with marker [17] instead of [50]), so a systematic verification of all rows would show whether the survey's comparisons are reliable.

Watch

Extended reading notes

Core claim

The paper's central claim is that the current state of private transformer inference is best understood not as a contest between homomorphic encryption and secure multi-party computation, but as a layered design problem: each transformer component imposes a different cryptographic cost, and each system can be described by which layer it optimizes and under which setup. The paper argues that in two-party setups, large matrix multiplications are a dominant bottleneck because secure multiplication requires extra privacy protection; in dealer-assisted and three-party setups, that bottleneck moves to non-linear layers, which now account for most of the runtime in both MPC and HE systems. It further claims that accuracy preservation is achieved mainly by replacing non-linear functions with crypto-friendly approximations (low-degree polynomials, Taylor, Maclaurin, Chebyshev, or Fourier series, and look-up tables) and then recovering accuracy through knowledge distillation. The paper concludes by proposing evaluation guidelines, arguing that fair comparison requires reporting communication volume, runtime, accuracy loss, and the security model together.

Load-bearing premise

The survey's usefulness depends on its tables faithfully representing the cited systems' security models, runtimes, and reference markers; if those entries are wrong or misattributed, the comparative conclusions drawn from the survey would be misleading.

Editorial extensions

If this is right

  • If the survey's classification holds, future systems in the two-party setting should focus on jointly optimizing MatMul and non-linear layers, since neither alone determines end-to-end cost.
  • If the per-layer breakdowns are accurate, optimizing Softmax, GeLU, and LayerNorm will produce larger end-to-end gains than further MatMul speedups for dealer and three-party systems.
  • If the evaluation guidelines are adopted, reported numbers across studies become comparable, because current tables differ in network bandwidth, input size, and setup, making cross-paper comparison unreliable without normalization.
  • If the approximation-and-distillation trend continues, accuracy preservation will carry an extra training cost that must be included in any resource comparison, not just inference runtime.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not stated in the paper, but a reader could infer that the setup choice encodes a trust assumption: two-party systems avoid extra trust but pay for matrix multiplication, while dealer and three-party systems push that cost onto an assumed-honest helper; the two-party direction is therefore the harder test for the field.
  • Not stated in the paper, the per-layer tables imply a testable ordering under identical network conditions: an HE-only system will show near-zero communication but the longest runtime, a two-party hybrid will sit in the middle, and a three-party or dealer system will show the lowest runtime only if a helper is available.
  • Not stated in the paper, the proposed evaluation guidelines could be turned into a community benchmark that normalizes communication per token, per-layer runtime, and GLUE accuracy loss, which would convert the survey's qualitative comparisons into reproducible numbers.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript is a survey of private transformer inference (PTI). It covers background on transformer architecture and cryptographic primitives (MPC, HE), reviews roughly thirty PTI systems from 2022–2024, categorizes them by setup (2PC, 2PC-Dealer, 3PC), and discusses protocols for linear layers (MatMul) and non-linear layers (Softmax, GeLU, LayerNorm). The abstract and introduction promise three contributions: a comprehensive review, a breakdown of challenges and typical solutions, and proposed evaluation guidelines for resource efficiency and privacy guarantees. The review portion is structured around comparison tables and per-layer discussions, but the manuscript is incomplete: Section 8 is empty, several cross-references appear as unresolved 'Section ??', and the promised evaluation guidelines do not materialize in the text.

Significance. If the survey were completed and made internally consistent, it would be a useful reference for researchers working on private transformer inference. The paper has several strengths: it collects recent work into a single narrative, reproduces the standard cryptographic background, gives explicit approximation formulas for Softmax, GeLU, and LayerNorm, and provides links to open-source implementations. The paper does not claim a novel cryptographic derivation; its value is survey-level. However, the current significance is substantially undercut by the missing conclusion, unresolved cross-references, and citation errors in the comparison tables, because a survey's primary value lies in the reliability of its organization and tables.

major comments (4)
  1. [Section 8 / Abstract] The paper advertises evaluation guidelines in the abstract and introduction, but no such guidelines appear anywhere in the manuscript. Section 8, titled 'CONCLUSION', is empty, and the future-directions section is referenced as 'Section ??' in the Introduction and in Section 5. This is a missing contribution, not a presentation issue: a reader cannot use the paper for one of its two advertised central claims.
  2. [Tables 3, 4, 8, 9] Several comparison tables mislabel cited systems, which undermines the survey's core comparative function. Table 3 lists BOLT as reference [26] rather than [43]; Table 4 lists SecFormer in both the 2PC group and the 2PC-Dealer group, and lists Curl as [17] even though reference [17] is SIGMA and the bibliography entry for Curl is [50]; Tables 8 and 9 label NEXUS as [43] rather than [64]. Because the tables are the main deliverable for comparing systems, these errors are load-bearing.
  3. [Section 7, Table 12] Table 12 mixes runtimes across different models (BERT-Base, GPT2-Base, LLaMA-7B, ViT-Base), different datasets, different input sizes, and different network settings (e.g., 5 Gbps with 1 ms latency, 3 Gbps with 0.8 ms, 100 Mbps with 80 ms) with no normalization, no stated methodology, and no hardware/software environment details. As presented, the table cannot support any cross-system ranking of resource efficiency, yet Section 7 claims to compare experimental results.
  4. [Sections 4.1, 5, and 6.2] The manuscript contains multiple unresolved cross-references and broken exposition: Section 4.1 says 'we first introduce a secure inference system setup in Section ??', Section 5 says 'Section ?? first provides a breakdown', and the text after the MatMul discussion refers to attention equations as '(??)-(??)'. In addition, Section 6.2 contains the incomplete sentence 'Tech Tips: The function GeLU(𝑥) . Besides, polynomials are still available...'. These are not isolated typos but indicate that parts of the draft are unfinished, making the survey difficult to follow.
minor comments (6)
  1. [Section 3 title] The section title 'PIVACY THREATS IN SECURE INFERENCE' should be 'PRIVACY THREATS IN SECURE INFERENCE'.
  2. [Sections 2.3.2 and 5.2] The phrase 'secrete sharing' appears multiple times and should be 'secret sharing'.
  3. [Section 6.1] In the Softmax Tech Tips, the phrase 'to server as F(x)' should be 'to serve as F(x)'.
  4. [Section 4.3] The paragraph on 'Stuides [1, 32, 59]' contains a typo: 'Stuides' should be 'Studies'. Additionally, the sentence immediately following 'THE-X [7]' is a dangling fragment with no accompanying claim, so the discussion of client computation is incomplete.
  5. [Reference [65]] The reference title contains a typo: 'latency efficiefnt' should be 'latency efficient'.
  6. [Front matter] The copyright line reads '© 2018 Copyright held by the owner/author(s)' while the manuscript is an arXiv 2024 submission; this date appears inconsistent and should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the survey's content is drawn from external work and contains no fitted parameters or self-referential derivation chain.

full rationale

This paper is a literature survey; it introduces no novel protocol, derives no result from an assumed conclusion, and fits no parameters. Its equations (e.g., the attention, GeLU, and LayerNorm definitions in Sections 2 and 6) are standard textbook definitions reproduced from the cited literature, not predictions derived from the survey's own inputs. The survey's comparisons rely on reported numbers from external systems, and any inaccuracies in those tables (such as the BOLT/NEXUS citation mismatches noted in the manuscript) are correctness and completeness concerns, not circularity. There is no self-citation chain that supports a load-bearing claim, no ansatz smuggled in via the authors' prior work, and no renamed known result presented as a new derivation. The proposed evaluation guidelines are underdeveloped, but absence of content is not circular reasoning. Therefore the circularity score is 0.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The survey contains no new derivations. It defines standard background concepts (private inference, secret sharing, homomorphic encryption) using prior sources, and it reports results from other papers. No free parameters are fitted, no nonstandard axioms are introduced, and no new entities are postulated. The only promised novelty, the evaluation guidelines, is absent and therefore cannot be ledgered.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Private Transformer Inference." pith.science (2026). https://pith.science/paper/DCMPTJEH

@misc{pith2026241208145,
  author       = {Pith},
  title        = {Pith review of: A Survey on Private Transformer Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DCMPTJEH}},
  note         = {Machine review of arXiv:2412.08145}
}
read the original abstract

Transformer models have revolutionized AI, enabling applications like content generation and sentiment analysis. However, their use in Machine Learning as a Service (MLaaS) raises significant privacy concerns, as centralized servers process sensitive user data. Private Transformer Inference (PTI) addresses these issues using cryptographic techniques such as Secure Multi-Party Computation (MPC) and Homomorphic Encryption (HE), enabling secure model inference without exposing inputs or models. This paper reviews recent advancements in PTI, analyzing state-of-the-art solutions, their challenges, and potential improvements. We also propose evaluation guidelines to assess resource efficiency and privacy guarantees, aiming to bridge the gap between high-performance inference and data privacy.

Figures

Figures reproduced from arXiv: 2412.08145 by the authors.

Figure 1
Figure 1. Structure and workflow of a Transformer [ [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives

    cs.CR 2025-07 conditional novelty 4.0 of 10

    A structured survey of PPML efficiency optimizations, grouped into protocol, model, and system levels, with comparisons and future directions.

Reference graph

Works this paper leans on

69 extracted references · 45 canonical work pages · cited by 1 Pith paper

  1. [26]

    Brian Knott, Shobha Venkataraman, Awni Hannun, Shubho Sengupta, Mark Ibrahim, and Laurens van der Maaten. 2021. Crypten: Secure multi-party computation meets machine learning. Advances in Neural Information Processing Systems 34 (2021), 4961–4973

  2. [43]

    Qi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng, and Thomas Schneider. 2024. BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 130–130. Manuscript submitted to ACM A Survey on Private Transformer Inference 23

  3. [17]

    Kanav Gupta, Neha Jawalkar, Ananta Mukherjee, Nishanth Chandran, Divya Gupta, Ashish Panwar, and Rahul Sharma. 2023. SIGMA: secure GPT inference with function secret sharing. Cryptology ePrint Archive (2023)

  4. [50]

    Manuel B Santos, Dimitris Mouris, Mehmet Ugurbil, Stanislaw Jarecki, José Reis, Shubho Sengupta, and Miguel de Vega. 2024. Curl: Private LLMs through Wavelet-Encoded Look-Up Tables. Cryptology ePrint Archive (2024)

  5. [64]

    Jiawen Zhang, Jian Liu, Xinpeng Yang, Yinghao Wang, Kejia Chen, Xiaoyang Hou, Kui Ren, and Xiaohu Yang. 2024. Secure Transformer Inference Made Non-interactive. Cryptology ePrint Archive (2024)

  6. [1]

    Yoshimasa Akimoto, Kazuto Fukuchi, Youhei Akimoto, and Jun Sakuma. 2023. Privformer: Privacy-preserving transformer with mpc. In 2023 IEEE 8th European Symposium on Security and Privacy (EuroS&P) . IEEE, 392–410

  7. [2]

    Toshinori Araki, Jun Furukawa, Yehuda Lindell, Ariel Nof, and Kazuma Ohara. 2016. High-throughput semi-honest secure three-party computation with an honest majority. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security . 805–817

  8. [3]

    Donald Beaver. 1992. Efficient multiparty protocols using circuit randomization. In Advances in Cryptology—CRYPTO’91: Proceedings 11 . Springer, 420–432

Show all 69 references
  1. [4]

    Elette Boyle, Geoffroy Couteau, Niv Gilboa, and Yuval Ishai. 2018. Compressing vector OLE. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security . 896–912

  2. [5]

    Nishanth Chandran, Divya Gupta, Aseem Rastogi, Rahul Sharma, and Shardul Tripathi. 2019. EzPC: Programmable and efficient secure two-party computation for machine learning. In 2019 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 496–511

  3. [6]

    Dake Chen, Yuke Zhang, Souvik Kundu, Chenghao Li, and Peter A Beerel. 2023. RNA-ViT: Reduced-Dimension Approximate Normalized Attention Vision Transformers for Latency Efficient Private Inference. In 2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD) . IEE...

  4. [7]

    Tianyu Chen, Hangbo Bao, Shaohan Huang, Li Dong, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. The-x: Privacy-preserving transformer inference with homomorphic encryption. arXiv preprint arXiv:2206.00216 (2022)

  5. [8]

    Yuntian Chen, Xianjia Meng, Zhiying Shi, Zhiyuan Ning, and Jingzhi Lin. 2024. SecureTLM: Private inference for transformer-based large model with MPC. Information Sciences 667 (2024), 120429

  6. [9]

    Edward Chou, Josh Beal, Daniel Levy, Serena Yeung, Albert Haque, and Li Fei-Fei. 2018. Faster cryptonets: Leveraging sparsity for real-world encrypted inference. arXiv preprint arXiv:1811.09953 (2018)

  7. [10]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  8. [11]

    Yuanchao Ding, Hua Guo, Yewei Guan, Weixin Liu, Jiarong Huo, Zhenyu Guan, and Xiyong Zhang. 2023. East: Efficient and accurate secure transformer framework for inference. arXiv preprint arXiv:2308.09923 (2023)

  9. [12]

    Ye Dong, Wen-jie Lu, Yancheng Zheng, Haoqi Wu, Derun Zhao, Jin Tan, Zhicong Huang, Cheng Hong, Tao Wei, and Wenguang Cheng. 2023. Puma: Secure inference of llama-7b in five minutes. arXiv preprint arXiv:2307.12533 (2023)

  10. [13]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint...

  11. [14]

    David Evans, Vladimir Kolesnikov, Mike Rosulek, et al. 2018. A pragmatic introduction to secure multi-party computation. Foundations and Trends® in Privacy and Security 2, 2-3 (2018), 70–246

  12. [15]

    Craig Gentry. 2009. Fully homomorphic encryption using ideal lattices. In Proceedings of the forty-first annual ACM symposium on Theory of computing. 169–178

  13. [16]

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press

  14. [18]

    Meng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing, Guowen Xu, and Tianwei Zhang. 2022. Iron: Private inference on transformers. Advances in neural information processing systems 35 (2022), 15718–15731

  15. [19]

    Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh. 2006. A fast learning algorithm for deep belief nets. Neural computation 18, 7 (2006), 1527–1554

  16. [20]

    Xiaoyang Hou, Jian Liu, Jingyu Li, Yuhan Li, Wen-jie Lu, Cheng Hong, and Kui Ren. 2023. Ciphergpt: Secure two-party gpt inference. Cryptology ePrint Archive (2023)

  17. [21]

    Hai Huang and Yongjian Wang. 2024. SecBERT: Privacy-preserving pre-training based neural network inference system. Neural Networks 172 (2024), 106135

  18. [22]

    Zhicong Huang, Wen-jie Lu, Cheng Hong, and Jiansheng Ding. 2022. Cheetah: Lean and fast secure{Two-Party} deep neural network inference. In 31st USENIX Security Symposium (USENIX Security 22) . 809–826

  19. [23]

    Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351 (2019)

  20. [24]

    2018.{GAZELLE}: A low latency framework for secure neural network inference

    Chiraag Juvekar, Vinod Vaikuntanathan, and Anantha Chandrakasan. 2018.{GAZELLE}: A low latency framework for secure neural network inference. In 27th USENIX security symposium (USENIX security 18) . 1651–1669

  21. [25]

    Marcel Keller. 2020. MP-SPDZ: A versatile framework for multi-party computation. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security. 1575–1590

  22. [27]

    Nishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta, Aseem Rastogi, and Rahul Sharma. 2020. Cryptflow: Secure tensorflow inference. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 336–353

  23. [28]

    Dacheng Li, Hongyi Wang, Rulin Shao, Han Guo, Eric Xing, and Hao Zhang. 2022. MPCFORMER: FAST, PERFORMANT AND PRIVATE TRANS- FORMER INFERENCE WITH MPC. In The Eleventh International Conference on Learning Representations

  24. [29]

    Shaohua Li, Kaiping Xue, Bin Zhu, Chenkai Ding, Xindi Gao, David Wei, and Tao Wan. 2020. Falcon: A fourier transform based approach for fast and secure convolutional neural network predictions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  25. [30]

    Xuanqi Liu and Zhuotao Liu. 2023. Llms can understand encrypted prompt: Towards privacy-computing friendly transformers. arXiv preprint arXiv:2305.18396 (2023)

  26. [31]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)

  27. [32]

    Yanxin Liu and Qianqian Su. 2024. PPTIF: Privacy-Preserving Transformer Inference Framework for Language Translation. IEEE Access (2024)

  28. [33]

    Natasha Lomas. 2023. Italy orders ChatGPT blocked citing data protection concerns. TechCrunch, March 31 (2023)

  29. [34]

    Wen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li, Jian Liu, Kui Ren, Cheng Hong, Tao Wei, and WenGuang Chen. 2023. Bumblebee: Secure two-party inference framework for large transformers. Cryptology ePrint Archive (2023)

  30. [35]

    Brady D Lund and Ting Wang. 2023. Chatting about ChatGPT: how may AI and GPT impact academia and libraries? Library hi tech news 40, 3 (2023), 26–29

  31. [36]

    Jinglong Luo, Yehong Zhang, Zhuo Zhang, Jiaqi Zhang, Xin Mu, Hui Wang, Yue Yu, and Zenglin Xu. 2024. SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC. In Findings of the Association for Computational Linguistics ACL 2024 . 13333–13348

  32. [37]

    2023.{SecretFlow- SPU}: A Performant and{User-Friendly} Framework for{Privacy-Preserving} Machine Learning

    Junming Ma, Yancheng Zheng, Jun Feng, Derun Zhao, Haoqi Wu, Wenjing Fang, Jin Tan, Chaofan Yu, Benyu Zhang, and Lei Wang. 2023.{SecretFlow- SPU}: A Performant and{User-Friendly} Framework for{Privacy-Preserving} Machine Learning. In 2023 USENIX Annual Technical Conference (USE...

  33. [38]

    Cecily Mauran. 2023. Whoops, Samsung workers accidentally leaked trade secrets via ChatGPT. Mashable [online]. Dostupné z: https://mashable. com/article/samsungchatgpt-leak-details (2023)

  34. [39]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016)

  35. [40]

    Microsoft and OpenAI. 2023. Bing Chat. (2023). https://www.bing.com/search

  36. [41]

    Jungho Moon, Dongwoo Yoo, Xiaoqian Jiang, and Miran Kim. 2024. THOR: Secure Transformer Inference with Homomorphic Encryption.Cryptology ePrint Archive (2024)

  37. [42]

    OpenAI. 2022. ChatGPT. (2022). https://openai.com/blog/chatgpt

  38. [44]

    Dongjin Park, Eunsang Lee, and Joon-Woo Lee. 2024. Powerformer: Efficient privacy-preserving transformer with batch rectifier-power max function and optimized homomorphic attention. Cryptology ePrint Archive (2024)

  39. [45]

    Hongyuan Qu and Guangwu Xu. 2023. Improvements of Homomorphic Secure Evaluation of Inverse Square Root. In International Conference on Information and Communications Security . Springer, 110–127

  40. [46]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  41. [47]

    Deevashwer Rathee, Dacheng Li, Ion Stoica, Hao Zhang, and Raluca Popa. 2024. MPC-Minimized Secure LLM Inference.arXiv preprint arXiv:2408.03561 (2024)

  42. [48]

    Deevashwer Rathee, Mayank Rathee, Rahul Kranti Kiran Goli, Divya Gupta, Rahul Sharma, Nishanth Chandran, and Aseem Rastogi. 2021. Sirnn: A math library for secure rnn inference. In 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 1003–1020

  43. [49]

    Lorenzo Rovida and Alberto Leporati. 2024. Transformer-based language models and homomorphic encryption: An intersection with bert-tiny. In Proceedings of the 10th ACM International Workshop on Security and Privacy Analytics . 3–13

  44. [51]

    Sijun Tan, Brian Knott, Yuan Tian, and David J Wu. 2021. CryptGPU: Fast privacy-preserving machine learning on the GPU. In2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 1021–1038

  45. [52]

    Ahmed Tlili, Boulus Shehata, Michael Agyemang Adarkwah, Aras Bozkurt, Daniel T Hickey, Ronghuai Huang, and Brighter Agyemang. 2023. What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart learning environments 10, 1 (2023), 15

  46. [53]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  47. [54]

    Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962 (2019)

  48. [55]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  49. [56]

    Sameer Wagh, Shruti Tople, Fabrice Benhamouda, Eyal Kushilevitz, Prateek Mittal, and Tal Rabin. 2021. FALCON: Honest-Majority Maliciously Secure Framework for Private Deep Learning. Proceedings on Privacy Enhancing Technologies

  50. [57]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461 (2018)

  51. [58]

    Weize Wang and Yi Kuang. 2024. CipherFormer: Efficient Transformer Private Inference with Low Round Complexity.arXiv preprint arXiv:2403.16860 (2024)

  52. [59]

    Yongqin Wang, G Edward Suh, Wenjie Xiong, Benjamin Lefaudeux, Brian Knott, Murali Annavaram, and Hsien-Hsin S Lee. 2022. Characterization of mpc-based private inference for transformer-based models. In 2022 IEEE International Symposium on Performance Analysis of Systems and So...

  53. [60]

    Tianshi Xu, Lemeng Wu, Runsheng Wang, and Meng Li. 2024. PrivCirNet: Efficient Private Inference via Block Circulant Transformation. arXiv preprint arXiv:2405.14569 (2024)

  54. [61]

    Andrew C Yao. 1982. Protocols for secure computations. In 23rd annual symposium on foundations of computer science (sfcs 1982) . IEEE, 160–164

  55. [62]

    Chenkai Zeng, Debiao He, Qi Feng, Xiaolin Yang, and Qingcai Luo. 2024. SecureGPT: A Framework for Multi-Party Privacy-Preserving Transformer Inference in GPT. IEEE Transactions on Information Forensics and Security (2024)

  56. [63]

    Wenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong, Wen-jie Lu, Jin Tan, Runsheng Wang, and Ru Huang. 2023. Mpcvit: Searching for accurate and efficient mpc-friendly vision transformer with heterogeneous attention. In Proceedings of the IEEE/CVF International Conference on Compute...

  57. [65]

    Yuke Zhang, Dake Chen, Souvik Kundu, Chenghao Li, and Peter A Beerel. 2023. Sal-vit: Towards latency efficiefnt private inference on vit using selective attention search with a learnable softmax approximation. In Proceedings of the IEEE/CVF International Conference on Computer...

  58. [66]

    Chuan Zhao, Shengnan Zhao, Minghao Zhao, Zhenxiang Chen, Chong-Zhi Gao, Hongwei Li, and Yu-an Tan. 2019. Secure multi-party computation: theory, practice and applications. Information Sciences 476 (2019), 357–372

  59. [67]

    Mengxin Zheng, Qian Lou, and Lei Jiang. 2023. Primer: Fast private transformer inference on encrypted data. In 2023 60th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–6

  60. [68]

    Itamar Zimerman, Allon Adir, Ehud Aharoni, Matan Avitan, Moran Baruch, Nir Drucker, Jenny Lerner, Ramy Masalha, Reut Meiri, and Omri Soceanu. 2024. Power-Softmax: Towards Secure LLM Inference over Encrypted Data. arXiv preprint arXiv:2410.09457 (2024)

  61. [69]

    Itamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov, Omri Soceanu, and Lior Wolf. 2023. Converting transformers to polynomial form for secure inference over homomorphic encryption. arXiv preprint arXiv:2311.08610 (2023). Manuscript submitted to ACM 24 Yang et al. A MPC SETT...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.