REVIEW 4 cited by
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report evaluates the performance impact of enabling Trusted Execution Environments (TEE) on NVIDIA Hopper GPUs for large language model (LLM) inference tasks. We benchmark the overhead introduced by TEE mode across various LLMs and token lengths, with a particular focus on the bottleneck caused by CPU-GPU data transfers via PCIe. Our results indicate that while there is minimal computational overhead within the GPU, the overall performance penalty is primarily attributable to data transfer. For the majority of typical LLM queries, the overhead remains below 7%, with larger models and longer sequences experiencing nearly zero overhead.
Forward citations
Cited by 4 Pith papers
-
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing
Under GPU-CC, LLM serving losses come from a serialized VM–GPU bridge, not compute; simple scheduling and loader changes recover most of the gap on Blackwell.
-
Hardware Mechanisms to Dynamically Throttle AI Performance
Dynamic microarchitecture throttling of GPU memory resources can cut LLM inference performance by up to 80% with low hardware overhead, giving architects a continuous, hardware-enforced AI capability control.
-
Performance of Confidential Computing GPUs
Confidential GPU inference under model swapping is significantly slower than non-confidential, and the gap is driven by model loading, not by inference compute.
-
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
Confidential mode on an H100 under Intel TDX adds roughly 21-28% time-to-first-token overhead and 17-21% global token-throughput loss for two LLMs, and causes earlier saturation for the larger model.
Discussion (0). Continue with ORCID to comment.