REVIEW 16 cited by
Silent Data Corruptions at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Silent Data Corruption (SDC) can have negative impact on large-scale infrastructure services. SDCs are not captured by error reporting mechanisms within a Central Processing Unit (CPU) and hence are not traceable at the hardware level. However, the data corruptions propagate across the stack and manifest as application-level problems. These types of errors can result in data loss and can require months of debug engineering time. In this paper, we describe common defect types observed in silicon manufacturing that leads to SDCs. We discuss a real-world example of silent data corruption within a datacenter application. We provide the debug flow followed to root-cause and triage faulty instructions within a CPU using a case study, as an illustration on how to debug this class of errors. We provide a high-level overview of the mitigations to reduce the risk of silent data corruptions within a large production fleet. In our large-scale infrastructure, we have run a vast library of silent error test scenarios across hundreds of thousands of machines in our fleet. This has resulted in hundreds of CPUs detected for these errors, showing that SDCs are a systemic issue across generations. We have monitored SDCs for a period longer than 18 months. Based on this experience, we determine that reducing silent data corruptions requires not only hardware resiliency and production detection mechanisms, but also robust fault-tolerant software architectures.
Forward citations
Cited by 16 Pith papers
-
Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models
WorldMove migrates a live multi-GB world-model cache bit-identically within one interactive block, and an admissibility condition over state, dirty rate, and horizon decides when fleet moves are legal.
-
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
A co-designed GPU FFT library that is competitive with cuFFT and adds fused, low-overhead online fault tolerance via two-side ABFT.
-
How Far Are We from Detecting Flaky Tests? On the Limits of Code-Based Detection
After removing fix-commit shortcuts and enforcing project-disjoint evaluation, CodeBERT flakiness detectors collapse to majority baselines on developer-confirmed flaky tests with rerun-confirmed non-flaky labels.
-
An Efficient Fault-Tolerance Scheme for CKKS Computation on CPUs
A checksum-based consistency check detects single-bit hardware faults in CPU-based CKKS encrypted computation with 6.0–8.4% runtime overhead (average 6.8%), a 4.9× reduction versus direct checksum protection.
-
Kwai Keye-VL 1.5 Technical Report
Keye-VL-1.5 combines similarity-based Slow-Fast video token allocation with progressive context extension and iterative RL, reporting leading video-understanding results among 8B-scale multimodal models.
-
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
Test escapes causing silent data corruption occur at roughly 5,000 parts per million, at least 10 times above industrial targets, across compute chips in large data centers.
-
Kwai Keye-VL Technical Report
Kwai Keye-VL shows that a five-mode chain-of-thought cold-start plus mix-mode reinforcement learning can push an 8B multimodal model to strong short-video and general vision-language performance.
-
Oobleck: Low-Compromise Design for Fault Tolerant Accelerators
Splitting accelerators into modular stages with software fallbacks preserves 1.7x to 5.16x speedup after a single fault at lower area cost than redundancy.
-
Evaluating Different Fault Injection Abstractions on the Assessment of DNN SW Hardening Strategies
Fault injection at the application level and at the instruction level ranks DNN software hardening techniques differently, sometimes reversing the winner entirely.
-
On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization
For CKKS homomorphic encryption, single-bit client-side errors show two resilience patterns, and multiplication dominates whenever it is used.
-
Towards a future space-based, highly scalable AI infrastructure system design
Space-based AI compute is argued feasible via close-formation laser-linked satellites, radiation-survivable TPUs, and launch costs projected below $200/kg by the mid-2030s.
-
Understanding the Error Sensitivity of Privacy-Aware Computing
Single-bit flips in CKKS ciphertexts can silently corrupt decrypted outputs, and the RNS and NTT optimizations used in practice amplify the damage.
-
Logical Maneuvers: Detecting and Mitigating Adversarial Hardware Faults in Space
A sensor-triggered recovery framework lets a partially damaged RISC-V soft processor on an FPGA keep running by recompiling around failed ALU units or relocating via partial reconfiguration.
-
Enhancing Neural Network Robustness Against Fault Injection Through Non-linear Weight Transformations
Applying bounded nonlinear functions (tanh, softsign, arctan) to DNN weights makes models tolerate random bit flips at BER 1e-5 with only a few points of accuracy loss.
-
Harden Deep Neural Networks Against Fault Injections Through Weight Scaling
Scaling weights up before storage and down after loading reduces the damage of simulated bit-flips, improving fault-injected ImageNet Top-1 accuracy by up to roughly 54 points in Q2.5 ResNet50.
-
When and Where Faults Matter: A Study of Transient Errors in CKKS Multiplication
In unoptimized CKKS multiplication, a bit flip that corrupts both partial-product uses of c0 or c1 is mathematically masked, while a flip hitting only one use leads to silent data corruption.
Discussion (0). Continue with ORCID to comment.