Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:21:54.402884Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 8 inbound Pith citation observations for arXiv:2502.04563.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:21:54.402884Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:40.140731Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:39:49.219978Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25d8dd53-25ba-43ec-a7d4-be3170ff06a2 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Abadi, P
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0470d58-a380-4d5e-8ee8-315f51b48d1b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale AMD optimizes EPYC mem- ory with NUMA
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d06c896b-3398-4e5e-a8e7-746a266e8586 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eff9cfd-e659-4b94-8b1e-4066cd92f195 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale AMD XDNA adaptive architecture, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 742007e0-5522-46e4-996e-3ce97166d261 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a586aa1-3995-4491-9c92-c6b7a35089e3 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale A cellular computer to implement the kalman filter algorithm
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f87e73b-9c25-4847-a0a0-88c36b568547 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale GEMM with collective operations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7fcae573-c9b7-48d5-986a-3eb957f46366 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale 100× defect tolerance: How cerebras solved the yield problem, 2022
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd086f67-45ca-4832-863c-d548b1c0a75c · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Benchmark GEMV collectives, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93a1ac70-a6a1-4db0-bd7f-c15c37182fbd · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Chen et al
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20925e04-0168-4aeb-b218-7342da6405d6 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Dongarra, and David W
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ce62af5f-36c3-466d-8322-98213bff1d98 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale FlashAttention-2: Faster attention with bet- ter parallelism and work partitioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e1116a39-a746-4da6-9d10-71a02bc6e9c8 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c896dff-558c-4cfc-887f-76577c15e453 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale SambaNova’s new AI chip and the quest for efficiency, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ddec9caf-42b6-4a52-9843-c3892b099d91 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41dc756-1749-428b-95a1-e60654c83af9 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daba1a71-d00e-4c44-902c-76a3a09321f7 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Tensor processing units for machine learning: An introduction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf3f1b9b-1e0a-4e8d-aef5-7ba1c7500d90 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Tenstorrent Blackhole and Metalium for standalone AI processing, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b9d7094-e21e-4fca-9992-9ed8118786d5 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Chiplet/interposer co-design for power delivery network optimization in heterogeneous 2.5-d ICs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4c3aa38-ec3c-462c-9c3c-285aeccaf0f1 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Efficient memory man- agement for large language model serving with Page- dAttention
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a8a7eed-8ec5-4575-8258-75f879ef2283 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale TSMC bets big on advanced packaging,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 15ae97b3-9de6-4bc8-b5bc-f9c843ecce6b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale ReSA: Reconfig- urable systolic array for multiple tiny DNN tensors
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 938e0880-b2d4-4051-8d0a-1cde62124697 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Gonzalez, and Ion Stoica
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a3b6c9a1-8982-4f73-afb5-f040108f4d58 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Cerebras architecture deep dive: First look in- side the hardware/software co-design for deep learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3404f0b1-556f-4f6e-a040-5a5a0d357fc6 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Scaling deep learn- ing computation over the inter-core connected intelli- gence processor with T10
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c965d35-ad95-44c5-a24a-96e795d5bf0c · outbound
WaferLLM: Large Language Model Inference at Wafer Scale TENET: A framework for modeling tensor dataflow based on relation-centric notation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e05af19e-01ff-4c0e-9d1f-06c1182f0435 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Near- optimal wafer-scale reduce
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b5aac418-d9df-4089-8594-2be3ba6f090b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Rammer: Enabling holistic deep learning compiler optimizations with rTasks
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb14a3af-ccb0-4ffa-acc8-ee16d3ecd3ea · outbound
WaferLLM: Large Language Model Inference at Wafer Scale An electrical-thermal co-simulation model of chiplet hetero- geneous integration systems
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0101c665-bc87-4693-9697-f61b2b2f69d5 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Introducing MTIA: Meta’s next-generation training and inference accelerator for AI, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20579df0-e4f2-4d9c-846a-2e7b51bce463 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Azure Maia: For the era of AI from silicon to software to systems, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c4cb968-8a7f-4070-aeb5-4438eb11bcdf · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Efficient large-scale language model training on GPU clusters using Megatron-LM
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d480844b-3c4b-439b-b772-cc2e5c572f31 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Openai o3 and o4-mini system card
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90057448-1fab-4adf-8a3b-16a570348c21 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Paszke, S
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d81b6e30-9f5d-41ee-80a0-c32d3033cc86 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Efficiently scal- ing transformer inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d04338a6-1bb9-4568-81d7-ed81316bc1fb · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Rock et al
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74b22d8f-7f75-4219-a107-04924cb34d4d · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Welder: Scheduling deep learning memory access via tile-graph
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b10a1373-0988-4f0c-ae07-fc0fc98692fb · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Souri, Kaustav Banerjee, Amit Mehrotra, and Krishna C
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30ced605-e430-4616-8c8c-9d8435161f23 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Cerebras and g42 break ground on condor galaxy 3, an 8 exaflops ai supercom- puter
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 422088e2-f153-4688-a943-b462f2629ee8 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Cerebras powers perplex- ity sonar with industry’s fastest ai inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8efa9c7-ba37-4cdb-aa11-b3ef4eec0c36 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale DOJO: The microarchitecture of Tesla’s exa-scale com- puter
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e17cdad2-50ef-4d0d-90fa-2bbf750dc6d8 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 40c3d0d4-f478-450c-9b99-612ea00643f7 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 742b4138-2316-460a-b01c-dd851f2b601b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Cerebras brings instant inference to mistral le chat
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07cfb9f5-a679-4ea9-b980-1deced62566d · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Ladder: Enabling efficient low- precision deep learning computing through hardware- aware tensor transformation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af81a517-e45c-4079-aafd-53f9fb64d5d9 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Application defined on-chip networks for heteroge- neous chiplets: An implementation perspective
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 379300a7-5c26-4a55-9fc8-46e539135281 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Static random-access memory,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 602e7d7c-5fc5-4814-8bed-ef523d2f193b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Wafer-scale integration, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 95bd58e5-84c1-4e3c-97f3-311bf62a7bb4 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale LoongServe: Efficiently serving long-context large language models with elas- tic sequence parallelism
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9886b644-2fa8-4fa8-8684-38addff6a143 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale DIS- TAL: the distributed tensor algebra compiler
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 65ffef5b-a088-48a2-8594-b09bd93fdf4b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Zhao et al
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 236a53b9-f82b-4e9c-808d-cdf3dd650f2c · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Alpa: Automating inter-and intra-operator parallelism for dis- tributed deep learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f80f57b5-2bf4-4127-b992-3ad075767420 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Sglang: Efficient execution of structured language model programs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6798210-4e09-4008-bd73-09b5bee2e050 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale FlexTensor: An automatic schedule exploration and optimization framework for tensor com- putation on heterogeneous system
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 176952f1-a9c6-49f3-a494-d069cd3feace · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 030b42b6-6bc5-488c-b103-2c12a9405591 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Exploring TensorRT to improve real-time inference for deep learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f35b1d5-92db-43a3-bc23-f1fc7adcb85b · outbound
WaferLLM: Large Language Model Inference at Wafer Scale ROLLER: Fast and efficient tensor compilation for deep learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd737c22-2cc1-4ee6-9899-f7ab408cc806 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0db404a6-c439-4996-8afc-9de7e5e0b5d9 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ded79cf-5a9e-4430-8a3e-3cebcfdfdde6 · outbound
WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff374140-96f0-4959-b4d6-62e81ed3eb2e · inbound
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations WaferLLM: Large Language Model Inference at Wafer Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286877e8-cfc0-4a97-9286-4b127b755e8c · inbound
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search WaferLLM: Large Language Model Inference at Wafer Scale
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e01058-9d02-427d-bac0-7691930b8048 · inbound
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques WaferLLM: Large Language Model Inference at Wafer Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a428dd-11fb-4344-855c-5420538801e3 · inbound
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive WaferLLM: Large Language Model Inference at Wafer Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a570285-7e65-4e33-9fd7-962efce7241f · inbound
Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing WaferLLM: Large Language Model Inference at Wafer Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 602676aa-5bc4-4c26-8bcc-5f0d44667e6a · inbound
MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference WaferLLM: Large Language Model Inference at Wafer Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation abcb77bf-4119-43d9-8063-6be776ac7911 · inbound
SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems WaferLLM: Large Language Model Inference at Wafer Scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9cdecb2-3a00-45fa-bb60-e2ba77924a45 · inbound
SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems WaferLLM: Large Language Model Inference at Wafer Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.