Pith. sign in

Paper Citation Record · LEDGER

WaferLLM: Large Language Model Inference at Wafer Scale

As of 12 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 8 inbound Pith citation observations for arXiv:2502.04563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04563 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:21:54.402884Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:40.140731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:39:49.219978Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved6
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25d8dd53-25ba-43ec-a7d4-be3170ff06a2 · outbound

This paper cites Abadi, P.

WaferLLM: Large Language Model Inference at Wafer Scale Abadi, P

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.272941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.140958Z digest=sha256:329a70b9f79512fff7a5b23bb01f9e2f39562715ceee38251c39bb9984ea47ac

Observation c0470d58-a380-4d5e-8ee8-315f51b48d1b · outbound

This paper cites AMD optimizes EPYC mem- ory with NUMA.

WaferLLM: Large Language Model Inference at Wafer Scale AMD optimizes EPYC mem- ory with NUMA

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.257422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.146156Z digest=sha256:00ea1e812a0e493917578ab15994adfd193cf610fe511cf12f2bab6cd0c6c36f

Observation d06c896b-3398-4e5e-a8e7-746a266e8586 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

WaferLLM: Large Language Model Inference at Wafer Scale Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.150702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.150702Z digest=sha256:db188371441a19e7b74499e564d9d89c30432fa44db48c5246d057cec217c38c

Observation 7eff9cfd-e659-4b94-8b1e-4066cd92f195 · outbound

This paper cites AMD XDNA adaptive architecture, 2023.

WaferLLM: Large Language Model Inference at Wafer Scale AMD XDNA adaptive architecture, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.231836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.155440Z digest=sha256:bd0263320fa1f2b3e41c6e71b0579a76ee8226c4b1d63da76a675fc3a98ed8a0

Observation 742007e0-5522-46e4-996e-3ce97166d261 · outbound

This paper cites Language Models are Few-Shot Learners.

WaferLLM: Large Language Model Inference at Wafer Scale Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.160164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.160164Z digest=sha256:4b51ce65eb08cc147cb657db8613301a3296e4a3ee3cc96b8aa9d62117211f44

Observation 4a586aa1-3995-4491-9c92-c6b7a35089e3 · outbound

This paper cites A cellular computer to implement the kalman filter algorithm.

WaferLLM: Large Language Model Inference at Wafer Scale A cellular computer to implement the kalman filter algorithm

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.217042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.165762Z digest=sha256:cd8a4fc3875d083212bc916a5364bc0351767b8ec6fb23516a92d7c9c4aab0e7

Observation 1f87e73b-9c25-4847-a0a0-88c36b568547 · outbound

This paper cites GEMM with collective operations.

WaferLLM: Large Language Model Inference at Wafer Scale GEMM with collective operations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.202346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.170940Z digest=sha256:2ffbab69638853a34df8745f3d1bc049f7ed5c20cd2ea74c145060df28350c17

Observation 7fcae573-c9b7-48d5-986a-3eb957f46366 · outbound

This paper cites 100× defect tolerance: How cerebras solved the yield problem, 2022.

WaferLLM: Large Language Model Inference at Wafer Scale 100× defect tolerance: How cerebras solved the yield problem, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.188480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.175371Z digest=sha256:8bf452a490b10291839ec6490f1ed4d70bce3078186da50205596dee349b4904

Observation dd086f67-45ca-4832-863c-d548b1c0a75c · outbound

This paper cites Benchmark GEMV collectives, 2023.

WaferLLM: Large Language Model Inference at Wafer Scale Benchmark GEMV collectives, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.175515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.179810Z digest=sha256:4ac2b59b6bde597a6eef14164ace1978ba1750e56ca47550bb9d5eaf7a8bdca5

Observation 93a1ac70-a6a1-4db0-bd7f-c15c37182fbd · outbound

This paper cites Chen et al.

WaferLLM: Large Language Model Inference at Wafer Scale Chen et al

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.162009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.184170Z digest=sha256:483c898e99b56b2b2299f8b73b4f5f7ce9de88c41714da08f8cba21adcfeb83e

Observation 20925e04-0168-4aeb-b218-7342da6405d6 · outbound

This paper cites Dongarra, and David W.

WaferLLM: Large Language Model Inference at Wafer Scale Dongarra, and David W

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.148644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.188459Z digest=sha256:918c36f707af1c95c3c3d93d35b4dec58a7e15131575d92652caf0df2c03b004

Observation ce62af5f-36c3-466d-8322-98213bff1d98 · outbound

This paper cites FlashAttention-2: Faster attention with bet- ter parallelism and work partitioning.

WaferLLM: Large Language Model Inference at Wafer Scale FlashAttention-2: Faster attention with bet- ter parallelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.135307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.193322Z digest=sha256:9dfdac85b4ab9181d4175a490caa2ad792a0d83e421932aaaf82b352a6b4654a

Observation e1116a39-a746-4da6-9d10-71a02bc6e9c8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WaferLLM: Large Language Model Inference at Wafer Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.197682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.197682Z digest=sha256:a95005e383961049cd3e640ad16e6b3461dca5d192efb8bf95541903634d8ebb

Observation 9c896dff-558c-4cfc-887f-76577c15e453 · outbound

This paper cites SambaNova’s new AI chip and the quest for efficiency, 2023.

WaferLLM: Large Language Model Inference at Wafer Scale SambaNova’s new AI chip and the quest for efficiency, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.122223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.202748Z digest=sha256:b688a13ce10d416a6607670105ca363532578c435cd0800f651b1a6f535e3233

Observation ddec9caf-42b6-4a52-9843-c3892b099d91 · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

WaferLLM: Large Language Model Inference at Wafer Scale FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.206928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.206928Z digest=sha256:9e239eb02b0f7eda1deebd704421393b338a5464729f5ca1ad04f030b46c0f54

Observation d41dc756-1749-428b-95a1-e60654c83af9 · outbound

This paper cites OpenAI o1 System Card.

WaferLLM: Large Language Model Inference at Wafer Scale OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.211615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.211615Z digest=sha256:ec0fb971dd1bb7feb7d10d757d30199dab2f76e3fa717a22fda3464d645d05f5

Observation daba1a71-d00e-4c44-902c-76a3a09321f7 · outbound

This paper cites Tensor processing units for machine learning: An introduction.

WaferLLM: Large Language Model Inference at Wafer Scale Tensor processing units for machine learning: An introduction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.108718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.216334Z digest=sha256:da8dfcca2116e28c02da3c2b27aedd1df5cdc6a47497a969097e226c4bffe192

Observation cf3f1b9b-1e0a-4e8d-aef5-7ba1c7500d90 · outbound

This paper cites Tenstorrent Blackhole and Metalium for standalone AI processing, 2024.

WaferLLM: Large Language Model Inference at Wafer Scale Tenstorrent Blackhole and Metalium for standalone AI processing, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.094479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.220241Z digest=sha256:543339e56a883ebc284e436835fafd5a102f832deae4ccba2d9d66008dcf6d44

Observation 1b9d7094-e21e-4fca-9992-9ed8118786d5 · outbound

This paper cites Chiplet/interposer co-design for power delivery network optimization in heterogeneous 2.5-d ICs.

WaferLLM: Large Language Model Inference at Wafer Scale Chiplet/interposer co-design for power delivery network optimization in heterogeneous 2.5-d ICs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.080134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.224161Z digest=sha256:dfdbbb0ee303ebac7312518979a86c6c896a60bf1a9acf6fa0d539a6d56a6d57

Observation f4c3aa38-ec3c-462c-9c3c-285aeccaf0f1 · outbound

This paper cites Efficient memory man- agement for large language model serving with Page- dAttention.

WaferLLM: Large Language Model Inference at Wafer Scale Efficient memory man- agement for large language model serving with Page- dAttention

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.065697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.228140Z digest=sha256:e0c414010efe9cbb3d084e25078ca4186ba95e2474299d4608c71fd26410cf40

Observation 2a8a7eed-8ec5-4575-8258-75f879ef2283 · outbound

This paper cites TSMC bets big on advanced packaging,.

WaferLLM: Large Language Model Inference at Wafer Scale TSMC bets big on advanced packaging,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.050424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.232084Z digest=sha256:52fab92728dd4466923fd15b98473120e4a13b93e1dc2baf8b6c033c73d5fb09

Observation 15ae97b3-9de6-4bc8-b5bc-f9c843ecce6b · outbound

This paper cites ReSA: Reconfig- urable systolic array for multiple tiny DNN tensors.

WaferLLM: Large Language Model Inference at Wafer Scale ReSA: Reconfig- urable systolic array for multiple tiny DNN tensors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.020742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.241193Z digest=sha256:efa546b0c32e5b533dd51e93880087f9513db596dfae67bd6aef4cb9dee1e7ff

Observation 938e0880-b2d4-4051-8d0a-1cde62124697 · outbound

This paper cites Gonzalez, and Ion Stoica.

WaferLLM: Large Language Model Inference at Wafer Scale Gonzalez, and Ion Stoica

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:55.006231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.245086Z digest=sha256:e85a8eb2747c12d17540051abe9bc00c8bb73c92009d709006f9da034cf2bec7

Observation a3b6c9a1-8982-4f73-afb5-f040108f4d58 · outbound

This paper cites Cerebras architecture deep dive: First look in- side the hardware/software co-design for deep learning.

WaferLLM: Large Language Model Inference at Wafer Scale Cerebras architecture deep dive: First look in- side the hardware/software co-design for deep learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.992077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.248918Z digest=sha256:0431d553a06438ee82275d8ae6bec2fc96b1d69a5602fe6e013b3cf954ccf4c0

Observation 3404f0b1-556f-4f6e-a040-5a5a0d357fc6 · outbound

This paper cites Scaling deep learn- ing computation over the inter-core connected intelli- gence processor with T10.

WaferLLM: Large Language Model Inference at Wafer Scale Scaling deep learn- ing computation over the inter-core connected intelli- gence processor with T10

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.977486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.253125Z digest=sha256:5c470b3ffe3faf07bd288f28b6eaceee0d604482c8a16eaa195b75663d8c506e

Observation 8c965d35-ad95-44c5-a24a-96e795d5bf0c · outbound

This paper cites TENET: A framework for modeling tensor dataflow based on relation-centric notation.

WaferLLM: Large Language Model Inference at Wafer Scale TENET: A framework for modeling tensor dataflow based on relation-centric notation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.963893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.257355Z digest=sha256:86f8950df3060a461586e4ff485b4399e74fa322750ad44a116e2d81e79227c9

Observation e05af19e-01ff-4c0e-9d1f-06c1182f0435 · outbound

This paper cites Near- optimal wafer-scale reduce.

WaferLLM: Large Language Model Inference at Wafer Scale Near- optimal wafer-scale reduce

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.950407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.261817Z digest=sha256:f96add6969a5e2850d6377b715143276727947489b8f3a200be6fae0b14842c4

Observation b5aac418-d9df-4089-8594-2be3ba6f090b · outbound

This paper cites Rammer: Enabling holistic deep learning compiler optimizations with rTasks.

WaferLLM: Large Language Model Inference at Wafer Scale Rammer: Enabling holistic deep learning compiler optimizations with rTasks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.935653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.265936Z digest=sha256:6b0397fb9219006592939c6c77321390c88b90b793ed5e3c2ac81f92519cb4bc

Observation bb14a3af-ccb0-4ffa-acc8-ee16d3ecd3ea · outbound

This paper cites An electrical-thermal co-simulation model of chiplet hetero- geneous integration systems.

WaferLLM: Large Language Model Inference at Wafer Scale An electrical-thermal co-simulation model of chiplet hetero- geneous integration systems

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.921161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.270098Z digest=sha256:b67eb9791b8d714cceb9c4787992fcc009d7c6c439d04c742259abc96bbaeb21

Observation 0101c665-bc87-4693-9697-f61b2b2f69d5 · outbound

This paper cites Introducing MTIA: Meta’s next-generation training and inference accelerator for AI, 2024.

WaferLLM: Large Language Model Inference at Wafer Scale Introducing MTIA: Meta’s next-generation training and inference accelerator for AI, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.906816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.274213Z digest=sha256:82b81482ff96a1f7c872482fea3fdef04d989fb47dabdbbc61e07fdcb28be778

Observation 20579df0-e4f2-4d9c-846a-2e7b51bce463 · outbound

This paper cites Azure Maia: For the era of AI from silicon to software to systems, 2023.

WaferLLM: Large Language Model Inference at Wafer Scale Azure Maia: For the era of AI from silicon to software to systems, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.893136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.278152Z digest=sha256:919b8220051026a2b5ec1f77d6a844051d48939ba2097ca2b7c812005ff0db17

Observation 3c4cb968-8a7f-4070-aeb5-4438eb11bcdf · outbound

This paper cites Efficient large-scale language model training on GPU clusters using Megatron-LM.

WaferLLM: Large Language Model Inference at Wafer Scale Efficient large-scale language model training on GPU clusters using Megatron-LM

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.878875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.282071Z digest=sha256:433abefb5ead7358b4bbe304d7ff664501ae9349748386003c97fffda002fe87

Observation d480844b-3c4b-439b-b772-cc2e5c572f31 · outbound

This paper cites Openai o3 and o4-mini system card.

WaferLLM: Large Language Model Inference at Wafer Scale Openai o3 and o4-mini system card

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.863555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.286166Z digest=sha256:b0eb0e0a76b8f8f7de30297c5b140126e62a5d3c3327114e455a6caaa4053ee9

Observation 90057448-1fab-4adf-8a3b-16a570348c21 · outbound

This paper cites Paszke, S.

WaferLLM: Large Language Model Inference at Wafer Scale Paszke, S

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.840014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.294450Z digest=sha256:7c995950db54aa11beec1d7bda58f8c0c1b273e353be5b9a90ff950618caade4

Observation d81b6e30-9f5d-41ee-80a0-c32d3033cc86 · outbound

This paper cites Efficiently scal- ing transformer inference.

WaferLLM: Large Language Model Inference at Wafer Scale Efficiently scal- ing transformer inference

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.825719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.298804Z digest=sha256:f13cd1fdff4499572f1fd313891713f2effd4136933ec7a81323f25df44a9c41

Observation d04338a6-1bb9-4568-81d7-ed81316bc1fb · outbound

This paper cites Rock et al.

WaferLLM: Large Language Model Inference at Wafer Scale Rock et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.811162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.303421Z digest=sha256:195e98fe436efcbf8dc484dd8facf9bf4285ee30c5dc7e2eff5335be132d9ce5

Observation 74b22d8f-7f75-4219-a107-04924cb34d4d · outbound

This paper cites Welder: Scheduling deep learning memory access via tile-graph.

WaferLLM: Large Language Model Inference at Wafer Scale Welder: Scheduling deep learning memory access via tile-graph

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.796785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.307797Z digest=sha256:3770aca1dccde8d857fd5341648922e39e6d04f1770a190d30f0a2faae51b8fa

Observation b10a1373-0988-4f0c-ae07-fc0fc98692fb · outbound

This paper cites Souri, Kaustav Banerjee, Amit Mehrotra, and Krishna C.

WaferLLM: Large Language Model Inference at Wafer Scale Souri, Kaustav Banerjee, Amit Mehrotra, and Krishna C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.782502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.312229Z digest=sha256:36f5cfe82950210cc64e529945089b2422ff5dd323691305b7f1d721db9ce3fc

Observation 30ced605-e430-4616-8c8c-9d8435161f23 · outbound

This paper cites Cerebras and g42 break ground on condor galaxy 3, an 8 exaflops ai supercom- puter.

WaferLLM: Large Language Model Inference at Wafer Scale Cerebras and g42 break ground on condor galaxy 3, an 8 exaflops ai supercom- puter

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.768221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.316773Z digest=sha256:e77044509fe24045cb4e555e017b47572cc0d2e93ba9e92ba122afce17f7fb8a

Observation 422088e2-f153-4688-a943-b462f2629ee8 · outbound

This paper cites Cerebras powers perplex- ity sonar with industry’s fastest ai inference.

WaferLLM: Large Language Model Inference at Wafer Scale Cerebras powers perplex- ity sonar with industry’s fastest ai inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.754599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.321212Z digest=sha256:ff2b4584d487bf528218a1d8adc5b09d45039898f6c5d0db18d1bbda5fba64aa

Observation b8efa9c7-ba37-4cdb-aa11-b3ef4eec0c36 · outbound

This paper cites DOJO: The microarchitecture of Tesla’s exa-scale com- puter.

WaferLLM: Large Language Model Inference at Wafer Scale DOJO: The microarchitecture of Tesla’s exa-scale com- puter

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.741889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.325648Z digest=sha256:db5138e8d2ef4bee9c2bc5e934f94b59613e6288a340e5ce614973b14dc0283d

Observation e17cdad2-50ef-4d0d-90fa-2bbf750dc6d8 · outbound

This paper cites an unresolved cited work.

WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:21:54.728796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.330382Z digest=sha256:e69f05f4db1abee77e49fd2a2d98170d272363f07fccf905f1fa74da03b1a0b5

Observation 40c3d0d4-f478-450c-9b99-612ea00643f7 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

WaferLLM: Large Language Model Inference at Wafer Scale Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.715403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.334809Z digest=sha256:95dcd5d3bfd9d40b2f9f8b8bb06f01c46c269edef4536206be21a42390cad15d

Observation 742b4138-2316-460a-b01c-dd851f2b601b · outbound

This paper cites Cerebras brings instant inference to mistral le chat.

WaferLLM: Large Language Model Inference at Wafer Scale Cerebras brings instant inference to mistral le chat

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.702790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.339350Z digest=sha256:f53481266f35b7d940aca30d7d9570b06e916180806fbbd4ac32dc5a75b3dd51

Observation 07cfb9f5-a679-4ea9-b980-1deced62566d · outbound

This paper cites Ladder: Enabling efficient low- precision deep learning computing through hardware- aware tensor transformation.

WaferLLM: Large Language Model Inference at Wafer Scale Ladder: Enabling efficient low- precision deep learning computing through hardware- aware tensor transformation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.689708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.343841Z digest=sha256:d440e24648a0e96dc3c0a2d76279ffaa9e922f7bcb2c267f4690f19717c794b3

Observation af81a517-e45c-4079-aafd-53f9fb64d5d9 · outbound

This paper cites Application defined on-chip networks for heteroge- neous chiplets: An implementation perspective.

WaferLLM: Large Language Model Inference at Wafer Scale Application defined on-chip networks for heteroge- neous chiplets: An implementation perspective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.677137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.348251Z digest=sha256:f1dcfd1c9ae7875a335c4013d17c9077d290f6d3c3a4e2ae7ffebaf47d930843

Observation 379300a7-5c26-4a55-9fc8-46e539135281 · outbound

This paper cites Static random-access memory,.

WaferLLM: Large Language Model Inference at Wafer Scale Static random-access memory,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.663171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.352727Z digest=sha256:7eee470b23fe3d68290b37f308571f14b7841112279364e0ebbada01d10169b7

Observation 602e7d7c-5fc5-4814-8bed-ef523d2f193b · outbound

This paper cites Wafer-scale integration, 2024.

WaferLLM: Large Language Model Inference at Wafer Scale Wafer-scale integration, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.633328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.361889Z digest=sha256:6dcebba36dcdb4144aee36eae81db70f4722fe4b4a462029aaba162b2406721e

Observation 95bd58e5-84c1-4e3c-97f3-311bf62a7bb4 · outbound

This paper cites LoongServe: Efficiently serving long-context large language models with elas- tic sequence parallelism.

WaferLLM: Large Language Model Inference at Wafer Scale LoongServe: Efficiently serving long-context large language models with elas- tic sequence parallelism

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.618819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.366455Z digest=sha256:1e16c1d67be9c3cc2aed2db22a6ffa00a0f3c3a4e290838bfa5fac85a8684738

Observation 9886b644-2fa8-4fa8-8684-38addff6a143 · outbound

This paper cites DIS- TAL: the distributed tensor algebra compiler.

WaferLLM: Large Language Model Inference at Wafer Scale DIS- TAL: the distributed tensor algebra compiler

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.603662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.371071Z digest=sha256:7f0be6f03c27f35b04603f0e79163aec172754ce961a5e434a5969f13b63beff

Observation 65ffef5b-a088-48a2-8594-b09bd93fdf4b · outbound

This paper cites Zhao et al.

WaferLLM: Large Language Model Inference at Wafer Scale Zhao et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.588499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.375549Z digest=sha256:9bf2df3d1b6b73612e1629cd85dcda3a5695daf0e684592d452f7e6e12872846

Observation 236a53b9-f82b-4e9c-808d-cdf3dd650f2c · outbound

This paper cites Alpa: Automating inter-and intra-operator parallelism for dis- tributed deep learning.

WaferLLM: Large Language Model Inference at Wafer Scale Alpa: Automating inter-and intra-operator parallelism for dis- tributed deep learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.571727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.380046Z digest=sha256:d1a7d6a0e6a80ab6b1b6fc75f93dd7a0c0800145afbe41897e5dae3d073c4810

Observation f80f57b5-2bf4-4127-b992-3ad075767420 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.

WaferLLM: Large Language Model Inference at Wafer Scale Sglang: Efficient execution of structured language model programs

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.556246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.384445Z digest=sha256:330fc8d2cd0bd87a7cfdf6352f62accef22152c2e2c7ca3f7ddc055a15124ae8

Observation b6798210-4e09-4008-bd73-09b5bee2e050 · outbound

This paper cites FlexTensor: An automatic schedule exploration and optimization framework for tensor com- putation on heterogeneous system.

WaferLLM: Large Language Model Inference at Wafer Scale FlexTensor: An automatic schedule exploration and optimization framework for tensor com- putation on heterogeneous system

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.542960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.389298Z digest=sha256:28431c77c658f8c8219bc08286139c02989e791de89a195094750ba203b77359

Observation 176952f1-a9c6-49f3-a494-d069cd3feace · outbound

This paper cites Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

WaferLLM: Large Language Model Inference at Wafer Scale Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.529170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.393963Z digest=sha256:fb6563da0af87a6fb2b4fd4a11c8bfa38acc57535484f62ae25b9991247f03b0

Observation 030b42b6-6bc5-488c-b103-2c12a9405591 · outbound

This paper cites Exploring TensorRT to improve real-time inference for deep learning.

WaferLLM: Large Language Model Inference at Wafer Scale Exploring TensorRT to improve real-time inference for deep learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.515207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.398542Z digest=sha256:46c4be3e473d2b634762be10e4fcd5779747b08136279145cf67c5b2fe07a63d

Observation 7f35b1d5-92db-43a3-bc23-f1fc7adcb85b · outbound

This paper cites ROLLER: Fast and efficient tensor compilation for deep learning.

WaferLLM: Large Language Model Inference at Wafer Scale ROLLER: Fast and efficient tensor compilation for deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:21:54.501039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.402884Z digest=sha256:e5847c936a3a70e12f3cbd0be66de5bfc948a538727cbd443bc5b05e9c63a320

Observation fd737c22-2cc1-4ee6-9899-f7ab408cc806 · outbound

This paper cites an unresolved cited work.

WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T22:21:55.035858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.237298Z digest=sha256:df4a96bec3e94136382b9e6a6f7c3e19ebce1e1a938034bd4efc3afd70ff53ec

Observation 0db404a6-c439-4996-8afc-9de7e5e0b5d9 · outbound

This paper cites an unresolved cited work.

WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T22:21:54.647903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T22:21:54.357471Z digest=sha256:8bab0180a160a2e9a42eb07a4ecb59eddcde0e0b91abc33ff15cc46d2d272aec

Observation 6ded79cf-5a9e-4430-8a3e-3cebcfdfdde6 · outbound

This paper cites an unresolved cited work.

WaferLLM: Large Language Model Inference at Wafer Scale Unresolved cited work

Reference 2025

Resolution
parse uncertain
no resolver link, observed 2026-08-08T22:21:54.290170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.290170Z digest=sha256:fcaf22aefb378acb195ad4e49652c66320e51a02695714f836440493a63237e3

Pith citing papers

Observation ff374140-96f0-4959-b4d6-62e81ed3eb2e · inbound

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations cites this paper.

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations WaferLLM: Large Language Model Inference at Wafer Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:34.964920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:34.964920Z digest=sha256:fec9edf768354828440e5fbe06b428e1e076d02047ff408a56d1311fb24d1601

Observation 286877e8-cfc0-4a97-9286-4b127b755e8c · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search WaferLLM: Large Language Model Inference at Wafer Scale

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:40.140731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:40.140731Z digest=sha256:dfce43328da2b7bc2f5f4b249764e214f1a31e2c49796b64d184c692644fa1d1

Observation e4e01058-9d02-427d-bac0-7691930b8048 · inbound

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques cites this paper.

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques WaferLLM: Large Language Model Inference at Wafer Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:36.837844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:13:36.837844Z digest=sha256:17e061dbee3430c005e822283351fb014f64d0c8c8325dafdff1fbf2a9cb4806

Observation 11a428dd-11fb-4344-855c-5420538801e3 · inbound

ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive cites this paper.

ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive WaferLLM: Large Language Model Inference at Wafer Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:18:06.814541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:18:06.814541Z digest=sha256:858795436d71c25f6bb3b440ac83d92e20af9b363bc1e45a4e2d624bd74174eb

Observation 3a570285-7e65-4e33-9fd7-962efce7241f · inbound

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing cites this paper.

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing WaferLLM: Large Language Model Inference at Wafer Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:21:31.806071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T02:57:11.715370Z digest=sha256:5f5bef84e2c337b55ec60415562edbe50fdfe385f1ae0be636e9b5808cb888a2

Observation 602676aa-5bc4-4c26-8bcc-5f0d44667e6a · inbound

MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference cites this paper.

MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference WaferLLM: Large Language Model Inference at Wafer Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:39:49.221490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T06:26:59.984686Z digest=sha256:de13e9545d0771b4c6dd182342138810395183cd304a970cdc05d203d88b3018

Observation abcb77bf-4119-43d9-8063-6be776ac7911 · inbound

SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems cites this paper.

SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems WaferLLM: Large Language Model Inference at Wafer Scale

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:04:32.876266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T08:57:23.580353Z digest=sha256:8db5a67f6f80b2530556f31de7db3e2473884ff2fbf72362fb26cb4941a421ff

Observation e9cdecb2-3a00-45fa-bb60-e2ba77924a45 · inbound

SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems cites this paper.

SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems WaferLLM: Large Language Model Inference at Wafer Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T11:19:44.400939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:19:44.400939Z digest=sha256:d51f752f3618bd5475cf5f12f6ee0767225c05762935ce84f1683026dd737222