Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T20:33:14.055182Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2605.25655.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T20:33:14.055182Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6d2edf1c-7b5c-4f01-a649-3537258fff8e · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers LLaMA: Open and Efficient Foundation Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95e28b52-88d2-4668-90fd-4cf9793b9d6f · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Qwen Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b25a901-7741-4fb2-8349-057b0166afb2 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b24729db-6e86-48c9-b3d3-1f65efdda61e · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient memory management for large language model serving with pagedattention,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28ca0df-cba0-4d5a-bac2-94ae370e5abf · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers TensorRT-LLM,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657778f7-9fef-4f4c-89ee-d74f0b859b80 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82dc052-3a63-447a-b7f9-a44ca9487c34 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Large- scale parallelization and optimization of lattice qcd on tianhe new generation supercomputer,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e9a3747-e92f-4425-857a-33790a644fc7 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Mt-3000: a heterogeneous multi-zone processor for hpc,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0119f3-4f5a-4d2c-b148-846aa8e24575 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Performance analysis of cuda, openacc and openmp programming models on tesla v100 gpu,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae047b4-1459-4805-b46b-f74eb8462b31 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Mpi: a standard message passing interface,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a6cd98-85fd-4ac9-8cfd-58928f2157a2 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Attention is all you need,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4df453-538a-457d-ad6a-318341307ef8 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers BERT: A Review of Applications in Natural Language Processing and Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c4244fa-f83a-492e-9ffd-eeb09a8c6d1a · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Language models are unsupervised multitask learners,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7d1c02-3047-42c2-8ab1-b67ecc721b36 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Language mod- els are few-shot learners,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa8fd6e-d620-4b6c-bb35-add9719f99f9 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1dea15d5-e615-4a35-a884-d6ba80c9ff79 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21f14377-3571-4e4f-9ba5-b10509f94842 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ebbd727-232e-4a78-bfff-41660976d5a1 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Qwen2 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1f198ea-f712-4f87-bab2-ab647243519b · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddbe6671-80a4-4826-a793-95ae20a5bb46 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a76898-7317-486d-9e37-429289e25055 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Sparsegpt: Massive language models can be accurately pruned in one-shot,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f327fb1e-5677-4a0f-910b-04398689bbb4 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Llm-pruner: On the structural pruning of large language models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3dc798f-e95f-4069-b52d-59445f591d66 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers MiniLLM: On-Policy Distillation of Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f542ca8-bdeb-465e-ba21-48a3f60379c6 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888d6e6f-091a-4a3a-85ba-1992102554e9 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Longformer: The Long-Document Transformer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 542292c2-31d1-45dd-bcbd-b8c389648252 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Linformer: Self-Attention with Linear Complexity
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d95be2c3-ca5a-46fc-acc9-0b2073717876 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Reformer: The Efficient Transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45e8b6c0-c945-460d-948b-a0152440e980 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3ca5121-d7de-485e-b279-63a70b105c69 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3237c88-7dea-4707-bf75-c063fed24f3e · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient Streaming Language Models with Attention Sinks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1090996-0c2c-44cb-bd2f-5cf32274a2a9 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Orca: A distributed serving system for{Transformer-Based}generative models,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233a8c46-11e2-40e9-9cd8-bc11ec110e61 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Text Generation Inference,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bad1ad4-8492-4b17-8cd0-4099a8cf1dc4 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Flexgen: High-throughput generative inference of large language models with a single gpu,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5361c5e1-7e74-4f72-afda-c3f52ffbf880 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889c0a49-7b04-4a19-8809-a10db6997573 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient processing of deep neural networks: A tutorial and survey,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 894cf718-d9e1-42dc-bc74-dea31e4d690d · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Roofline: an insightful visual performance model for multicore architectures,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e96e60-b30d-4773-abe3-6f9f0800e040 · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Optimizing general matrix multiplications on modern multi-core dsps,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6629f8c6-9f15-4ee2-a633-2beda16cb7db · outbound
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers He served as the chief scientist of China National High Technology Program on high perfor- mance computing for 20 years
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.