Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:01:46.443737Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2501.11779.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:01:46.443737Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T10:07:06.141581Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T10:13:17.805330Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd8de5fb-360c-4b60-9eb0-a7750f47ec22 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9afcc265-6c9e-493f-b264-bc249066982e · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b17145cb-ea74-4204-b576-b5c5fa1b91f1 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 40834a54-9776-447b-9d81-0578cc68fa6c · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5476b1d9-7ca1-40a4-a9fc-ce9d313715b6 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 32ae0556-c6f6-4648-9fc9-bef2c07adc3f · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 185deb44-aa19-4381-b0c9-6c2c9ad6812b · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d6f5c94e-829c-40e4-b784-ebd37a967a15 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a96ce76-ce01-404f-91b7-de152b17b688 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49fd9485-fe74-4915-ab5c-509fc4535d98 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference IEEE Standard for Floating-Point Arithmetic
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30508ef3-964c-4a89-b862-55a835c64b15 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b852a1a-d85d-48e5-ac61-781431ef0e37 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d873e317-48e0-4e95-9548-57e999c07e5a · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b10e56c-ab3a-48e8-8081-68882a03d8eb · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 821c8ffe-0d58-480e-ad1a-3e3f0e051c72 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Language Models are Few-Shot Learners
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0f0fc0-f990-4b19-974a-1ff4ed8a7f96 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc4d592-14b0-4872-b016-572b6b7205cb · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e3309a-8922-4408-8530-5963028a63ff · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gonzalez, Ion Stoica, and Eric P
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f7acc176-41e2-4c67-900c-1bfb10329ef2 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0020f9a-f8d6-4595-8b60-7ddc5010b49c · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa93b2c9-582e-40a1-a862-d63a7b72ce3c · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference The Llama 3 Herd of Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de14574-7611-41a9-bc56-159929147cf3 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df405801-7510-4393-ba9d-c170c58d0343 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 230a61d0-3247-457e-8bf9-041081624c0e · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference A Review of Sparse Expert Models in Deep Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5db946e-4b6a-4c3e-a57b-2ba309b685e3 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b71d3ac2-652b-4b38-bf8e-b82d796d912d · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be11884-41b4-47eb-aaa8-98915c69c71e · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Hydragen: High-Throughput LLM Inference with Shared Prefixes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b7bc4e-ddf8-4658-af62-ab28c98b8f48 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Mixtral of Experts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8976462-c76a-4b5f-abbb-98e1e5806d2d · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd323758-0973-4522-b625-69ca3526214e · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gonzalez, Hao Zhang, and Ion Stoica
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0114ff0f-e917-4c51-826f-57473f90e5c3 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e6f5535-8ba1-432d-8af5-80c971140973 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa98683e-5b9a-43f0-a904-6183b0d6c34c · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5166c2-c1a0-4b1c-adc1-74aa3e589515 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b20b16e-4b1b-4b6b-84d2-eeba8a108ed3 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Can Foundation Models Wrangle Your Data?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6328fcdd-c8a6-4999-ae95-e56846b34da7 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Very Deep Transformers for Neural Machine Translation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe44fe3-76cd-49eb-aa0b-1d355d017a14 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77e6827-4ea1-4d3f-a9bc-6b04e3dcacaf · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a12cfd8b-e8f4-4433-b6db-008b8cc6c386 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d608c8dd-c737-4c05-acef-536a9115555d · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed95a91c-db04-4f62-89b6-c23584fb5e00 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008b4344-7dd3-45f0-807d-d483848decd0 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Splitwise: Efficient generative LLM inference using phase splitting
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09dea148-68ba-4b46-ab31-f4962826ad67 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4d17f1-51ec-4698-af02-cf6ce691bef2 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a63bae5-8be2-49fb-8f40-8dc294552bbb · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3109990-33ca-4a29-8056-87fc6b526f6d · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538f9376-388b-4a88-bd12-368ad2660701 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddef0a13-13fc-4950-85de-3fddc4607b09 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc9cdde-5e85-4042-8003-9455e23d8fd1 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Fast Transformer Decoding: One Write-Head is All You Need
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1388f066-5ab7-4032-800b-0946cd7feb8e · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gemma: Open Models Based on Gemini Research and Technology
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e96f27-0529-4ed8-9291-2434bfab177a · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference LLaMA: Open and Efficient Foundation Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8075934-419a-4ff1-824a-69fc162d6966 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1b74c8-5585-4af5-9e9d-9f80e3477958 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Hashimoto
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa9fc5da-241b-48a4-a391-f12f25d25746 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference https://github.com/tatsu-lab/stanford_alpaca
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6638c6f8-f207-4fbc-9d19-e9a5194ea8a8 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Streaming Language Models with Attention Sinks
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a341c62-7758-4ef6-9b7b-f3a906243954 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ead83c85-b29d-4a54-8b74-e380f9d34488 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e4e9b8-bffb-405d-a62d-5efe4dd4944f · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Attention Is All You Need
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aac2e18-c272-4954-a29b-1199c8c3ddca · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e6ae9f20-fcf8-4cfd-9b14-1d03b3498b32 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78efc752-fb3c-448b-8aa0-b3c6f03409f4 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246c3827-a89b-4aa2-9451-d0bfe498b018 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8781437-3a20-4210-8d3c-598c1a342079 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference ZeRO-Offload: Democratizing Billion-Scale Model Training
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641f29ab-e41c-47fd-9499-c9961485dcc7 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Petals: Collaborative Inference and Fine-tuning of Large Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ea8d80-6826-4545-b75a-8d3a2f494986 · outbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ab72be-fe40-44fe-beda-082c7f07b3c8 · inbound
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.