Pith. sign in

Paper Citation Record · LEDGER

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference

As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2501.11779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11779 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:01:46.443737Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:07:06.141581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:13:17.805330Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd8de5fb-360c-4b60-9eb0-a7750f47ec22 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.812432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.140974Z digest=sha256:ff18430a6f4ba4b8ed87ef51b4cb3b7e37d2218b91bab71fd627acd8a2546e52

Observation 9afcc265-6c9e-493f-b264-bc249066982e · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.797233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.146738Z digest=sha256:1093fc6b47a77396bc16f27289b4e50856900b33b2daf69d9aa8bd2c2da74db6

Observation b17145cb-ea74-4204-b576-b5c5fa1b91f1 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.782633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.152117Z digest=sha256:8e158fd2251345782ecf949d54f66001808b30a825793601b7bd8747b95020a9

Observation 40834a54-9776-447b-9d81-0578cc68fa6c · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.764911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.156915Z digest=sha256:4576c76a693bc589adaeda25b7c5ead6ce513581451e02d93a00efd5d5a32045

Observation 5476b1d9-7ca1-40a4-a9fc-ce9d313715b6 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.749908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.161851Z digest=sha256:86a088b7168a15adcae3f6e72bdb04110c1e58ac3e9f005ff6c44c800cacc0d8

Observation 32ae0556-c6f6-4648-9fc9-bef2c07adc3f · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.732416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.166848Z digest=sha256:5c7546891b376eaff01eb1daa0225e5de34e54273dc6262be009543d3ae46a7f

Observation 185deb44-aa19-4381-b0c9-6c2c9ad6812b · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.715213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.172598Z digest=sha256:3c0af652be9f2df8a4ea7cd1b8ba62790abab86a924e9e8353661bdbddeca342

Observation d6f5c94e-829c-40e4-b784-ebd37a967a15 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.698905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.177212Z digest=sha256:1b541de84026b22b6f463c3a778eb3371a94e3421fc0dd5452fed10fb716352c

Observation 5a96ce76-ce01-404f-91b7-de152b17b688 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.680983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.181969Z digest=sha256:9b3c31a532863105284650fa0ce29f6b3a173c1b357ffc34cbf3a7474db8d789

Observation 49fd9485-fe74-4915-ab5c-509fc4535d98 · outbound

This paper cites IEEE Standard for Floating-Point Arithmetic.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference IEEE Standard for Floating-Point Arithmetic

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.186890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.186890Z digest=sha256:4a44d2d7d32493210ba94077087bce1f03f5d77cbe026526a06a3c2f7cc516f8

Observation 30508ef3-964c-4a89-b862-55a835c64b15 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.191693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.191693Z digest=sha256:db6b50432b8ff271ca9dbb94d5a36190f5f0e77d9e18828b06a019f46ef2b0db

Observation 0b852a1a-d85d-48e5-ac61-781431ef0e37 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.196980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.196980Z digest=sha256:5259583a9a226504724c7814a698220f64222bd0728b07a9ee43c85b953144b2

Observation d873e317-48e0-4e95-9548-57e999c07e5a · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.202032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.202032Z digest=sha256:0fe50e036cadd9ce9606b0e74ad948fc0f79b305cbfc90f1a832f40ebd71c8d9

Observation 4b10e56c-ab3a-48e8-8081-68882a03d8eb · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.664194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.206986Z digest=sha256:a58a928971dae0ae6fd3d9c06544ff4bb314c72cca82bf0a17644aebdd7bec5e

Observation 821c8ffe-0d58-480e-ad1a-3e3f0e051c72 · outbound

This paper cites Language Models are Few-Shot Learners.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Language Models are Few-Shot Learners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.216356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.216356Z digest=sha256:2e231af177a57b58522a80c09f6dbf72aa405dcd865731f2cbb6034adc0a12de

Observation 6b0f0fc0-f990-4b19-974a-1ff4ed8a7f96 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.221091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.221091Z digest=sha256:13e58263f2fa913de4e4fdb6705c15208225e4420ab3adb7ca7767e105ff578a

Observation dfc4d592-14b0-4872-b016-572b6b7205cb · outbound

This paper cites Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.225490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.225490Z digest=sha256:4735da9c3ed87648ee4ae07937e9238f819225486c80f40d33d9e97117ffcbff

Observation 98e3309a-8922-4408-8530-5963028a63ff · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gonzalez, Ion Stoica, and Eric P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:01:47.647618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.230083Z digest=sha256:dcae8023a885f16e8d9270c6fe80ef890e041553c68b6975f49ab861d5c83dba

Observation f7acc176-41e2-4c67-900c-1bfb10329ef2 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.234652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.234652Z digest=sha256:a72bf649066bf10d443412e09acae2d6d4808393e1c91c2781acceceec0cc073

Observation a0020f9a-f8d6-4595-8b60-7ddc5010b49c · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.239316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.239316Z digest=sha256:c47d9c6de10d994066d1de4d96a7271c3cc0daa2fed1f3b0d1bd27192014ccd2

Observation fa93b2c9-582e-40a1-a862-d63a7b72ce3c · outbound

This paper cites The Llama 3 Herd of Models.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.247690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.247690Z digest=sha256:35bc7572338f74733eed328bb6eef6d8c28c38b2183b2278fafd7ded71362d21

Observation 6de14574-7611-41a9-bc56-159929147cf3 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.243579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.243579Z digest=sha256:c05087173df4e3d9c14b0bfd822cc77b26806b377029fa5bd8c04f5081847a8e

Observation df405801-7510-4393-ba9d-c170c58d0343 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.621766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.256388Z digest=sha256:94649afdbd4a332115d3c68a447041917302b90a49b56cce99aa7ed533722d08

Observation 230a61d0-3247-457e-8bf9-041081624c0e · outbound

This paper cites A Review of Sparse Expert Models in Deep Learning.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference A Review of Sparse Expert Models in Deep Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.251978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.251978Z digest=sha256:465d7284f846374da43d853a60976029f37829c988a361445d0f1f41c18029fa

Observation e5db946e-4b6a-4c3e-a57b-2ba309b685e3 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.606695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.264975Z digest=sha256:dc72163a535a5660b6366b15952c3dfadf67e4bf7f217c252b417a9f068494af

Observation b71d3ac2-652b-4b38-bf8e-b82d796d912d · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.260751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.260751Z digest=sha256:38fede9883605ef2b866f9782d4a77e7c8f1fbb87b26037fccbf467d2ef0788c

Observation 7be11884-41b4-47eb-aaa8-98915c69c71e · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.274340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.274340Z digest=sha256:9cd3b4d333fd91b1d9f2fc8f804446a3cbd3563f2dae351bedcbd080b48bb9dc

Observation d4b7bc4e-ddf8-4658-af62-ab28c98b8f48 · outbound

This paper cites Mixtral of Experts.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Mixtral of Experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.269431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.269431Z digest=sha256:4b40b8d8112623ceaf180fa34f5370e882e73860603f8af529ea5e6d64f12e35

Observation f8976462-c76a-4b5f-abbb-98e1e5806d2d · outbound

This paper cites AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.288769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.288769Z digest=sha256:d980b7dd696e4114298edf5fcc70440635f0f8b1e3b3cac87d298b6215a0ca9c

Observation bd323758-0973-4522-b625-69ca3526214e · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gonzalez, Hao Zhang, and Ion Stoica

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.279142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.279142Z digest=sha256:e7be9b794e9d369e4037f5c2b7657fb9dc154a778cbbfbc730d9350da00497a2

Observation 0114ff0f-e917-4c51-826f-57473f90e5c3 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.298972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.298972Z digest=sha256:0a46904f48feafe87d5cacf5adf75df03172bb0e22e16097b7f2ded2ab6d3583

Observation 8e6f5535-8ba1-432d-8af5-80c971140973 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.304481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.304481Z digest=sha256:c2e31e4d7b4439bae261525babea9493bd3094a784d85b96769dace5bfdb526d

Observation fa98683e-5b9a-43f0-a904-6183b0d6c34c · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.293818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.293818Z digest=sha256:7d76f41a7a1c2840ee9ee1bc4a4aca82b600dbd8e5c360c5cef7bf462d1fee18

Observation ee5166c2-c1a0-4b1c-adc1-74aa3e589515 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.313249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.313249Z digest=sha256:3e827f90bb5f2e6ba8964b8d3112f7e0c24b5c244f6a2168a4df05b3601505a0

Observation 9b20b16e-4b1b-4b6b-84d2-eeba8a108ed3 · outbound

This paper cites Can Foundation Models Wrangle Your Data?.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Can Foundation Models Wrangle Your Data?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.317445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.317445Z digest=sha256:d7ec61435e804a1bc1a50331b22c574ce11d21de7ac226e50a0a560e32bb213b

Observation 6328fcdd-c8a6-4999-ae95-e56846b34da7 · outbound

This paper cites Very Deep Transformers for Neural Machine Translation.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Very Deep Transformers for Neural Machine Translation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.309018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.309018Z digest=sha256:ba7c26871b2d9fee799e7ee8b8c69c175a34c1e1b897c7b1a7b305b9953f6228

Observation 6fe44fe3-76cd-49eb-aa0b-1d355d017a14 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.325426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.325426Z digest=sha256:67281541d32e255c794c97f75c51596d47d0cfade62229c838c9a6e4ee86715b

Observation c77e6827-4ea1-4d3f-a9bc-6b04e3dcacaf · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.566549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.329244Z digest=sha256:6d2f12cfdfce97d39c8171d88996588c24a86c59f57df6c1c0716fb615bdd867

Observation a12cfd8b-e8f4-4433-b6db-008b8cc6c386 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.581844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.321446Z digest=sha256:f199c798456ae00a1403f5b6bf860bfcf4b864cd70bf565cad393652aa9774af

Observation d608c8dd-c737-4c05-acef-536a9115555d · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.551053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.338347Z digest=sha256:7275d8a62a0caed07ce0dbfceaf028722a7e6ec8374c25656b7f59d032dd8543

Observation ed95a91c-db04-4f62-89b6-c23584fb5e00 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.342837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.342837Z digest=sha256:03f6518b738f4026ec44938a10d105db4102a6beb691d66e1b114076634b7e50

Observation 008b4344-7dd3-45f0-807d-d483848decd0 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Splitwise: Efficient generative LLM inference using phase splitting

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.333688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.333688Z digest=sha256:c38797411df764269cebcc4fca1cb233c26a646e6db3efb0a696a5604ddf04ef

Observation 09dea148-68ba-4b46-ab31-f4962826ad67 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.352066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.352066Z digest=sha256:121470964bd7d35dce46baf14c1b0b0e750f5da28e34cc785001359bd9c3d6d7

Observation 4a4d17f1-51ec-4698-af02-cf6ce691bef2 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.511723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.362128Z digest=sha256:f3ce642acd10d4f63dcd99bead486dfcf45e764327e93e299467bd450f41bd7f

Observation 8a63bae5-8be2-49fb-8f40-8dc294552bbb · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.347288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.347288Z digest=sha256:dee8b636a0df4466a4ff515c871cc055c351593b5204acdbdebe8059360d7438

Observation e3109990-33ca-4a29-8056-87fc6b526f6d · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.371711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.371711Z digest=sha256:d93a3f61c2f357dcbe8c59e67163ffbcd04cdbc316afde8b0dc53be06cb9583b

Observation 538f9376-388b-4a88-bd12-368ad2660701 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.376642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.376642Z digest=sha256:7c7c1d35669f38cf600bb3ccb64562f14053bc088b74e94400281bad7c82da95

Observation ddef0a13-13fc-4950-85de-3fddc4607b09 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.381691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.381691Z digest=sha256:9ed4fe0fc11555b4b7c1c474da0248b27eb7d6df6f883f137099449d5f0f8d2e

Observation afc9cdde-5e85-4042-8003-9455e23d8fd1 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Fast Transformer Decoding: One Write-Head is All You Need

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.366806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.366806Z digest=sha256:1a5d30f911d6ae6897c692bb85c85ca9a1b218a7dc08180250e3659424cdf81d

Observation 1388f066-5ab7-4032-800b-0946cd7feb8e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.395770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.395770Z digest=sha256:ab889d2108d8497eb3fd49b0f5ccda9dbbba962956b542bd4ac781f59cc06d30

Observation e6e96f27-0529-4ed8-9291-2434bfab177a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.400525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.400525Z digest=sha256:0714bb3df49a370d9f70ae81580c860d173e89422a8d7d78ffc5824a75561c3b

Observation c8075934-419a-4ff1-824a-69fc162d6966 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.405679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.405679Z digest=sha256:970ddffd1fe7c75f489e96ea1337b14d6791f3cf88c01e4637dba64cdd6a5102

Observation 7b1b74c8-5585-4af5-9e9d-9f80e3477958 · outbound

This paper cites Hashimoto.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Hashimoto

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:01:47.495870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.386417Z digest=sha256:cfd476eea1d4a5f7815ba23ef6e887fe24555b0967713aba4443004e46a2d440

Observation aa9fc5da-241b-48a4-a391-f12f25d25746 · outbound

This paper cites https://github.com/tatsu-lab/stanford_alpaca.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference https://github.com/tatsu-lab/stanford_alpaca

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:01:47.480502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.391253Z digest=sha256:351f66389178627a316ec8ad1853dc4dd85ceef7bd9da5ad3cc5800871c0faf3

Observation 6638c6f8-f207-4fbc-9d19-e9a5194ea8a8 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Streaming Language Models with Attention Sinks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.419673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.419673Z digest=sha256:98ed0721d41d0973a521594b97a72fd49c4d83cf0d74ef8e4c4823c8dcbc97b6

Observation 1a341c62-7758-4ef6-9b7b-f3a906243954 · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.446422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.424728Z digest=sha256:a671d27cfcefa78d4f23df71844a9700fe68c0af4ea98c5f46cabf27adbb7dba

Observation ead83c85-b29d-4a54-8b74-e380f9d34488 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.429279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.429279Z digest=sha256:a3294ff74afabfa3fde91dfcaccaff1d811cb962b87af5397f3e87ac25132357

Observation 94e4e9b8-bffb-405d-a62d-5efe4dd4944f · outbound

This paper cites Attention Is All You Need.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Attention Is All You Need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.410490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.410490Z digest=sha256:9c7222c2146c3709ceb6b69092dae0e2032504bd777e7798595290d43f839962

Observation 6aac2e18-c272-4954-a29b-1199c8c3ddca · outbound

This paper cites an unresolved cited work.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:01:47.463238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:01:46.414954Z digest=sha256:2294bdd50aa368c071aa9ffd7b3eec0d2b1c2507423042d4c135a8016dad3ee8

Observation e6ae9f20-fcf8-4cfd-9b14-1d03b3498b32 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.443737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.443737Z digest=sha256:581b0db446068ac03ad08c4e7ec3f1b44dd8a53f25082fdd9c0c9f626d14cf68

Observation 78efc752-fb3c-448b-8aa0-b3c6f03409f4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.434131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.434131Z digest=sha256:3aa2a66580ddea27c7c3dbaa59b4651ec2a918a06a299bcd3574baf523000d8f

Observation 246c3827-a89b-4aa2-9451-d0bfe498b018 · outbound

This paper cites Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.438859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.438859Z digest=sha256:f6e9af70c821308dceb43a41f5a2b60990cf315306087326b15f03b076732c93

Observation f8781437-3a20-4210-8d3c-598c1a342079 · outbound

This paper cites ZeRO-Offload: Democratizing Billion-Scale Model Training.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.357295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.357295Z digest=sha256:c0f9733a3aed0d7f1aacbd6a1930bd4027c883ba087597cd7fe482268d0bd7d5

Observation 641f29ab-e41c-47fd-9499-c9961485dcc7 · outbound

This paper cites Petals: Collaborative Inference and Fine-tuning of Large Models.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Petals: Collaborative Inference and Fine-tuning of Large Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.211564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.211564Z digest=sha256:06d0526ad02a38fe80aa0df31cb56f1c936c9c8d75a1854bde8f6637b7d95881

Observation a5ea8d80-6826-4545-b75a-8d3a2f494986 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.284071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.284071Z digest=sha256:920f313b186db3199e3727d9bf8802dcd66c6471bb7d45a59ac6a698e9b1f699

Pith citing papers

Observation 92ab72be-fe40-44fe-beda-082c7f07b3c8 · inbound

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference cites this paper.

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference Glinthawk: A Two-Tiered Architecture for Offline LLM Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.806801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T10:07:06.141581Z digest=sha256:030b958e4ba37740383e12757a4f16d6cb208cf426217e3b3924b09ca6f7faf7