Pith. sign in

Paper Citation Record · LEDGER

Hydragen: High-Throughput LLM Inference with Shared Prefixes

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2402.05099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05099 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:39:40.700629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:39:46.505898Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8dc0df40-83bd-4002-b1a2-e7021f3f8de3 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.176768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:75ddfe415272bf07457ac497aae901c4cf98b9f1b9c175771959fc45d9f744af

Observation f7378c09-19ef-4bdd-99d8-7d21e8faacae · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:42:23.616982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:c30a4e4faa3a564ab0e9c82c0515de5a652e4349416bc28edd30655f2aa67f7e

Observation 29ad9e26-4fd4-4c5b-bb39-9adf2a7d1edd · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.956282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:9b986d9ab9590974372022b282ba39dda508a2c2943622253f9fc836beca321f

Observation d023f8a6-482b-4e87-afbf-0cc32b1bcc64 · inbound

Auditing Prompt Caching in Language Model APIs cites this paper.

Auditing Prompt Caching in Language Model APIs Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T11:39:40.700629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:39:40.700629Z digest=sha256:92643c494de8e3a51d1e97a4d8bd722a7791255ef69413b5a26a8d92e4a9ae46

Observation 4b210226-1ba2-46f0-94b8-7912b1eae50a · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.260984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.260984Z digest=sha256:9e3e88dbb6689e3a5f32255228262bca62e501c1eed7a6426dc09de4d2ba9e20

Observation 6acb8aa0-8d0e-42d9-9f0f-43b53db11b11 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.093224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.093224Z digest=sha256:d839b80a7944e9a6a28009df18cc74f1c47c7dd1719c2b3bafaa2d5c2b0de8e8

Observation f92c6351-36f5-467d-8658-d8e1baa35229 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.723671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.723671Z digest=sha256:567848eb6285f650b30f21e58d4f736599e568970b222ddf2cbe109c9c29dcba

Observation d4bd292e-1473-4a07-954c-7d0590f9378e · inbound

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions cites this paper.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.050106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.050106Z digest=sha256:3c75a9002e7519a087dd85054ad3389d2bbfcb0f37736e5658ff6294c2b2ba3b

Observation b29408a1-2498-4e35-9655-7eebfbf5dd8d · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.784846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:fb5cb6c9b8e2ab3c52e25788d139d827da71253d407e30b0e7b71192ce56f795

Observation 871fe2b3-6d42-4b7d-9d84-4e326f4e8d49 · inbound

SEMA-SQL: Beyond Traditional Relational Querying with Large Language Models cites this paper.

SEMA-SQL: Beyond Traditional Relational Querying with Large Language Models Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:17.460759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T05:11:46.991452Z digest=sha256:776165ee8b21be37cef8a74b5c726ecdd19df3e8912d1faa747fcdaf62675e34

Observation 215da78f-f1e6-4f31-bf4d-dd1394a001fc · inbound

SEMA-SQL: Beyond Traditional Relational Querying with Large Language Models cites this paper.

SEMA-SQL: Beyond Traditional Relational Querying with Large Language Models Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.559212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:53:11.604043Z digest=sha256:b6d82986a296811ab544aff33b67297b608b97fa5ba2ea74ed30e42f1cffbb23

Observation cdedbf31-6a5a-4d3f-bd19-7c5b09b0f9e7 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:27.025286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:29:28.916831Z digest=sha256:8c784a8dfb5e1111768444a7500628be65974063630567d6bc68c1907478053e

Observation bf46b8a7-09f8-4d97-b749-180b2cc8fe59 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:47:41.948204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:43:56.767714Z digest=sha256:f242a5274e0a377c3b07742cbe24551ffc0b3a6caa45eb9963d6dcba68c7c4ff

Observation f1d616d3-9ae6-4017-9f98-73d82614dd82 · inbound

Towards Distributed Inference of LLMs on a P2P Network cites this paper.

Towards Distributed Inference of LLMs on a P2P Network Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:15:07.948549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T23:13:57.705287Z digest=sha256:479af491cbffade79c1eb0dc6b6726ab23a3574ac4bc36d3d4147b660c5b178d

Observation 2da98e4e-03c3-407d-82ff-7497334a5223 · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.915609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:36e4032120c516736734158c8fb566a0847f680c123af570b9866934a7e3f98f

Observation 1f20be0c-1054-4a50-b7b3-0d28561a3184 · inbound

One Generator, Any Process: LLM-Conditioning for the LHC cites this paper.

One Generator, Any Process: LLM-Conditioning for the LHC Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:46.507262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T07:53:57.250401Z digest=sha256:7634896d217380740e0cc76beeb56b334a58ac1473fd4a6acb1143a493f53b73

Observation 1157fc1f-5e5f-40d0-85bc-7aee4a5b97ca · inbound

One Generator, Any Process: LLM-Conditioning for the LHC cites this paper.

One Generator, Any Process: LLM-Conditioning for the LHC Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:14:36.110959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T10:13:09.503522Z digest=sha256:fa1715bfe4f55dc69916605dcf204372926c9d5ec00c587939fd5d6ea7b69f61

Observation ec8e962e-f30c-442b-9bf0-228e2aebb6a8 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:83feee948b8b910be1f1dab62b5b17492d079792d7012d57c4efd0ecc2f4c144