Pith. sign in

Paper Citation Record · LEDGER

Taming the Titans: A Survey of Efficient LLM Inference Serving

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.19720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19720 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:40.020798Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T16:37:23.046838Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b8329d69-3ab0-4e7d-9cc0-4ba684ce9bfc · inbound

Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT cites this paper.

Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:40.020798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:40.020798Z digest=sha256:958e81cd0fd504a45c22c4b9414a5662a25f2a415822c6b261e86393c4158bdd

Observation 0365177a-d57d-4ed1-87cb-4e1e494377b1 · inbound

SNLP: Layer-Parallel Inference via Structured Newton Corrections cites this paper.

SNLP: Layer-Parallel Inference via Structured Newton Corrections Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:38:16.792421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T12:36:41.055552Z digest=sha256:75dd45a3cb45509df859d762b12323a3b3107985e30b1d722528fdde949efa7d

Observation e44697e8-5558-40b7-be87-4cc4f1203112 · inbound

SNLP: Layer-Parallel Inference via Structured Newton Corrections cites this paper.

SNLP: Layer-Parallel Inference via Structured Newton Corrections Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.149925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:53:40.889001Z digest=sha256:a12b27c6654a125890cb562075cdb47daacbf8ca2364e6804af5ac1f1114ace6

Observation 6a33381b-bbf5-4aff-9cdd-32de3877eabe · inbound

Recency/Frequency Adaptive KV Caching for Large Language Model Serving cites this paper.

Recency/Frequency Adaptive KV Caching for Large Language Model Serving Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:39.033127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T13:24:44.218871Z digest=sha256:56e38681390fd2d8c24dbb66505792a8f5c390f735a008a582b11e57dd1af201

Observation c2416f8d-2cc5-4fbc-92d5-21919b12a54e · inbound

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers cites this paper.

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T13:03:39.236118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:03:39.236118Z digest=sha256:709b0a89add44b2f79063aca50d9c58f8b880bb0a518991fbb14ead3209baf88

Observation 0d37bf6c-7e9a-48db-aa6a-3820733bd029 · inbound

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers cites this paper.

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T16:21:05.570023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:21:05.570023Z digest=sha256:7ce7e2ed14904d704ef1af786bebd8071eed3c50271f6b48460d2d2ffed783be

Observation af013cce-dc30-4506-8d31-1418aaecfcf5 · inbound

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems cites this paper.

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems Taming the Titans: A Survey of Efficient LLM Inference Serving

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:37:23.048137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T16:30:09.561233Z digest=sha256:7e058bb56b8320a09d916bbb9317366c9022240d5ac377c2740c6d007b4db52c