Pith. sign in

Paper Citation Record · LEDGER

LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2404.09526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09526 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:47:58.835279Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.831808Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a683edef-25d9-4b83-8210-3fdc33d42ae9 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.537114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:a0eef4f760760c49974c96e78a914a63de0c58b001ae9fb1fa6e0f242c9824c1

Observation 27046d2a-4f07-405a-b45b-d0fc17e55721 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.094171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:891830c2c9c9e52b4339fa8d9426486e575a55ff4e854cf2ffaf9df1806845d7

Observation b3c7d940-e2f3-4f7d-a548-bd53046113a1 · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.835279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.835279Z digest=sha256:012a04efa5cb9f029ed044a609c0c5a6539540f8784389930e67610f04e8d9e7

Observation 8b265304-d579-481a-a31f-dab3b8fadf1e · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.204437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.204437Z digest=sha256:2876bb53839a3c68016b0fc18b3b5b9d2d22ecfd63bbf35627f2e8c1ef6d512d

Observation 9182bd48-f2ae-424d-a307-cd01d9864f44 · inbound

On Evaluating Performance of LLM Inference Serving Systems cites this paper.

On Evaluating Performance of LLM Inference Serving Systems LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.254256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.254256Z digest=sha256:cbca44b69963e75ba041edb147b402f167e5a6a051d636ebb3a4bceed3d6360c

Observation 78b0628f-0d5a-40f6-9387-393d7d3399e3 · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:42.984162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:310854eddfd2da6a8c4b4fb647e6c876c7f8062cbade0824e3237c80b78ffb1d

Observation 831585a0-01cd-4bcc-a8f3-36e4afa18fca · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.283971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:d314f3f71d9657dfa64693b92b75b472a56c21596d02ff1bc1df86c912631f78

Observation df7fd63f-68ca-4243-8bc5-37580f12f8be · inbound

Beyond Prediction: Tail-Aware Scheduling for LLM Inference cites this paper.

Beyond Prediction: Tail-Aware Scheduling for LLM Inference LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:57.833226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:01:34.655458Z digest=sha256:dae75e7711bb2a9844c25c8fb7e415882e68e40623099fafdef68ed783de5372