Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient and Practical GPU Multitasking in the Era of LLM

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2508.08448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08448 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.899552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:02:24.558075Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5cc4f5f-9dad-4df7-949f-8607647b2129 · inbound

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving cites this paper.

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:40:59.372832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:33:35.821999Z digest=sha256:ff5aafb20217cc4a5dac00b897ed04113ea8c143731deb0234d8904669059156

Observation c1565867-e009-4eb2-a196-a74bf2a558d4 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.350204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:12a7c8b080b064dfa17050b4a4642ceacdcccd81788a21a2006da55059c3274b

Observation f861fc33-b81a-48f0-ac73-82ebe9c41a0d · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.318914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:b22b5eaa2295019363516e68367d3599baa43db0e87030bee435b34a1ee14fc3

Observation e3d61e0b-d6b7-46e0-8b5e-788e6e1d3500 · inbound

Lodestar: An Online-Learning LLM Inference Router cites this paper.

Lodestar: An Online-Learning LLM Inference Router Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.559538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T16:55:38.764173Z digest=sha256:1db9ebde945567563e892b79c0ded63ffc0f8f7168546860edeb8497fa6aa23b

Observation b6bf6a03-1957-45cf-af65-9246a48b9bd8 · inbound

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving cites this paper.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.899552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.899552Z digest=sha256:78e7c27b9a8b9c7a286dc3d5bf93dee072ca240f0befa62bddcef1270f97ab10