Pith. sign in

Paper Citation Record · LEDGER

ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.01228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01228 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:40:57.624458Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:27:06.032587Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 605cb783-6603-4532-b2bd-4b797d720249 · inbound

IC-Cache: Efficient Large Language Model Serving via In-context Caching cites this paper.

IC-Cache: Efficient Large Language Model Serving via In-context Caching ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.376047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.376047Z digest=sha256:df9e01affeba742651253b6004d3ff8d50fb8b75b8f3c6cac9933f1cfb217fcc

Observation 7e4aa127-30ae-446a-853a-4649aeb001e4 · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.461110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.461110Z digest=sha256:76f907449f91fe321de986770b2f65e011cb995e3a4f02d8af39c629a9e7a8d6

Observation 62162d4d-2118-4406-a6b9-2c6519d98415 · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:14d33a840df93315edc014dac2ed249c665f16cacc79da75c4445437e2fb447b

Observation 0f833576-156b-48df-8ba7-51561d1fe526 · inbound

The Energy Cost of Execution-Idle in GPU Clusters cites this paper.

The Energy Cost of Execution-Idle in GPU Clusters ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.593039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:04:25.951890Z digest=sha256:69214ca49ecf118b61f163e1891d29eae358bfcffdb24c8cda444710a9ad15ef

Observation 8023bd15-7971-4431-81cc-67daf1548ef8 · inbound

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start cites this paper.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.482759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:df3231b52d5caabbfe48656df31b61aadb89f6579cff3b3c8f82c14be23ee160

Observation 2f976b28-39a3-4f1e-b366-726fe491f225 · inbound

Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC cites this paper.

Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.238853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:01:57.651581Z digest=sha256:87251d2f21856984948d75a99af9e7c16f06a20f2d03a59d4f9667b9f25f449f

Observation 8bff6817-1084-4762-9086-23cd2ab78642 · inbound

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate cites this paper.

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:31:00.663351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:05:08.438659Z digest=sha256:2f78fd349c4bb713f9aaf405b37d37e346a52762f3cfdda32c8fa31ef7709c2e

Observation 3b16e5a1-23ba-4061-8a41-381ba605d4ce · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.389848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:72e7cc5e85c398abf6644c9fe7624e14382c769b9e5a427c7e15c45e7fb2fd97

Observation dfbeedee-3cf8-47fd-8389-fc1a6f40e3ee · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.282201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:c437b6a0a2a8a144d99405f7a038efa3de761f8a0b1dcac7373ea1700df7f91d

Observation f51cf09a-edc0-4789-a698-73f684fd32a0 · inbound

RW-TTT: Batched Serving for Request-Owned Test-Time Training State cites this paper.

RW-TTT: Batched Serving for Request-Owned Test-Time Training State ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:53:28.852555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:46:48.954003Z digest=sha256:26f2f2ebcedecd0a620ebc3f02ff1588093d1e6c0cd6edae5d1873efd511160b

Observation 63db9f70-be66-4694-a3be-7087fb5f7f77 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:27:06.034079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:8a1520d58d0f31a03f3ab65f5a68689c780cbec9d117f49af1b00a5b919f08c7

Observation e461ba60-53a1-4a7c-81f8-699586e97199 · inbound

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters cites this paper.

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:40:57.624458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:40:57.624458Z digest=sha256:40758257161af54f8518fc3ac48525f3faf1642d51fb70542b651710c9878bf3