Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2404.14527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:36.145048Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T12:48:11.776825Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6a74bdda-0d75-4dff-865e-b99caede7018 · inbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3d6901-d154-496e-8359-f7975f57057b · inbound
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a2eecc-af91-47a5-88c0-6e4ffbf90283 · inbound
Efficient Remote KV Cache Reuse with GPU-native Video Codec M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97e75627-d6ae-46f1-8472-6c7f90e7cd77 · inbound
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f024742-6abb-4e2d-92a6-a634314d7a49 · inbound
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e3f47cc7-f059-4aed-968b-c22b5dd047f3 · inbound
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f492d09-168c-45c6-96ab-72f045356c97 · inbound
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ad6e453-c711-4d7e-ba5f-f05139539f09 · inbound
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcbeaa8a-9fdf-4afe-a9ae-927bd4195003 · inbound
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 742057a3-94f1-46c6-bc02-d221dcd12285 · inbound
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b493d96c-0e03-4d16-af54-e4bdde93ac33 · inbound
Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40952af5-02c5-45d4-a60c-9b9150df9633 · inbound
Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f32c728-b661-4c6f-9b57-2dd4e9c074c9 · inbound
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25261bff-4760-48e5-b787-0482be52aac9 · inbound
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.