Pith. sign in

Paper Citation Record · LEDGER

M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2404.14527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.14527 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:36.145048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:48:11.776825Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a74bdda-0d75-4dff-865e-b99caede7018 · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.145048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.145048Z digest=sha256:74006508d15106c8c44217d1e07229c946da68ce6b555354392cbb94a4a15dd3

Observation 8e3d6901-d154-496e-8359-f7975f57057b · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.694454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.694454Z digest=sha256:011d481bec29469174707c5856e463ba512aee5c61d5c0c57fa97a9fde169090

Observation c0a2eecc-af91-47a5-88c0-6e4ffbf90283 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.819213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:39ea2adf4e02f4b9156bf429949737a51976672fe165222d5d39acd4f505b545

Observation 97e75627-d6ae-46f1-8472-6c7f90e7cd77 · inbound

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project cites this paper.

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:12.168226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:40:27.945478Z digest=sha256:5301446a1a6593bd3de744b9db05656faf4741a3ee6da38d0f9fec529e13fd84

Observation 9f024742-6abb-4e2d-92a6-a634314d7a49 · inbound

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving cites this paper.

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:40:59.322868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:33:35.821999Z digest=sha256:6682791d9c2eaeb4ffde4a0c515e9abedc1ded51d621f857e7bfe54ce294ad4e

Observation e3f47cc7-f059-4aed-968b-c22b5dd047f3 · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:07.099032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:5f29420e1d887a1ebfa8b858557135a6961499b6b5cf3fdaaab4893fe469d606

Observation 2f492d09-168c-45c6-96ab-72f045356c97 · inbound

GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources cites this paper.

GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:32:43.688066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T19:31:43.015127Z digest=sha256:40ae2180db7d00eaa0cdbe504b32ba13a648df61775225f20b3f71af3fa7d617

Observation 5ad6e453-c711-4d7e-ba5f-f05139539f09 · inbound

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG cites this paper.

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.302834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T02:10:57.582345Z digest=sha256:49e35a2e0292ffd4e97ad27058798c51910b8d0022649077ec6afe45933b5e41

Observation fcbeaa8a-9fdf-4afe-a9ae-927bd4195003 · inbound

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI cites this paper.

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.634824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:30:56.324289Z digest=sha256:d1569a1558a8844f8b2de33cd3c9a552e04c651c124ff09d7160f817aadab6db

Observation 742057a3-94f1-46c6-bc02-d221dcd12285 · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.637024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:e4ac21e74a2eabfac17a372139d718a900ee76527f0e329bf9b2a7d701ffe81a

Observation b493d96c-0e03-4d16-af54-e4bdde93ac33 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.778144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:c0bf42fa92eef52f414eccfe0ddf197fa685c59727b77640462b618bc4fa79bb

Observation 40952af5-02c5-45d4-a60c-9b9150df9633 · inbound

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams cites this paper.

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.069860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T23:53:23.807636Z digest=sha256:8a45c18108379c843369c68af9769af0723acb95773ea125de876d4ef2ee2d75

Observation 8f32c728-b661-4c6f-9b57-2dd4e9c074c9 · inbound

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management cites this paper.

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:52:51.924998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:52:51.924998Z digest=sha256:341de980d861ee1a4999b5956ccb54011dc44435d333d6a52516c9dd5e057607

Observation 25261bff-4760-48e5-b787-0482be52aac9 · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:12.373433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:12.373433Z digest=sha256:a845106aa00f2eebe38629ca58a76fd2be51ed3b8a7e3a8f07304f22e14ab416