Pith. sign in

Paper Citation Record · LEDGER

M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2404.14527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.14527 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:02:12.350975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:48:11.776825Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9e43f7f6-cef1-4f01-8f69-2651a8771961 · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.350975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.350975Z digest=sha256:5391538cb70be17dc57edeaa4afef3f7dd79f5d3245d30e5a16cdf803741786c

Observation 82b08817-acbb-4dd0-8f64-f230718144d6 · inbound

An Inquiry into Datacenter TCO for LLM Inference with FP8 cites this paper.

An Inquiry into Datacenter TCO for LLM Inference with FP8 M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T16:47:23.539248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:47:23.539248Z digest=sha256:14a3519a07e6c2c5855ae4887fb91ff847a004305bb2d715bf02c163f07b8a97

Observation a53e3198-57e8-440d-903e-54aa16bc353f · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.505463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.505463Z digest=sha256:79ee5dc45874ccd13dff51ee39339b7094061873c31080d5b510c455219f2046

Observation 0875813d-2c66-4ca6-bea5-74fc4b222a31 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.853443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.853443Z digest=sha256:3965657ad418612df1d1e6899ad28ff3df33c0cd4edab27a3eca8ac31595c4c9

Observation 6a74bdda-0d75-4dff-865e-b99caede7018 · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.145048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.145048Z digest=sha256:b9b349791ca761b71877db3a5e6af99917e6c99e740b70cd768284f60f3e8f75

Observation 8e3d6901-d154-496e-8359-f7975f57057b · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.694454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.694454Z digest=sha256:011d481bec29469174707c5856e463ba512aee5c61d5c0c57fa97a9fde169090

Observation c0a2eecc-af91-47a5-88c0-6e4ffbf90283 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.819213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:82cc1d189da4ae4df9954920f72a4473c4f1651884d6f048f7cfb95fbddfe8e5

Observation 97e75627-d6ae-46f1-8472-6c7f90e7cd77 · inbound

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project cites this paper.

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:12.168226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:40:27.945478Z digest=sha256:84d477e22b58a0fc8cba3e39e1fb81a320333875724de1e585e00f37c5c80b1e

Observation 9f024742-6abb-4e2d-92a6-a634314d7a49 · inbound

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving cites this paper.

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:40:59.322868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:33:35.821999Z digest=sha256:954d4c096dc20d589dfb90763a056b516a2993ffbf95907cff3239987acabca7

Observation e3f47cc7-f059-4aed-968b-c22b5dd047f3 · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:07.099032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:5a2cea0a2a55570001157a8e5b7e6297ec708d88da24610e4b23ddf2aae3309d

Observation 2f492d09-168c-45c6-96ab-72f045356c97 · inbound

GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources cites this paper.

GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:32:43.688066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T19:31:43.015127Z digest=sha256:01a4285a36a6dbe736cf1a93e1d5921c98b54e4349778c82cbdef79b5e1bb53e

Observation 5ad6e453-c711-4d7e-ba5f-f05139539f09 · inbound

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG cites this paper.

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.302834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T02:10:57.582345Z digest=sha256:a84f1b6f613b3ab79c3ba4ef3340a0b2dddcfb9092a5c419436baeefd59cb937

Observation fcbeaa8a-9fdf-4afe-a9ae-927bd4195003 · inbound

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI cites this paper.

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.634824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:30:56.324289Z digest=sha256:2496041daa3ef10bc9cfa02ca58acb4c746c445e80927ad0460cf76b916f4810

Observation 742057a3-94f1-46c6-bc02-d221dcd12285 · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.637024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:5899893f7cf08b64aa40cf6ebe89c19d55d63cdb8615a63fa824c42b6054bc4f

Observation b493d96c-0e03-4d16-af54-e4bdde93ac33 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.778144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:7ea8e821a4eaa4b430400fcf94b37b09a64601957a1fbf112f358c2a77b9f271

Observation 40952af5-02c5-45d4-a60c-9b9150df9633 · inbound

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams cites this paper.

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.069860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:53:23.807636Z digest=sha256:fee4f45c47380102ae98c6e9dbc8ec34c037ecfcbd89e542af78469520807f9a

Observation 8f32c728-b661-4c6f-9b57-2dd4e9c074c9 · inbound

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management cites this paper.

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:52:51.924998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:52:51.924998Z digest=sha256:01448213e82dfc989bdaba6424f94e0f1f478ab36a583f6e48fa502bd50750f2

Observation 25261bff-4760-48e5-b787-0482be52aac9 · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:12.373433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:12.373433Z digest=sha256:a180ea0c94a3fa273eabaa3f9a05c5dd754c702da64bf535400246a0c3de5ee2