Pith. sign in

Paper Citation Record · LEDGER

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

As of 20 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.06709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06709 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:08.498820Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:45:32.847427Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved5
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 70838d38-4447-49df-a558-80af7a61cc6b · outbound

This paper cites Brown, B.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Brown, B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.653172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.213684Z digest=sha256:20720e8582667fcbb95ad85bf7db96ebfd9c07e3e6ae329bfe15896066bc0661

Observation e9b3d3ab-052a-46ce-b479-6114c41f6f03 · outbound

This paper cites GPT-4 Technical Report.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.221119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.221119Z digest=sha256:a54a0ee6907126d53a520f25e3c23baf964ed2db0582a5dcda005adbf036f1c2

Observation 2642f22a-d560-412d-a3b7-041a69b91997 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.227140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.227140Z digest=sha256:6d2650c69ac424be9b1f55891e15e4a37e8a9d4f7a7d9372d9a45edf3ac7542b

Observation c8b27bae-ca69-4c32-aab8-045993c3a9e1 · outbound

This paper cites Characterization of large language model development in the datacenter,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Characterization of large language model development in the datacenter,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.623471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.234544Z digest=sha256:a7a3535f9416da2e65eea080062ca4d56e3b428cd2e7c3ce86b693daadb33bfd

Observation a9b7d5f8-4f80-4113-9703-d45e9911c44a · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Orca: A distributed serving system for Transformer-Based generative models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.600018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.240643Z digest=sha256:d61d08b78f0089dc2fdb0857e8ac2808b5161937b847bfd86b29eed45ae96061

Observation 5a53d823-1a65-4e4a-859f-72126ad12211 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Splitwise: Efficient generative llm inference using phase splitting,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.576977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.248315Z digest=sha256:e6eaca4fd68518ac43884ac50d6ca689273d0b80ca201ca01b0954313cb42c54

Observation d3401967-7f6b-4079-82e4-878eb18686a4 · outbound

This paper cites Optimus: Warming serverless ml inference via inter-function model transformation,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Optimus: Warming serverless ml inference via inter-function model transformation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.552481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.254850Z digest=sha256:e1a4d439bf6c0b8b7b1cf81240a4eb0f9a6d289984eae10ba970fd8032660843

Observation 995dcd7f-b0f1-4a81-8ef3-f9f7b29ef31b · outbound

This paper cites Otas: An elastic transformer serving system via token adaptation,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Otas: An elastic transformer serving system via token adaptation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.528062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.262315Z digest=sha256:cbb1e25216ce175e62fd67c68796bd69dddae6f716055daf60aa76693324ccbe

Observation c51ead8a-cd79-482f-9ad7-581802a112da · outbound

This paper cites Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.507764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.267931Z digest=sha256:ca5bccfde725e40f655c011ae9e8544870b1cff886d159bf2527cad748f85236

Observation 49e0f30b-0397-499f-9da2-81df2420cc4a · outbound

This paper cites Efficiently scaling transformer inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficiently scaling transformer inference,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.484082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.276900Z digest=sha256:1dacc968733bb64d2147ca8086e1339812adeff9cc17ee42e749b00b4d901d87

Observation b98744f7-ea30-4b0f-8c29-375784e0e37e · outbound

This paper cites LongloRA: Efficient fine-tuning of long-context large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LongloRA: Efficient fine-tuning of long-context large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.461417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.287438Z digest=sha256:62669ba6f0247a306ae599affd2e7054c46dd44a698692c7d7f8c268ac0e8680

Observation c4a6f511-e017-42b4-aa5f-014cac17659c · outbound

This paper cites Efficient memory management for large lan- guage model serving with pagedattention,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient memory management for large lan- guage model serving with pagedattention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.433817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.293309Z digest=sha256:138ff4e80c978eaff946803c12e7cc989c565eb2cfbaeda9f83e60211f6b2e99

Observation 36fe6834-c623-4a4a-b6ee-baf3aa39282c · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.408044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.298553Z digest=sha256:c099e2b9dcbb7f9b250b022f0eef256afd245f98ced4d9249279ffcbb8203c9b

Observation d00bb195-ef32-4aaa-85a2-30ee35428d73 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.386345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.305038Z digest=sha256:3847cbdad2b11be9f07b9c1b4364053f44d7bc56f76fa3f795de9196dac275ce

Observation f871c23b-05d6-4ad8-9056-f3344e591c07 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient streaming language models with attention sinks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.359107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.312308Z digest=sha256:cfae5fcaecf09c582e5777c7815b2cd2fab1f4341c43eb9afef916b20ce09d59

Observation c370e43a-eb65-47c6-908d-b817162cd06d · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.317601Z digest=sha256:78a11e7d9101d6b296cdbf5594b6df181a151c926696f7c1d45c50168b9aa819

Observation f6791e28-2a64-42b7-901d-bd756b1bba73 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.323613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.323613Z digest=sha256:fd47557c4aef5f3da280aef520f60819605d26d72332a6555dc6740ac0a1ad4f

Observation d7d8c44a-c3e6-421b-9e35-83e93800a937 · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLMs,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Model tells you what to discard: Adaptive KV cache compression for LLMs,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.320919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.335736Z digest=sha256:e1fc155f4c85c352e81a75ca420a6aba41f286e2771e7ee1f798f938ef5708ee

Observation 92d5fa76-fd6b-45b5-8513-c5cddeeb8fac · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Flexgen: high-throughput generative inference of large language models with a single gpu,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.289777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.342949Z digest=sha256:b0124003e0b937ed37a360c1c6ccfd59daa565a896ca5d2e0b2406f41c094016

Observation 21a3dae1-09ea-4da9-9f75-686ecf9d61e7 · outbound

This paper cites Infinigen: Efficient generative in- ference of large language models with dynamic kv cache management,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Infinigen: Efficient generative in- ference of large language models with dynamic kv cache management,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.265498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.349106Z digest=sha256:eedb434ff14a571473c4837aed125df6f72a4fdfccda5d40611ebad13f314a5f

Observation 1668b1ff-fff7-482d-9975-177f6099fe7d · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.241225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.356264Z digest=sha256:1c741b1417ff503a95aad9c02c78fbc7fab2bc69a5b6ed9a09c5397bc1ff8615

Observation 17a3a078-f1d3-4e2a-a5b9-b39cccd494d5 · outbound

This paper cites Deepspeed- inference: enabling efficient inference of transformer models at unprece- dented scale,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Deepspeed- inference: enabling efficient inference of transformer models at unprece- dented scale,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.218362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.363065Z digest=sha256:b0f206d474655c3d3345b0f41a7f2d37912b99cd3038cc4f9a06061141d8f365

Observation e7f660de-3832-4453-8725-a33dbb518132 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Llumnix: Dynamic scheduling for large language model serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.193397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.369100Z digest=sha256:077e52ce80a281bdce5ebb418884e8d63753b8119db8d5e8f64f226f8a42f524

Observation 9057a506-236a-438c-90e5-4d64452995a6 · outbound

This paper cites Serverlessllm: Low-latency serverless inference for large language models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Serverlessllm: Low-latency serverless inference for large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.173370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.374194Z digest=sha256:8409742f50cf732eb5c4d4f2d061ab62f54015c6f8b38561f11117b85852b1fa

Observation 19df26f4-727a-4b8a-b54f-a3a7326ff70c · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Turbotransformers: an efficient gpu serving system for transformer models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.146756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.379805Z digest=sha256:4eb108b633df0e71d204bd82ea2972cc3d4b0230ecfb23c77cf989e2629d837c

Observation 7e83071a-3a68-4526-94d9-b8d3d42306be · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.122150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.385167Z digest=sha256:8fefacceeedcd33c185ccf5cf116e62b0f4c75a37cf18016fab343f2b97cb5b5

Observation 3cc9f462-3e53-4d19-876a-a09fa998e019 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.086283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.394284Z digest=sha256:c8eb7d477a99dd09b4d74b79f817460bb67409b93ae7189b3f2ba465b375440c

Observation a9042608-ac4d-4cde-97b3-35f7cd5fc350 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.401041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.401041Z digest=sha256:18478dc48aaea2c819cdd892a110ea9099cbb78b75514dad9c23b041a1ef8f73

Observation 3405f68f-6281-44eb-ba11-2d40878f4747 · outbound

This paper cites dLoRA: Dynam- ically orchestrating requests and adapters for LoRA LLM serving,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management dLoRA: Dynam- ically orchestrating requests and adapters for LoRA LLM serving,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.061030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.406766Z digest=sha256:d71fd8d7cf156868f37ae62fb2b10f2c09005e688f194d4d8f37889fcc4b52fa

Observation 7546fca3-7b42-40bd-9489-c5b6ea9e92ce · outbound

This paper cites LMSYS- chat-1m: A large-scale real-world LLM conversation dataset,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LMSYS- chat-1m: A large-scale real-world LLM conversation dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.033618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.411772Z digest=sha256:ff819de117446ad802c9a9cb46066cc5265d2b04581a96cb87e973f49b38ca50

Observation 2f10efc9-75f1-4584-8455-1b3a8a0032fb · outbound

This paper cites Wildchat: 1m chatGPT interaction logs in the wild,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Wildchat: 1m chatGPT interaction logs in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:09.014848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.417714Z digest=sha256:c432222a82cc49ad971195357dbd28fb635f67bf8c332d427c403733f90b709b

Observation 207f5195-686a-466c-a0d1-ae5395559b42 · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Judging LLM-as-a-judge with MT-bench and chatbot arena,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.995713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.423055Z digest=sha256:cc3b3ed86b7c311a5b54e97fc6de52cc083ce81719e2699218ee3957a1cf809d

Observation 7b768cf5-06c4-4aab-8dc5-4266b15924aa · outbound

This paper cites Koala: A dialogue model for academic research,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Koala: A dialogue model for academic research,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:08.429322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.429322Z digest=sha256:121e1470b44cebfc7752c313dbe71e615554fedc5f1148f6673b35b7378567ba

Observation 37831d52-d151-4bfd-b0d0-10ecd0a38cdc · outbound

This paper cites Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.965068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.435506Z digest=sha256:da3650a4d2589de23ae8c4dcc43cb315dfe8ef33f82cdfe66963ad306c3b2142

Observation 53e40a07-cfdc-4642-8716-f2bc105a4cce · outbound

This paper cites an unresolved cited work.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Unresolved cited work

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:58:08.946688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.442068Z digest=sha256:b1ef65f687a1090136466fe2483ec07c59248f615b8f02edb53a7847f57db27a

Observation f145516f-1315-48cf-b63d-90fa4948dcbb · outbound

This paper cites Adaptive resource provi- sioning for the cloud using online bin packing,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Adaptive resource provi- sioning for the cloud using online bin packing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.928833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.448959Z digest=sha256:e25ec7a767b2479c57f1e3d3b020d32ba389eda9ccd410d5248b76a9a375418d

Observation 4abe209a-994f-41df-9d29-5bf5a9aac970 · outbound

This paper cites Efficient online strategies for renting servers in the cloud,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient online strategies for renting servers in the cloud,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.908341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.454685Z digest=sha256:385d6383109e5fdd9938a1c75fb653b3927bfaa77c72e5e2097560a270d7e9df

Observation 71d93163-586b-4e8f-8e20-fef44de7b11d · outbound

This paper cites Powernap: eliminating server idle power,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Powernap: eliminating server idle power,

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-10T20:58:08.460910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:08.460910Z digest=sha256:a2bb1d7092c800be2c568abc661c59935a1c6265d4fcced0de131d8143144508

Observation 99c54bf0-4ef5-498b-9932-3f39f792bc10 · outbound

This paper cites Easy, fast, and cheap llm serving for everyone,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Easy, fast, and cheap llm serving for everyone,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.885991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.466106Z digest=sha256:398efb90b05015a94781dc397a2b3c0a314e4d907135bd73022942f748dbe20b

Observation e58c2e49-e9eb-416a-a8f6-336b99cdc3ec · outbound

This paper cites Ray: a unified framework for scaling ai and python applications,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Ray: a unified framework for scaling ai and python applications,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.859845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.474077Z digest=sha256:fc2ba78c5b1cadb7ed1542faeea3213866554ac138088e896419cfcbc73869c2

Observation 8ecb158a-0e6c-4f84-a2c1-6053038575de · outbound

This paper cites Gloo: Collective communications library with various primitives for multi-machine training,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Gloo: Collective communications library with various primitives for multi-machine training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.810006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.486606Z digest=sha256:4a9f53da4c1014e98385f1324ee529d362c345ce1cceb3a20b476432b11923a3

Observation a36461d3-72c8-4209-b669-8f95285db950 · outbound

This paper cites Openai platform document,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Openai platform document,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.790533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.491510Z digest=sha256:e193cdf0dce9418a12993cec24ea5b52f58fec1e3b531a0242553bec76507b51

Observation 8f967afb-874f-4d4b-957e-8dee2255b2e4 · outbound

This paper cites Anthropic platform document,.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Anthropic platform document,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.769470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.498820Z digest=sha256:bf4c3e7dd6f5c7a42b94c83791f3828ea2acee66c0fdeb6ec884107ec4e23b3c

Observation 4a2d838e-b267-4d3b-8d9e-a7e1500860a8 · outbound

This paper cites Available: https://github.com/ray-project/ray.

Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Available: https://github.com/ray-project/ray

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:08.836802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:58:08.479825Z digest=sha256:c90f334a6ef2bd3565ba97fa3e19c6f96f3aec61c682d169d2cbf42b23fcf0ec

Pith citing papers

Observation 01ad7691-2c93-438b-9311-2acf41d527b1 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 260

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.523271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.523271Z digest=sha256:adcd4fba053c763c27c6dcd224e3aeb196fad9607760ff41240d9aa1fd1a5d74

Observation 80459404-6387-4874-b2a6-aef638b43d33 · inbound

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning cites this paper.

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:45:32.847427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:45:32.847427Z digest=sha256:ac8228605212aad99cff26a358e6670f0a9f269b75822ee6399e82ef0c92a649

Observation 0b31f61f-9a83-4899-bc23-8de4f04438e7 · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:46.874022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:b9732d945d47f48857d1b377dd716ffc9f89cd2c2052ea339dab4c3b1ca18f6f

Observation 87c72b18-1423-4970-92e9-e9b572cdfcb1 · inbound

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs cites this paper.

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:55:56.631328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:55:56.631328Z digest=sha256:3cbcb0786d71e645aa02911265a592896be516245e6c9fd0dbc36942039c6dab