Pith. sign in

Paper Citation Record · LEDGER

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 6 inbound Pith citation observations for arXiv:2505.03756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03756 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:01.881688Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:47:00.295144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T22:05:06.251185Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved34
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d77a65af-c5f5-4158-8fbc-7d35818c28b1 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.692789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.692789Z digest=sha256:8b7db709c28f40cfd098ccec8bd49042047ab4ddd84bfa100d758b6971f715b8

Observation f34dadd4-cb6a-456b-a374-2f65acd35538 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.697534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.697534Z digest=sha256:6c2b38e0582092f50cb92b2ce9004278d1ab18c674e28f66fbac164cdc864729

Observation 70718582-7ee3-4eda-a965-d2cc5ed84ee7 · outbound

This paper cites Introducing apple’s on-device and server foun- dation models, 2025.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Introducing apple’s on-device and server foun- dation models, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.538980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.701797Z digest=sha256:1a3be30690eaa44faf54416a4a500f22bd53d2dcad80ebe532fae9fe26c96e09

Observation f4768bd2-f376-46a5-a399-07b5d37ae985 · outbound

This paper cites The Costly Dilemma: Generalization, Evaluation and Cost-Optimal Deployment of Large Language Models.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Costly Dilemma: Generalization, Evaluation and Cost-Optimal Deployment of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.705679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.705679Z digest=sha256:0a1a996c359677c50b607dc0595f72b1a1729ede1c05fdf9f9006aa66e3f9cc2

Observation a72d59f2-eeb1-45fc-b5cb-bfa4892ebc69 · outbound

This paper cites Taskmaster-1:toward a realistic and diverse dialog dataset.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Taskmaster-1:toward a realistic and diverse dialog dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.528286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.709512Z digest=sha256:8442bf034346abdc34a213d9bb57709c0c64a3eafc25a5c230eda5d331b45deb

Observation 1410a725-0eb3-4e76-8d83-afebd3df86f7 · outbound

This paper cites Punica: Multi-tenant lora serving.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Punica: Multi-tenant lora serving

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.517409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.713583Z digest=sha256:b2f61a3c4293736063f1398a63b39516778bcd8fe3d364d8c41a3cd9bea36017

Observation 4c54425e-6112-41a9-9e66-d24d841ff385 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.717527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.717527Z digest=sha256:ba160db6f7d775b76e92e9f3174ca3cdfbdd91cdf2220de7b44c78517c28faf7

Observation 48f314f6-b18b-4715-9eae-df39be8b22a8 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Palm: Scaling language modeling with pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.721637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.721637Z digest=sha256:19153c53299898dc3eb626635eab2a72eab560a7f126bddf4087874fbf386c19

Observation e5c3a67f-f606-4125-99ec-2ff937571ba5 · outbound

This paper cites Introduction to tpus, 2023.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Introduction to tpus, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.498115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.724828Z digest=sha256:c08b9f4d859fb26e4e408c7f17334598e9f3defbd5fed13bfe0aeff879b2a5d9

Observation 2b3a5b57-94c8-458f-8990-c692ddf520e8 · outbound

This paper cites sglang: A fast serving framework for large language models and vision language models., 2024.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management sglang: A fast serving framework for large language models and vision language models., 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.486403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.728100Z digest=sha256:f188620b4cf92d2d1e8af632433d67b99ade8b1b0c5ead2ab1314fdc920978ec

Observation 2fe70934-7afc-4584-a05b-49a4a7ce9db3 · outbound

This paper cites High bandwidth memory, 2023.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management High bandwidth memory, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.475515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.731413Z digest=sha256:c475157e09c19b18cfcb1b450ff550724b2257974464082d810bb470d70000ee

Observation edc0a231-ccfd-4d35-84b1-a2034f778590 · outbound

This paper cites Trie, 2023.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Trie, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.464723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.734605Z digest=sha256:bb193961044acd3fa89baa9d68a568d44cd27d8bfa7123e406b5ed7be2416f7d

Observation 8e80c258-b45c-41fb-89d2-9f22bce18848 · outbound

This paper cites Nvidia a100 tensor core gpu, 2024.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Nvidia a100 tensor core gpu, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.454006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.738280Z digest=sha256:18c3dd302d7415f3201bea4c434862cc414f6d373461b56500f849589dba1dcc

Observation 4167d71a-fbd4-4219-9538-4efcdb01e611 · outbound

This paper cites Qlora: Efficient finetuning of quan- tized llms.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Qlora: Efficient finetuning of quan- tized llms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.741620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.741620Z digest=sha256:e6194abfe54c46ebb83ac3ca2ac03c7d2a7915531770383b6eaa7581d36bc116

Observation f39b5a8a-f1b8-4a92-a694-5d52b6a1857c · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Gpt-3: Its nature, scope, limits, and consequences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.744787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.744787Z digest=sha256:c4ee500045bbd92df10455700e0b260ba1e4cf63cf3cceb463442e1bb27a4a52

Observation 5eb411b4-7279-4a1f-9c5d-93c1778190cd · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.748120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.748120Z digest=sha256:cf9773a25a2756742cc3d1280739ec5ab7e94774676bcf17540ac49d2b2e54a0

Observation b9486e9e-d77f-4486-85da-e34bb823c605 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Prompt cache: Modular attention reuse for low-latency inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.428587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.751777Z digest=sha256:6b8db42eb74530ddd3bfeb08227036fc0f4db80e0a7214d0eb76f28e241c59c9

Observation ae4ccf8e-dd1a-4b43-a95e-5831a7b1955a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.755128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.755128Z digest=sha256:82a6fd4ea293382d9d8376f670710ab953558eb191ba04527ba607fec02d7cbe

Observation aa2c5c4d-d506-407e-adc0-556294f77415 · outbound

This paper cites LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.758544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.758544Z digest=sha256:2853503d29c42e1883266c2a8de114b12dbbeee48ffb269203278530c93c1f6c

Observation e365f6ed-7408-49b3-a184-184c30990a42 · outbound

This paper cites Chameleon: Adaptive caching and scheduling for many- adapter llm inference environments.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chameleon: Adaptive caching and scheduling for many- adapter llm inference environments

Reference 21

Resolution
verified exact
raw_fallback, observed 2026-08-16T11:57:02.215502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.762195Z digest=sha256:937966d1ffa5874569796bb52315378b99db37c9787a6a431bcb01c17e1168df

Observation 3c5753f6-b4ac-4728-bc25-464a14aa2bc9 · outbound

This paper cites Efficient memory man- agement for large language model serving with page- dattention.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Efficient memory man- agement for large language model serving with page- dattention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.765927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.765927Z digest=sha256:5a06f9f21e21704d6b71d8850c625897296c738e6d063282e79a2b96ed40c945

Observation 6cc91c4d-7915-4744-8885-301a44c8a8af · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.769214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.769214Z digest=sha256:6c9ab6baeab6cade7a631402f8c38264b3ae2076630344bc45c756f6f002c3c4

Observation 28475c05-f47c-42e4-969c-d3453d406675 · outbound

This paper cites CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.772247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.772247Z digest=sha256:b3e58d8343260f70d4801345bebcec4a3d8a67e0484fbbe2288ae75c5b994af1

Observation 7fff902a-b521-4faa-a448-88f410897a0b · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.775659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.775659Z digest=sha256:986418a2d62fd665b9e82854d40a56081e28e33605b86577b86c5cec7d1b727f

Observation 13b5eff5-021e-4f5a-9f8e-6772002df314 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.779138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.779138Z digest=sha256:b0c58f6c376f01d908cb2923035d41298bf146ed3fe675e333947349900879b1

Observation eb81c90b-306d-477c-9f2e-4d716beb127a · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.782606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.782606Z digest=sha256:2d90765522f73bb99ed8edc64b11a9d53714ce3d7826c8c21e346dfadedd92d5

Observation f0398a13-59c0-4c8b-be41-69a2b40332ce · outbound

This paper cites Instruct-tune llama on consumer hardware using alpaca-lora, 2023.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Instruct-tune llama on consumer hardware using alpaca-lora, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.411273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.786129Z digest=sha256:43027f5079be4ce06599d0e65915123ef9e92fc70377df075361d2a19867dc32

Observation eef52f8a-75f9-4197-9171-bbebf9c66ade · outbound

This paper cites Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.789502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.789502Z digest=sha256:66ddd116b836c77c2f7ee90c4d77337e760864ce6ee8492cb6051e10c5849a02

Observation 26d9b191-8666-40f4-9cc5-16cb73f88a6b · outbound

This paper cites Language Models are Few-Shot Learners.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Language Models are Few-Shot Learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.793126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.793126Z digest=sha256:118f1c8ee673609603b939aaa2b43f3a4cc991053752c6352faf729708d25328

Observation 3f6700a8-b82d-4592-a1b5-7e9fd585d7b7 · outbound

This paper cites The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:57:02.043357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.796567Z digest=sha256:7fe2a63726b03a5681d8d772600b59a78ff314e3cb248596533f4e0a64bc1120

Observation 1a54de16-150e-42eb-a09b-f09f61482d29 · outbound

This paper cites Chatgpt, 2020.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chatgpt, 2020

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T11:57:02.400603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.800176Z digest=sha256:2dfc70b6b851400899abf88d6bc11a87b72662a7239bd4a3c88cb79944010f56

Observation e426f054-1e11-4d0a-a0cb-c532dd4c1ebf · outbound

This paper cites torch.stream — pytorch 2.0.1 documentation, 2023.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management torch.stream — pytorch 2.0.1 documentation, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.389230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.803652Z digest=sha256:1c712b71a8ac030fb6a5570dd74b3d847ccb096ca73e74fde6f7bb69ed062598

Observation bd27a102-1e9d-4a79-bff1-db2cc4a90ac6 · outbound

This paper cites Moon- cake: Kimi’s kvcache-centric architecture for llm serv- ing.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Moon- cake: Kimi’s kvcache-centric architecture for llm serv- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.377850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.808007Z digest=sha256:cf3ee41d6a1fd9c3c131d4117a7fc8847820360dee8a38d872c7cfb36b7b35af

Observation e806169b-dbc6-4605-81f9-914cbc8931d9 · outbound

This paper cites Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.366647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.811318Z digest=sha256:bbe8f2415d578217d9e8599b2983d100cf3043c9928ed7d3f3a13c7ee0b232f0

Observation 53bc2ed7-fd1f-43e7-ac63-079b8163d743 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.815067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.815067Z digest=sha256:f35a79df2525c20d225da10c19d25d298c2cc4b0c1ab84a3aa1d78e096615290

Observation f2197089-74ee-4b9d-8be4-289cfa67e394 · outbound

This paper cites Slora: Scalable serving of thousands of lora adapters.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Slora: Scalable serving of thousands of lora adapters

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.355674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.818586Z digest=sha256:3824ae7c851f7e82849282e71f830acd418325434d9ffbb83148a463df68a562

Observation 29180fcf-06f0-49f7-b4fb-cbbf98c2855d · outbound

This paper cites Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.822147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.822147Z digest=sha256:2c6f3b2ac678a9f7a65d00b51da6f78185bd871ac8b73fe5f71a879459c6d85b

Observation 6bfd6808-f4ef-406d-9234-86793841d94b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.825709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.825709Z digest=sha256:e817a346d220092b51d05fe9b94ecb261cdd6706176a13d70aaab5e69b5c86a0

Observation de06853e-ebe8-446f-b414-d009d559bd45 · outbound

This paper cites Discovering finance keywords via continuous-space lan- guage models.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Discovering finance keywords via continuous-space lan- guage models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.344777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.828982Z digest=sha256:39243c2e8feb0f2c8f2e2609f295dd7299d8088de51bd881f6ea8f2530e4e3a6

Observation 71c9a847-92c9-4836-9e38-d7dcd8350fd1 · outbound

This paper cites vllm: A high-throughput and memory-efficient inference and serving engine for llms.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management vllm: A high-throughput and memory-efficient inference and serving engine for llms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.333594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.832167Z digest=sha256:2f31a5849c3535d2a5d9629d8ac124b4f444cee40764b8be151092b35e794be2

Observation b0b5f971-9565-49b9-a289-803c287750b0 · outbound

This paper cites LoRA-Pro: Are Low-Rank Adapters Properly Optimized?.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA-Pro: Are Low-Rank Adapters Properly Optimized?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.835511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.835511Z digest=sha256:61bd1b5a48dd1cd9c9aaea3d38e751265d39f692e1b21a5bd2b6de696e5962c0

Observation c4704b4f-dfea-4f00-bff9-862ee3e83cee · outbound

This paper cites {dLoRA}: Dynamically or- chestrating requests and adapters for{LoRA}{LLM} serving.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management {dLoRA}: Dynamically or- chestrating requests and adapters for{LoRA}{LLM} serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:57:02.321911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:57:01.839142Z digest=sha256:8663faf9551398e545e76fa26907cb95cdb4580bdd7d7f153b74f79d8f8336a1

Observation bd84af8d-a272-4a4e-8be5-ebd876a11f20 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.842609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.842609Z digest=sha256:76bca632ac7e0523b8c959d12808427a9932eee0edd2573815164c501ba44f68

Observation 5f863be8-a115-493b-86e9-1b1388c262ce · outbound

This paper cites Orca: A distributed serving system for transformer-based generative mod- els.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Orca: A distributed serving system for transformer-based generative mod- els

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.846060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.846060Z digest=sha256:b0e5a303be2c941cb6d62e51133ac019c3bb565e4ec549ade71cde5e48e83c19

Observation 4c247a40-3f34-474e-9abf-c3d79632cc3c · outbound

This paper cites Stateful Large Language Model Serving with Pensieve.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Stateful Large Language Model Serving with Pensieve

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.849601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.849601Z digest=sha256:1a8b1644784e331b3948d1f8c2465cacbde7606edf8e3afd81f089ae1d020867

Observation 4a2844c1-3634-45de-b2a6-b6c669d62409 · outbound

This paper cites Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.853284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.853284Z digest=sha256:452b0f18df9b1837d5e6d573ef368cff828cf3f29adcaf451bf6e57d613d3fa3

Observation 704ea79e-3b71-4257-89f0-78163f6d5e8b · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.857107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.857107Z digest=sha256:5e5792247a306b45a9b19782946c15a5a67717ee7d6c1d1135b357414a373718

Observation 111a0041-f61d-4710-8eb1-2b8f4258e3c5 · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.860891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.860891Z digest=sha256:f9f784a37868aa3461430a53ff1d75be6302596693a876bc2e68f83322543a26

Observation 2ea5a426-71ac-40ec-8787-d347aea43be1 · outbound

This paper cites Faster and cheaper serverless computing on harvested resources.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Faster and cheaper serverless computing on harvested resources

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.864874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.864874Z digest=sha256:7017359bb0ea09364fa06e843832db177092e3d946c1c996ec801192a36060b5

Observation af570e82-87f7-43d2-a34f-3514cbb703b2 · outbound

This paper cites LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.868044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.868044Z digest=sha256:6344063ec4e77ad3603a7eac5ff36abad2e0328467313fe4b8b0108831436eed

Observation be2da8b7-7aae-430a-80f0-7c12da66791f · outbound

This paper cites Judging llm-as- a-judge with mt-bench and chatbot arena.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Judging llm-as- a-judge with mt-bench and chatbot arena

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.871664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.871664Z digest=sha256:7f637c8c7bb43d075d3a9ef41d4d79808599ed13f192e5e3d34ae89fc3c4fc5c

Observation 4000ee5f-5a5a-4768-85d1-4e767c690e33 · outbound

This paper cites Efficiently programming large language models using sglang.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Efficiently programming large language models using sglang

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.874842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.874842Z digest=sha256:ea3beb5599c47c2fe3318725f57e4a27a4afa803fe97edc66658d52386add816

Observation eab5219c-9d8c-4cfc-97d3-ab0f472ae1f9 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.877980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.877980Z digest=sha256:f0c719d4eff43adb5325a9af6c917b452814f1b78feaba407173e749fcbf93e3

Observation 7075f73d-61e6-4f70-85b1-d89fbb0c83ae · outbound

This paper cites Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis.

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:01.881688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:57:01.881688Z digest=sha256:574d0bd008ab152df3e2fb8e96cb72b45c685e59c7b8984dc681d5fce81a616f

Pith citing papers

Observation f7de1d06-576c-40c4-9f24-4a0533830e9c · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.967305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:d9638b8a88445b1682eef5333021d6d4fc70d44b0247e7bd6467a32b4fc4d6e9

Observation 73fcc33d-1d8e-4967-b5f1-28bf0cada3ec · inbound

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models cites this paper.

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:55:59.287808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:23:40.872418Z digest=sha256:db2006b728f229c473c480962fd2ab97aea88641de3ec2699a50191328e7acaf

Observation 17a2c2dc-19e2-408e-b1c8-aedd740e52ba · inbound

POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving cites this paper.

POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.269274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:49:43.418576Z digest=sha256:30b0cf7ff81e65815b4a13def4aa9de8fc451327af59949da3b8296b4b038332

Observation 4c19f3af-b35e-44ce-9b08-b052bb18a265 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:51.818338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:25:12.407148Z digest=sha256:7fadfea8a0e492ab1feca7e6de8e87da3fc0f3ce7a4b81442a91437805b29b10

Observation 9980b588-d301-4673-af44-6c390837a1b1 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:06.252808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T21:47:00.295144Z digest=sha256:3c00703515acea604a8822ce2e93715754bbfa5832ed6d686a58bb94f3e8c6aa

Observation d4ac3560-e4c6-4f14-a0bf-fe5cb1507393 · inbound

PreFT: Prefill-only finetuning for efficient inference cites this paper.

PreFT: Prefill-only finetuning for efficient inference Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.657350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T02:10:14.721584Z digest=sha256:3d17e5b0d8f10b5ff30bd367f253f97971ba8324e400b4af2237966ce4f8475a