Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T16:21:19.946822Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2607.28150.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T16:21:19.946822Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a1b21e5-e5f4-4521-aeb6-f6de3eb6b26f · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phi-4-reasoning Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e1e90f-7664-46b6-9f7b-327600ed9d06 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ea6a80-45d0-4643-a3d6-1f1627dbc47a · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bd9d1a-57e5-47ab-a13b-5e2095273bfe · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LongBench: A bilingual, multitask benchmark for long context understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f88231-f561-42ce-b31d-034983fb7af3 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acceb86b-4b7c-4af9-9190-945eabe05800 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c871b3-a6b3-431d-bc2a-69a82736f6ea · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Retroinfer: A vector storage engine for scalable long-context LLM inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16178e22-2491-464b-88e5-19fd8f266fd0 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Elastic GPU service instance families
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c0429d-1695-4288-acce-c52a2a710416 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd7a11f3-3dd1-4e26-8a27-dbf2bf135df9 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50a4781-be92-40b3-a776-e8fc18801422 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5257c786-1809-4850-8017-37a35a81bf17 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9706a4c-a753-493a-aaba-c22ae4175ac1 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a4b3a7-8577-4ef3-9b2a-abe8b57b95cf · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef8f0b6-e997-48f8-89f3-4205d6870e1e · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DeepSeek-V3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab829f9a-1c01-4adf-9ac4-70e779731ecc · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eeae7a4-03c7-47c3-9a1c-e632c27c51a4 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The design and operation of CloudLab
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4ba4ca-34bf-4642-ae95-a4168b81eb4e · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Graham, Artem Y
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f34fcd-5a60-45f1-81f3-51f9404a14e6 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfec739-9496-4adb-8558-49c9238dc7a1 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fast state restoration in LLM serving with HCache
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e69d8b-16b4-4014-adb1-f4fafcaa5a53 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Weaver: Efficient multi-llm serving with attention offloading
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8146844-e22f-4abb-8533-1663cabdc776 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Hybrid multi- document summarization using pre-trained language models.Expert Syst
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c050fe89-1eab-40a2-805b-e11282e526ba · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219f1b3b-220d-4a35-92e8-f0d330fcdfa8 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db2f6316-a3f1-415c-84ab-0a00d025d108 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec1bfb5d-b426-49a8-90d9-e39889c655cf · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768defe2-4ffd-475a-85c2-35e87c35f7f0 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer OmniKV: Dynamic context selection for efficient long-context LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3109b47-3e74-44ff-8ce5-7f54b51ffd9e · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer WaferLLM: Large language model inference at wafer scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1493b0-fd7f-4b3e-9cfd-d922e82bcefc · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a198e9ad-aeb4-4d46-9291-79305597b009 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39c685a-19f3-414b-8a27-6b841a53c3f4 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient attentions for long document summarization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40def8a-c8e6-42a3-a7f2-cc4873d9822f · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa02f65-eb89-4e6b-bea6-7c9341a4a8a8 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ba4d0b-9423-4b9d-913c-c231645e266c · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24cbc52d-330e-4989-9f4a-90e53d334753 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient memory management for large language model serving with PagedAttention
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17372a46-8264-4d95-afef-b95fe357b59f · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InfiniGen: Efficient generative inference of large language models with dynamic KV cache management
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480830a5-bd9f-437d-af28-cd77bd7bdc71 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9168d745-635f-4f6e-8fa4-8e125894b1df · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cachegen: KV cache compression and streaming for fast large language model serving
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d4893e6-b746-483c-bb52-f4cf6674b763 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Helix: Serving large language models over heterogeneous GPUs and network via max-flow
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a679fda-2331-4e5e-9665-2264f296e165 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78d0d49-af54-4119-9d9a-f3f00887bd17 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Heterogeneity-aware cluster scheduling policies for deep learning workloads
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bf4129-a72d-4573-bac6-9efb850ce811 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f1b082-5cf6-48bb-9c06-8d658b5b644d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InstAttention: In-storage attention offloading for cost-effective long-context LLM inference
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe3057a-892d-4051-bb43-c6f60a269d71 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Splitwise: Efficient generative LLM inference using phase splitting
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27130887-cb0f-40a2-b12a-76381da4c0e3 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer vAttention: Dynamic memory management for serving llms with- out PagedAttention
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e26cee-d669-45e7-882f-c3eb61e3177f · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ea45bf-cee7-42e9-9983-df5588d67c2d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d950655-6fe7-4926-aac1-f3e6f96b0c3f · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Code Llama: Open Foundation Models for Code
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447a7dc3-5d2d-423e-bdfa-5ff0092eab9e · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8d85d9-7e40-4039-838e-17fa95eaebbb · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Recommended GPU instances
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c047fda4-07cc-4f71-b038-ff25ef0f7af2 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, and Ion Stoica
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af42044-f506-4a87-9fec-2e54abad43cf · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlexGen: High-throughput generative inference of large language models with a single GPU
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 707ee9f0-01b3-45d7-a105-d5db55c090ac · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Maguire Jr., and Dejan Kostic
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc46c3f-e186-4485-85f6-8d25c5de5894 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ed69c4-7aa6-4a59-bc1c-b0c462225e9b · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer PowerInfer: Fast large language model serving with a consumer-grade GPU
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f82ed7-c4e8-4c3d-b5f3-aeb6c480360b · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Preble: Efficient distributed prompt scheduling for LLM serving
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d28ecb4-9dd4-4959-a4b1-fb00d6b26d62 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53cef9d2-a140-470a-a865-0969cd1daecd · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer QUEST: query-aware sparsity for efficient long-context LLM inference
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b885dc-3922-4b61-ba37-07abbaaa4c40 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gemma 3 Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ec2726-551d-4284-97bd-3eac364d487d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The Llama 3 Herd of Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690116a3-8c48-4e8c-a601-00df14918396 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Qwen3 Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44aee88-2c5d-4a16-9047-addd2b89c0bd · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07710b8f-929d-4d72-b195-898d83b6990d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d9eb3b-868a-40b7-b2fa-190b65c8cbee · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24c399bc-9622-4b97-9ed9-f5c1554eaaa7 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30eca377-2f8e-4662-b1aa-9842445cd7ed · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa59009-37cc-4db8-96e1-f6165a693b8d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb05c1d-4a7a-4e15-88a7-17e8ff8fff54 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29c98f4-00f5-4917-b2c9-7d948f907eca · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient streaming language models with attention sinks
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dbbf075-439e-47e9-bfd3-d104192ffc0d · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer CacheBlend: Fast large language model serving for RAG with cached knowledge fusion
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f05b22b4-7696-4abe-826e-4e0913b26c81 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlashInfer: Efficient and customizable attention engine for LLM inference serving
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64eb608-11df-4107-9838-aa70d32b32af · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d251480-4d71-4901-8605-8624723641b2 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Orca: A distributed 17 serving system for transformer-based generative mod- els
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555c11d3-bd72-4d90-bf87-6e816a092f19 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Stateful large language model serving with Pensieve
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1192aa68-5c5d-4b93-8287-a82c84c26301 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402fccba-1f55-4d3d-b334-3a0acd67166a · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c82d628-7573-425e-8245-af6f02bfb635 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Rethinking database high availability with RDMA networks.Proc
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377907d7-b3b5-4727-91b4-89842814a8f0 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPU checkpoint/restore made fast and lightweight
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4208fbf-b923-4b7d-8b31-2468997733fb · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Jenga: Effective memory manage- ment for serving LLM with heterogeneity
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81aa742-8004-4f47-8637-cb295962f17c · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer BlitzScale: Fast and live large model autoscaling with O(1) host caching
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 492e4f93-fd0e-4e3b-9c76-b72f1f4028fc · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HACK: homomorphic acceleration via compression of the key- value cache for disaggregated LLM inference
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1106db-6561-4d4d-9c62-25479bce809b · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Barrett, Zhangyang Wang, and Beidi Chen
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54039205-e481-42a2-873b-a180b6ef5095 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, Clark W
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86c767e-91d2-438f-83f9-fb8f8c5dcbe5 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5dbca3c-7b69-4390-b99b-ed9965171a49 · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NanoFlow: Towards optimal large language model serving throughput
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6bb2120-4bfe-4d09-ab2a-7f7d53b1ff4e · outbound
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work
Reference 964
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.