Pith. sign in

Paper Citation Record · LEDGER

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2607.28150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28150 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:21:19.946822Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved86
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a1b21e5-e5f4-4521-aeb6-f6de3eb6b26f · outbound

This paper cites Phi-4-reasoning Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.065335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.065335Z digest=sha256:4d2ccbffa4476fad73d1e85ac72fd43d8f5bfab8ab4390c542a8c6d21725cfe0

Observation 36e1e90f-7664-46b6-9f7b-327600ed9d06 · outbound

This paper cites Gulavani, Alexey Tumanov, and Ramachandran Ramjee.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gulavani, Alexey Tumanov, and Ramachandran Ramjee

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.156641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.156641Z digest=sha256:488667a7b9d7af2b669b4633b2a420830cc408fc4fdcc27297da124244f01bbf

Observation 92ea6a80-45d0-4643-a3d6-1f1627dbc47a · outbound

This paper cites Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.216603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.216603Z digest=sha256:046a63823d2926f1454a033f476c8a7dee95f97e7bbb4cea1f9d03c9d34c55ba

Observation b1bd9d1a-57e5-47ab-a13b-5e2095273bfe · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LongBench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.259742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.259742Z digest=sha256:dd44d715090bcff4261b3bd07daedefcefa88909d46a05999ad16fbb1c51f013

Observation 90f88231-f561-42ce-b31d-034983fb7af3 · outbound

This paper cites TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.360134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.360134Z digest=sha256:affd1e67c48234423b7849d9f7c487a303a484e1ce00e51b7436268234c322cb

Observation acceb86b-4b7c-4af9-9190-945eabe05800 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.479848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.479848Z digest=sha256:5a822b01da0b5d0f1f0cd284a1e9df34cca259daeea1e45c8b4320a94321525b

Observation 20c871b3-a6b3-431d-bc2a-69a82736f6ea · outbound

This paper cites Retroinfer: A vector storage engine for scalable long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Retroinfer: A vector storage engine for scalable long-context LLM inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.584938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.584938Z digest=sha256:3e04a35ccd85a177224081ce6b3a945c5255b1bab90a9528521c144290e3c536

Observation 16178e22-2491-464b-88e5-19fd8f266fd0 · outbound

This paper cites Elastic GPU service instance families.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Elastic GPU service instance families

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.721313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.721313Z digest=sha256:3b774b7472802263b4bd534a8933bd1f3b0762d694ee73bec263be6bc90ca0ad

Observation 18c0429d-1695-4288-acce-c52a2a710416 · outbound

This paper cites eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.833863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.833863Z digest=sha256:e0fcdfdf550395b189ce26d3148eaeccfe84af99e1f9fb326ecbb38a5d0d3fa5

Observation bd7a11f3-3dd1-4e26-8a27-dbf2bf135df9 · outbound

This paper cites A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.959928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.959928Z digest=sha256:c49010a1b078e6c063adc15bb35c823e8d6a04eec38f0eefbb9ebc1e58c1fc08

Observation c50a4781-be92-40b3-a776-e8fc18801422 · outbound

This paper cites Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.064781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.064781Z digest=sha256:f893e506b2cf19188e2e583432c2813e1bb449f64ea44e6f8ef108e8794da5c5

Observation 5257c786-1809-4850-8017-37a35a81bf17 · outbound

This paper cites NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.135063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.135063Z digest=sha256:10301bfb8193e91e6f012fb168ff073dc60fa3bdca563ea3d0702c9cdf999315

Observation b9706a4c-a753-493a-aaba-c22ae4175ac1 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.200993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.200993Z digest=sha256:a13e9fb2585dfc84376fd27785a04ae02aa9074dbba6878885296e482c9f2461

Observation 16a4b3a7-8577-4ef3-9b2a-abe8b57b95cf · outbound

This paper cites Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.262681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.262681Z digest=sha256:ee63034209f514026b0e2a5e04f3cf36dd609218dead483c38f2b91ac137b568

Observation 7ef8f0b6-e997-48f8-89f3-4205d6870e1e · outbound

This paper cites DeepSeek-V3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.321893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.321893Z digest=sha256:c05e6820c2e95cbe25be46882b94e36f70bd7d821e6e71b22347a98bc19e10fc

Observation ab829f9a-1c01-4adf-9ac4-70e779731ecc · outbound

This paper cites Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.361409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.361409Z digest=sha256:3c8aab99c2fdad87376d9a953d368eece84c80f6597c44308058176f45b4a1ad

Observation 8eeae7a4-03c7-47c3-9a1c-e632c27c51a4 · outbound

This paper cites The design and operation of CloudLab.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The design and operation of CloudLab

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.412761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.412761Z digest=sha256:fc43852f612c204f84fd21d5da726d37ffd99c6da6aee08a189722acceccd747

Observation 7f4ba4ca-34bf-4642-ae95-a4168b81eb4e · outbound

This paper cites Graham, Artem Y.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Graham, Artem Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.463354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.463354Z digest=sha256:6a3c8235869327144b26ee0ff214e9cb16cb6889de817226d8d5c248477ae100

Observation 68f34fcd-5a60-45f1-81f3-51f9404a14e6 · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.530574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.530574Z digest=sha256:fcccc6b1adf4bc8c2b14d317d21fa2388ff5356000ce75addb1771cbc9ec009b

Observation bdfec739-9496-4adb-8558-49c9238dc7a1 · outbound

This paper cites Fast state restoration in LLM serving with HCache.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Fast state restoration in LLM serving with HCache

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.606072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.606072Z digest=sha256:8b6b1ffc905a476b5ab9db37e5c6b1d172939dbf5c56fb03de3604fe332035b5

Observation 91e69d8b-16b4-4014-adb1-f4fafcaa5a53 · outbound

This paper cites Weaver: Efficient multi-llm serving with attention offloading.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Weaver: Efficient multi-llm serving with attention offloading

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.665888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.665888Z digest=sha256:c73f5782779f3bad94eba0d2e63310229cefb6f0e0e027a36da3fd04e0c3bf60

Observation f8146844-e22f-4abb-8533-1663cabdc776 · outbound

This paper cites Hybrid multi- document summarization using pre-trained language models.Expert Syst.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Hybrid multi- document summarization using pre-trained language models.Expert Syst

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.762133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.762133Z digest=sha256:7d62483bb9a34ed5fe1de1ab4366cba915fc794f8bffa004c3e7fe32a69d1cbf

Observation c050fe89-1eab-40a2-805b-e11282e526ba · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.813892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.813892Z digest=sha256:0706df0a9fe65c80e3dbf56c3e9b9f2e099748dde7ff0b85f4f7e882dc155110

Observation 219f1b3b-220d-4a35-92e8-f0d330fcdfa8 · outbound

This paper cites HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:13.882652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:13.882652Z digest=sha256:875c5f0a2d715b83c9471af368e9565e2be7375155d57b0672420a325e80102e

Observation db2f6316-a3f1-415c-84ab-0a00d025d108 · outbound

This paper cites Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.049340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.049340Z digest=sha256:a50ba304b2ec8b95dc5d54e5a5f780fb2abbeca24b6aabbc0f4a9a8b5d60ce29

Observation ec1bfb5d-b426-49a8-90d9-e39889c655cf · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.163297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.163297Z digest=sha256:ac59470348199164eb034c6bb0ace4585494937e3ecc551d5231f836e635fcf9

Observation 768defe2-4ffd-475a-85c2-35e87c35f7f0 · outbound

This paper cites OmniKV: Dynamic context selection for efficient long-context LLMs.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer OmniKV: Dynamic context selection for efficient long-context LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.336284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.336284Z digest=sha256:ccb9737fc92d5d6027e752707f90004a185904cca82f5e5ae2187b43ced40d37

Observation e3109b47-3e74-44ff-8ce5-7f54b51ffd9e · outbound

This paper cites WaferLLM: Large language model inference at wafer scale.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer WaferLLM: Large language model inference at wafer scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.452112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.452112Z digest=sha256:ed8671ec6de5be1bd61e06ea9e21525547c1062f79aeff8b71dabd9381077e8a

Observation 3c1493b0-fd7f-4b3e-9cfd-d922e82bcefc · outbound

This paper cites Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.600544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.600544Z digest=sha256:87623c38406261dbae66068dc0b6c8494ba695aae5243677e72d70b038983b70

Observation a198e9ad-aeb4-4d46-9291-79305597b009 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.757015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.757015Z digest=sha256:045efbc3709c0036d32004538d77bc79473d137cf65f7982fbca46090343e5c1

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:db327048d4a31beb38ed992ba841939bc96a6155b08c848c3c9c1e253c8e7aa0

Observation d39c685a-19f3-414b-8a27-6b841a53c3f4 · outbound

This paper cites Efficient attentions for long document summarization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient attentions for long document summarization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.985128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.985128Z digest=sha256:8dd5dacb1df913651c01be738d0e02310f4fcde20df6ebe26c604423a481070d

Observation b40def8a-c8e6-42a3-a7f2-cc4873d9822f · outbound

This paper cites Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.132083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.132083Z digest=sha256:d87cf407e3e5bc1409116f94eb0e97267909c4fc1968eb9aa44762a6b86983ff

Observation 8fa02f65-eb89-4e6b-bea6-7c9341a4a8a8 · outbound

This paper cites Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.245943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.245943Z digest=sha256:d9b2531b4b71c0216aa99351b0fba67a0b969d723b46bb0fcdab846dc597f26e

Observation f7ba4d0b-9423-4b9d-913c-c231645e266c · outbound

This paper cites Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.349109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.349109Z digest=sha256:17fba8b9b9fb424d390ee60bf222436dffe8e29b3e6fc76f2194ab036b3f031f

Observation 24cbc52d-330e-4989-9f4a-90e53d334753 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient memory management for large language model serving with PagedAttention

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.510728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.510728Z digest=sha256:ee300c10c8bebcc869a22eba814ae576d05d22b1a2545c6feec19b59e9e838d7

Observation 17372a46-8264-4d95-afef-b95fe357b59f · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.620746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.620746Z digest=sha256:cce6e92c6369a41d05a60baf9d6a95a0de01b9189d2edcd2fca7ea5f64e5dba9

Observation 480830a5-bd9f-437d-af28-cd77bd7bdc71 · outbound

This paper cites ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.687846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.687846Z digest=sha256:a4736e8b99bac7717653630f9d0fa308358382ead1d0ec991463dfe8c04ffc2f

Observation 9168d745-635f-4f6e-8fa4-8e125894b1df · outbound

This paper cites Cachegen: KV cache compression and streaming for fast large language model serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Cachegen: KV cache compression and streaming for fast large language model serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.781253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.781253Z digest=sha256:b7867cc8bd64b8d890be37b619b5bd8b43a106ffb82f9428bf30c0e8d7b86e3e

Observation 6d4893e6-b746-483c-bb52-f4cf6674b763 · outbound

This paper cites Helix: Serving large language models over heterogeneous GPUs and network via max-flow.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Helix: Serving large language models over heterogeneous GPUs and network via max-flow

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.839066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.839066Z digest=sha256:457cfaa5d8acd8b73f601edb86d1b6d09b541732db2346be63495263a26cfb8f

Observation 4a679fda-2331-4e5e-9665-2264f296e165 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:15.924508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:15.924508Z digest=sha256:a85e21820767864134d7e0240220c480bb88b5c8e3915748c4b01eb973b37993

Observation d78d0d49-af54-4119-9d9a-f3f00887bd17 · outbound

This paper cites Heterogeneity-aware cluster scheduling policies for deep learning workloads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Heterogeneity-aware cluster scheduling policies for deep learning workloads

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.042849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.042849Z digest=sha256:be4b7cb22d4caefcb6b0bf18ec8000ff95de99f81509fdecc35951b82380441e

Observation 71bf4129-a72d-4573-bac6-9efb850ce811 · outbound

This paper cites GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.128417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.128417Z digest=sha256:4e5160dddb29b2ab4558b87dfe4d6fd28fe24548cf63aacec23d1aff00c4bfa7

Observation b7f1b082-5cf6-48bb-9c06-8d658b5b644d · outbound

This paper cites InstAttention: In-storage attention offloading for cost-effective long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer InstAttention: In-storage attention offloading for cost-effective long-context LLM inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.241941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.241941Z digest=sha256:b0a94ff4c2a10f3efec8acccb5b3b9cedc0678b8c3638b49a4392a80c2b086fd

Observation abe3057a-892d-4051-bb43-c6f60a269d71 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Splitwise: Efficient generative LLM inference using phase splitting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.318560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.318560Z digest=sha256:57a5796b1483131e23de71440740a26052724d05681781875d115f8c562e0311

Observation 27130887-cb0f-40a2-b12a-76381da4c0e3 · outbound

This paper cites vAttention: Dynamic memory management for serving llms with- out PagedAttention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer vAttention: Dynamic memory management for serving llms with- out PagedAttention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.451886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.451886Z digest=sha256:9425c5887ec4e3cf80a6ebb856d5e48983f059c80ef0c899517daeb8cab70c00

Observation b5e26cee-d669-45e7-882f-c3eb61e3177f · outbound

This paper cites Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.537189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.537189Z digest=sha256:e6463dc405b920af2ca88c5ec638f59cb8a8aa6cbde9c02ffcbd9a66fb3c5cb0

Observation d6ea45bf-cee7-42e9-9983-df5588d67c2d · outbound

This paper cites Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.641220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.641220Z digest=sha256:2c47d4c43b4b6e0d6c9fc5153ea9ae5e485da31692dd1f512cfc7844aa055b61

Observation 0d950655-6fe7-4926-aac1-f3e6f96b0c3f · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Code Llama: Open Foundation Models for Code

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.739282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.739282Z digest=sha256:cde03375e3c286aeee7d0c473201d47d92b209131468c4009bf59055cae0a445

Observation 447a7dc3-5d2d-423e-bdfa-5ff0092eab9e · outbound

This paper cites Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.813479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.813479Z digest=sha256:18ac156911aed6a23ca8e6afbbece2594bf29c95bd41f1d59bcf6f85e503b5b5

Observation ee8d85d9-7e40-4039-838e-17fa95eaebbb · outbound

This paper cites Recommended GPU instances.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Recommended GPU instances

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.899455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.899455Z digest=sha256:b72d3b9c80772dd691e9b679dcece2a6fc5b51341cf33631ee64a6d7a6418214

Observation c047fda4-07cc-4f71-b038-ff25ef0f7af2 · outbound

This paper cites Gonzalez, and Ion Stoica.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, and Ion Stoica

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:16.985377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:16.985377Z digest=sha256:3c5d870f2c440fe0dd8ba9588ff28d5f7ece53168f809039a9c6de89978f64df

Observation 9af42044-f506-4a87-9fec-2e54abad43cf · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single GPU.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlexGen: High-throughput generative inference of large language models with a single GPU

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.076373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.076373Z digest=sha256:42d0406a9334bf63dafad5e35687531b6f8eb72fa3570ffe3196f09ae39eff7d

Observation 707ee9f0-01b3-45d7-a105-d5db55c090ac · outbound

This paper cites Maguire Jr., and Dejan Kostic.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Maguire Jr., and Dejan Kostic

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.144547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.144547Z digest=sha256:cc33e96993a32e36f3c335a8aee7507b00ed54dae934c1a21e04d26b8725f1a4

Observation dbc46c3f-e186-4485-85f6-8d25c5de5894 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.265620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.265620Z digest=sha256:debb9f500625718c2c89d4823a20218a0ef1b33047011b392e81a2f2f4b48094

Observation 89ed69c4-7aa6-4a59-bc1c-b0c462225e9b · outbound

This paper cites PowerInfer: Fast large language model serving with a consumer-grade GPU.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer PowerInfer: Fast large language model serving with a consumer-grade GPU

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.444363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.444363Z digest=sha256:ce5a8c7cb02a5325831d7f0c7570dd21c5a604fe50f3ea313ac4a80e1380791d

Observation 51f82ed7-c4e8-4c3d-b5f3-aeb6c480360b · outbound

This paper cites Preble: Efficient distributed prompt scheduling for LLM serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Preble: Efficient distributed prompt scheduling for LLM serving

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.514460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.514460Z digest=sha256:5d5e58139cace2dc14b3e89750204f0f9b1dc69d0aea4a2d548bebdfc1d01cfa

Observation 5d28ecb4-9dd4-4959-a4b1-fb00d6b26d62 · outbound

This paper cites Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.609483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.609483Z digest=sha256:176fc7369e73b7c005c81e789f47a91f35d912d6fbb8b2af5a313a3afaf1bb96

Observation 53cef9d2-a140-470a-a865-0969cd1daecd · outbound

This paper cites QUEST: query-aware sparsity for efficient long-context LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer QUEST: query-aware sparsity for efficient long-context LLM inference

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.673473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.673473Z digest=sha256:a5af18002e6755d1d75f4187a3723d33d288467b63ad49bb3b7e9c42c4b61d37

Observation 89b885dc-3922-4b61-ba37-07abbaaa4c40 · outbound

This paper cites Gemma 3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gemma 3 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.746835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.746835Z digest=sha256:c70325bca7f5667be167ea56074cbde657870ec01a3742173fe56e783f33a8c6

Observation c0ec2726-551d-4284-97bd-3eac364d487d · outbound

This paper cites The Llama 3 Herd of Models.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer The Llama 3 Herd of Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.878772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.878772Z digest=sha256:e0f096834761df59d4e5ca3d8311fccdfe52d4deabff23e4530c05b359717216

Observation 690116a3-8c48-4e8c-a601-00df14918396 · outbound

This paper cites Qwen3 Technical Report.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:17.977097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.977097Z digest=sha256:c6067f9677911184fd07d97673afc5a647c5556662565d06d4895bc5c9d2c816

Observation b44aee88-2c5d-4a16-9047-addd2b89c0bd · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.061323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.061323Z digest=sha256:b582458210ba43028d52bd9dc8e8c9bf815652fa99a42c4e4a5064d2e4819cc2

Observation 07710b8f-929d-4d72-b195-898d83b6990d · outbound

This paper cites KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.149540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.149540Z digest=sha256:dacf94c8174d4428be26d398684a301e1774124bfb34bea9bf867e1ee54ea187

Observation 20d9eb3b-868a-40b7-b2fa-190b65c8cbee · outbound

This paper cites From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.205939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.205939Z digest=sha256:1861ada923dae8e0ad1ba8499f832a4c17c16ff38bf1f6718a9f25bd057a3091

Observation 24c399bc-9622-4b97-9ed9-f5c1554eaaa7 · outbound

This paper cites Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.286738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.286738Z digest=sha256:aebd5c3f312419989a6920993770983ddd49b6a015da177902f755b963d11165

Observation 30eca377-2f8e-4662-b1aa-9842445cd7ed · outbound

This paper cites Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.371796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.371796Z digest=sha256:68cf86517633531d670bf655dd43d4033d3a2eb247af0b25d04684a9f12308bd

Observation faa59009-37cc-4db8-96e1-f6165a693b8d · outbound

This paper cites LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.481672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.481672Z digest=sha256:aeb06893757fc903447bdcb08369586a6f9f576dbc2f11e1dc6a4cf7dd94ddfa

Observation afb05c1d-4a7a-4e15-88a7-17e8ff8fff54 · outbound

This paper cites DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.554551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.554551Z digest=sha256:b0b4d91c9008ed2664e667083091aae21c86fce95d84d2d5cb747718e44baa67

Observation a29c98f4-00f5-4917-b2c9-7d948f907eca · outbound

This paper cites Efficient streaming language models with attention sinks.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Efficient streaming language models with attention sinks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.644504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.644504Z digest=sha256:eb9519fe430d1fb98dc6721bbcb23b4e556d0e1607ca917ec56bf07b36d3fe0f

Observation 6dbbf075-439e-47e9-bfd3-d104192ffc0d · outbound

This paper cites CacheBlend: Fast large language model serving for RAG with cached knowledge fusion.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer CacheBlend: Fast large language model serving for RAG with cached knowledge fusion

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.715514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.715514Z digest=sha256:664e0b5315001792a50d8f847c84a2a864e86d5a347b63ee2826e62865d0a875

Observation f05b22b4-7696-4abe-826e-4e0913b26c81 · outbound

This paper cites FlashInfer: Efficient and customizable attention engine for LLM inference serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer FlashInfer: Efficient and customizable attention engine for LLM inference serving

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.842799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.842799Z digest=sha256:ca92aca2485ed11b19555523d6fb909d1d2d9d8acd9e4b0166abda224ca9ea45

Observation e64eb608-11df-4107-9838-aa70d32b32af · outbound

This paper cites Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:18.946243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:18.946243Z digest=sha256:10d9e7f2c6d5ac03cb10d4499893aae1cc193ceb270dcbeebf8181c2a3d7a8cf

Observation 0d251480-4d71-4901-8605-8624723641b2 · outbound

This paper cites Orca: A distributed 17 serving system for transformer-based generative mod- els.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Orca: A distributed 17 serving system for transformer-based generative mod- els

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.040565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.040565Z digest=sha256:c7b7f54a08428996990189177766d7bdb0445b710c76150dd1a5fb980debf363

Observation 555c11d3-bd72-4d90-bf87-6e816a092f19 · outbound

This paper cites Stateful large language model serving with Pensieve.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Stateful large language model serving with Pensieve

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.124125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.124125Z digest=sha256:7792785f71e43c69c9db1ad241a479e2217d7cdf7cad6f586eefec1002c7f975

Observation 1192aa68-5c5d-4b93-8287-a82c84c26301 · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.211420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.211420Z digest=sha256:e7918caa3a450ed8251c697132a3a7074344916587a6764d5cad30b70fd91795

Observation 402fccba-1f55-4d3d-b334-3a0acd67166a · outbound

This paper cites Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.286688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.286688Z digest=sha256:622467be63b8c5e1385da5a88061b42a4f575fff0b0ebbf9aaf69a6529e169d4

Observation 4c82d628-7573-425e-8245-af6f02bfb635 · outbound

This paper cites Rethinking database high availability with RDMA networks.Proc.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Rethinking database high availability with RDMA networks.Proc

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.375120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.375120Z digest=sha256:730ecb983719eeccf3f7a00ab833e422eeac9dfee3f3b3ca6a637f433a00e0c8

Observation 377907d7-b3b5-4727-91b4-89842814a8f0 · outbound

This paper cites GPU checkpoint/restore made fast and lightweight.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer GPU checkpoint/restore made fast and lightweight

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.465535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.465535Z digest=sha256:f547f01bcebda71d02246327d8e71aec5d563ed95c48eae88e7ef85186c56c7c

Observation e4208fbf-b923-4b7d-8b31-2468997733fb · outbound

This paper cites Jenga: Effective memory manage- ment for serving LLM with heterogeneity.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Jenga: Effective memory manage- ment for serving LLM with heterogeneity

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.526988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.526988Z digest=sha256:3c94b5d9868aa1f5185e4382c037317ae1fac83748c66e99472b0a9efce5d3fc

Observation d81aa742-8004-4f47-8637-cb295962f17c · outbound

This paper cites BlitzScale: Fast and live large model autoscaling with O(1) host caching.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer BlitzScale: Fast and live large model autoscaling with O(1) host caching

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.589085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.589085Z digest=sha256:141ffd6e582fb54880d1efe9de378ff5656e25ebc4fb91dd4895b706ea89177c

Observation 492e4f93-fd0e-4e3b-9c76-b72f1f4028fc · outbound

This paper cites HACK: homomorphic acceleration via compression of the key- value cache for disaggregated LLM inference.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer HACK: homomorphic acceleration via compression of the key- value cache for disaggregated LLM inference

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.661302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.661302Z digest=sha256:b79f4899329eaef2e9dea701912e6fe594fe64774e4cca452777422ecad37525

Observation fa1106db-6561-4d4d-9c62-25479bce809b · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Barrett, Zhangyang Wang, and Beidi Chen

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.735381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.735381Z digest=sha256:cdb737959804647ccd2627f6c1cc7d882d037790be599add3ae905edb05ea4a4

Observation 54039205-e481-42a2-873b-a180b6ef5095 · outbound

This paper cites Gonzalez, Clark W.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Gonzalez, Clark W

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.788776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.788776Z digest=sha256:96d1d2cc3bfe1e64dfa100bcee0a6de8f00cc70b92ca59f6130946af79ecf8f2

Observation d86c767e-91d2-438f-83f9-fb8f8c5dcbe5 · outbound

This paper cites Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.869913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.869913Z digest=sha256:60e4f23565864793a12129d9030fcb6e98af4a9e4d45b008f3799d047a93479c

Observation d5dbca3c-7b69-4390-b99b-ed9965171a49 · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer NanoFlow: Towards optimal large language model serving throughput

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:19.946822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:19.946822Z digest=sha256:e1dce022dee9cfe24102c0bf529a6da27a40b2a12eab2d7606c426c79aae7bfc

Observation b6bb2120-4bfe-4d09-ab2a-7f7d53b1ff4e · outbound

This paper cites an unresolved cited work.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Unresolved cited work

Reference 964

Resolution
parse uncertain
no resolver link, observed 2026-07-31T16:21:17.363791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:17.363791Z digest=sha256:cc31adbede487b066d8d5bca96a3e5a0c7a5c11d6425cdfd66c2e993e7fc6a86

Pith citing papers

No inbound Pith citation observations are available.