Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:27.634595Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.01891.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:27.634595Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c8b61952-edf6-419b-8c18-0766b3e15927 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6ae754-fee4-4a2b-a942-2cc2210b866c · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5146295f-fdc9-4792-9e52-37b510fcd8fe · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Azure Public Dataset,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f222f9-27ed-4bad-b2e1-8f9b7fd4991a · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Biscale: Energy-efficient disaggregated llm serving via phase-aware placement and dvfs,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8891bc-fa31-4f8a-9b9d-e67073499dcd · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Reducing the carbon impact of generative ai inference (today and in 2035),
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e9e55b-1e0a-40e3-8c73-ad1defc5eff3 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Copilot,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7adb02b2-d8de-4fd5-82cc-15ff0deabe3f · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention-2: Faster attention with better parallelism and work partitioning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a4e0ba6-d21e-4b69-8d7d-915c9b3fcfe3 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b169a158-7018-48c4-90a4-f1e0e3fbd971 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Context Caching,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08e005f-87c2-4338-99d1-66eec8b4dc53 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Smoothoperator: Reducing power fragmentation and improving power utilization in large-scale dat- acenters,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b17282e-2cbb-4913-896d-4eaa09d08b5c · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457c84e7-5678-48f6-a755-cebf572dc9d2 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling throttll’em: Predictive gpu throttling for energy efficient llm inference serving,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e7ff0c-8bea-4d4e-9641-10d6b95e32f0 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Efficient memory management for large language model serving with pagedattention,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7303883-8393-45f5-a34f-296637f6ce26 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Greenllm: Slo-aware dynamic frequency scaling for energy-efficient llm serving,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794e38b3-7266-455a-84f7-dca308ad216b · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Lmcache: An efficient kv cache layer for enterprise-scale llm inference,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef64e5d-7e77-42b0-8889-997a0ab070de · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling The Mixtral-8x7B Large Language Model,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a549a92-6e5b-4dda-bd38-c611ca53b34b · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling NVIDIA NVML API,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43d3f0f-5404-4a5d-a9a7-df170b12c193 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling TensorRT-LLM,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59d7b1f-d478-4908-a202-92558e179c5f · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Chatgpt,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58744b07-d77d-4831-a327-7c569edba536 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Splitwise: Efficient generative LLM inference using phase splitting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab4d207-f061-4a53-b638-09625136db28 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Characterizing power management opportunities for llms in the cloud,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54002eba-c663-49c7-bb85-6036f9d4ec6c · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Scikit-learn: Machine learning in Python,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f16d29-84ed-4854-a1d6-e6d4afe4c52b · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling A study of generative large language model for medical research and healthcare,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab04deb-8913-4d50-9b27-0d582d93e81b · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Mooncake: A kvcache-centric disaggregated architecture for llm serving,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ebd65c-c8b8-42f3-8606-daea37da5e80 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Accelerating retrieval-augmented generation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cded0ae-83a2-4c83-b703-5000c6474660 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flexgen: High-throughput generative inference of large language models with a single gpu,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c39d644-88c3-475f-8410-d9ef367d75cc · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2e6577-78f2-483f-8cd8-6c321991c303 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Dy- namollm: Designing llm inference clusters for performance and energy efficiency,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f707d1-3c7e-4b6c-9620-10046c86d5c1 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Qwen3 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e741e68c-649e-4678-8028-0a09900ed3ec · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32477326-cf53-49a0-bba3-40945c07fecc · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Fast Distributed Inference Serving for Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d6a954-3afd-4f8a-9844-63303edc836d · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling xdeepserve: Model-as-a-service on huawei cloudmatrix384,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb5a811-269a-4111-b665-e4ca18f0ba36 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Orca: A distributed serving system for{Transformer-Based}generative models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613bd53a-614b-4575-9df9-2b3852c3a5e1 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention-4: Algorithm and kernel pipelining co-design for asym- metric hardware scaling,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130dfaaf-46ba-4f50-abe4-03c045b3c6c3 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling ZeroMQ API,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ee6a5e-891b-4c64-bad8-269bd69d9b41 · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SGLang: Efficient Execution of Structured Language Model Programs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c37b2b4-f27d-4ada-8ab9-0926f0f264ab · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2e0bd0-1e48-40aa-af6f-10cf8bedd45b · outbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Megascale-infer: Efficient mixture- of-experts model serving with disaggregated expert parallelism,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.