Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:55:12.532975Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2508.15881.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:55:12.532975Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.642188Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:57:51.549825Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6edc2b6b-c2c2-4309-9ccc-ad3b0e024c57 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hello GPT-4o, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa797d20-45f1-49d0-a11c-81497eb6b47f · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Claude 3.5 sonnet, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965d0c1e-df52-49f4-ba11-53487b71ee5c · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5385e1-57fc-40a2-9fbf-8c44ec7646e5 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Improving language understanding by generative pre-training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340b23cc-06af-434f-8069-893f1b047482 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f929d593-9219-4713-9a8b-eef18834e01c · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Palu: Compressing KV-Cache with Low-Rank Projection
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec240f3a-ad6f-416e-a183-1e108f27f7a8 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a51bea7-5360-48d8-995f-7b83e8aec3b8 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Transformers are Multi-State RNNs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7d8e1c-4d26-40de-a192-56676940206f · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d7552b-b5c9-49af-b04f-21d5462056a3 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e895fb-9fb8-4df6-b298-9f9d347da560 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5560a2c-a598-4adb-8bf2-445f2ebf7791 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Tensor-parallelism with partially synchronized activations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdaf4a06-f029-4ae2-8f28-43a37df27a65 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d1014f-f333-4cb1-9c37-0284d893ebb7 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e348d791-4afe-4c49-ae5b-11306a7fa47e · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dff6cf-7348-463e-be2b-411436252875 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TransMLA: Multi-Head Latent Attention Is All You Need
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 975fa342-94ee-49e6-911b-192190c4e5a3 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0065ab31-3123-4b57-b276-47b67095ef55 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Llama 3 model card, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c7fda1-37fe-49d8-a9a4-a5142ac201f7 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a655ef5-cf0b-481a-a0e3-b783519b25c6 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hardware-Efficient Attention for Fast Decoding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4903d4-8734-4d1d-8c41-e98c273649d9 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8e1ff1-35dd-4b7e-bb30-414854bdeb30 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ef57b3-bd3a-4e58-9d51-25808beb877d · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01c6fd89-3438-408f-9755-0f5b15e2ed3f · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference SnapKV: LLM Knows What You are Looking for Before Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ec6af8-00c3-42e1-aba7-a1a71f7b5545 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9828a0e-fcb4-4bdf-be76-030d6eaa0a87 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b418c4-ede0-4354-bc8b-8640b51de096 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f9f956-7c89-4b28-920a-cdc8de6a3b42 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Zsmerge: Zero-shot kv cache compression for memory-efficient long-context llms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f168205e-0307-4026-8c99-cef1defffce9 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Efficient long-context llm inference via kv cache clustering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d040f08-9e60-418b-ae63-6fe5022c1419 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0279807a-455a-42b3-9ac9-00da505464b2 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Inference-Friendly Models With MixAttention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60be5ed8-16ce-4a7e-941f-1cd3df3493bf · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62bd1eb-86d9-4cb8-8820-9fec1ed413d2 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d5022ab-4df1-4ddb-b02e-098ef4e89e67 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0fb4e1-5587-4190-a412-ac69e2e71b3b · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2434c60e-8e4b-408c-8625-af96bc073b3c · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129d056b-aa76-4449-8ef0-fd01fe67112e · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Effectively Compress KV Heads for LLM
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30df9cd3-da21-4b66-94cc-6e529e1f16c4 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406bcb6e-4a11-4d5f-a52a-d2c81af75543 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e4f6c27-5b95-414c-8f31-40f5ea2c1fff · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06792c92-3127-4413-b66c-d107532327c0 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06455c16-75ee-45b9-a55e-80ba7532ad83 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Large scale distributed deep networks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed51e7b-2b70-4475-a962-b76d80725d2e · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Horovod: fast and easy distributed deep learning in TensorFlow
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1925cf9e-10c1-4370-867c-5a8b2a871581 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7396cf27-1e60-4aa8-96ba-124365c1f946 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Pipedream: Generalized pipeline parallelism for dnn training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d8ba3e-4312-4cc6-a132-ffc9930d9915 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1771d614-49e7-4410-b557-3a488ece433a · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference An Efficient 2D Method for Training Super-Large Deep Learning Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16c73e13-a6d2-4f21-b730-3382a05f53b1 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Maximizing Parallelism in Distributed Training for Huge Neural Networks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc945350-74c7-40d7-a7fe-22216310527c · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sequence Parallelism: Long Sequence Training from System Perspective
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ef7060-1cc4-4fb6-a21a-afd6fa062cb5 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference NVIDIA dynamo, a low-latency distributed inference framework for scaling reasoning ai models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ef47a78-7883-4088-bfe2-a993c0cb07fa · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ae5a0c8-de1c-46c6-ae05-062650b6c4ca · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb093de-5471-4821-a6ec-228d975ac973 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1dab253-a3ef-4535-9418-2b625967db27 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Kimi K2: Open Agentic Intelligence
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986acead-968f-4d81-a5cd-48dded46847d · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Mea- suring massive multitask language understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62c1e467-c512-4891-a383-774e9c6755cc · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce74745c-520b-4f95-a0a1-213aede0a115 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference PIQA: reasoning about physical commonsense in natural language
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8dfad646-0691-4224-80fa-631e4c3e4c13 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a14dcce-2508-4c20-9d78-59ed1a8a7e84 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Can a suit of armor conduct electricity? A new dataset for open book question answering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f753b1f5-1f3a-4978-ae79-f8259aa643e9 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Winogrande: an adversarial winograd schema challenge at scale
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9d15e7c-55ff-4312-b58a-8c29caef111c · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Pointer sentinel mixture models, 2016
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcaf792-9df4-4cac-a2a0-7e6b2e80661b · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Smollm-corpus
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 150c8a9f-d5b0-4fb1-9061-5f6866b520a0 · outbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb906a33-09fe-4c7b-bf98-729123df890d · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f560a902-6b34-41e3-84c4-5b085714bdde · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6107027a-ed16-41c6-ab09-fbc260fc3daa · inbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.