Pith. sign in

Paper Citation Record · LEDGER

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2508.15881.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15881 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:55:12.532975Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.642188Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.549825Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact10
  • verified fuzzy7
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6edc2b6b-c2c2-4309-9ccc-ad3b0e024c57 · outbound

This paper cites Hello GPT-4o, 2024.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hello GPT-4o, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:05.969609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:05.969609Z digest=sha256:7e09f5d281fea820809656e1d36603c5940b6cd18fbf1754bda68e90dbe0bfc4

Observation fa797d20-45f1-49d0-a11c-81497eb6b47f · outbound

This paper cites Claude 3.5 sonnet, 2024.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Claude 3.5 sonnet, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.070966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.070966Z digest=sha256:f401e20a92d39cd8dd551481af7cc198d1257eef46b1dbded1bcd8fd542f98d3

Observation 965d0c1e-df52-49f4-ba11-53487b71ee5c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.170024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.170024Z digest=sha256:7c37b9deae76bc842382b9039d60a61c3b53637a042f482232b66a228dc15814

Observation 7f5385e1-57fc-40a2-9fbf-8c44ec7646e5 · outbound

This paper cites Improving language understanding by generative pre-training.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Improving language understanding by generative pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.200705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.200705Z digest=sha256:95f07c2b0c7c4e24bf696c71bf2b2623aa6e1998348beb4b9cf4eccdc02d3e9d

Observation 340b23cc-06af-434f-8069-893f1b047482 · outbound

This paper cites Language models are few-shot learners.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.281785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.281785Z digest=sha256:967a638fe184fd125df3367bd32ded7fb4eb7329b23234a2ce38e986bdd02046

Observation f929d593-9219-4713-9a8b-eef18834e01c · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.348432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.348432Z digest=sha256:fa5a0a52403b3bcb26b786f5e6b0b7b4f998d3a4bdff9670f1f1b71292d82cee

Observation ec240f3a-ad6f-416e-a183-1e108f27f7a8 · outbound

This paper cites xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.405836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.405836Z digest=sha256:6eead9517a3bee3c004a900ac70c665bc59098e0339794bdee2d8fcf5855dd8f

Observation 9a51bea7-5360-48d8-995f-7b83e8aec3b8 · outbound

This paper cites Transformers are Multi-State RNNs.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Transformers are Multi-State RNNs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.525852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.525852Z digest=sha256:8b6d8acd06ebcd7d0d175ea220d65461f38516e80d53eda0f4b20f66c6ec3e63

Observation 1c7d8e1c-4d26-40de-a192-56676940206f · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.637672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.637672Z digest=sha256:6b4ea8fb796769ef1155facbb7fc35d40e93ca4aafa4f7ccf018c759ad4e6450

Observation f9d7552b-b5c9-49af-b04f-21d5462056a3 · outbound

This paper cites Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.756104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.756104Z digest=sha256:2f4cfc027dea5318c23530a9933a6f89438448e042699039ae307eb279b23901

Observation d9e895fb-9fb8-4df6-b298-9f9d347da560 · outbound

This paper cites SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:15.253571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:06.872252Z digest=sha256:b81d979db6bfd6cf9174424ee0b42f5b612011ccf1542786afe7cd8350dec609

Observation d5560a2c-a598-4adb-8bf2-445f2ebf7791 · outbound

This paper cites Tensor-parallelism with partially synchronized activations.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Tensor-parallelism with partially synchronized activations

Reference 12

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:55:14.958729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:06.984488Z digest=sha256:742366eca466e6d4fa72ca016d0c1dbe14c7d0f2d0e0d8951339518c0ad80699

Observation cdaf4a06-f029-4ae2-8f28-43a37df27a65 · outbound

This paper cites Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.138663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.138663Z digest=sha256:4f19fdcf99af1f943d1f07245d9225b7c0ca1b455fa75477fca2cce35e14d163

Observation 90d1014f-f333-4cb1-9c37-0284d893ebb7 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.250063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.250063Z digest=sha256:a5ad4c5d15c53a1b2b7daf79b3c0f6040993273e59fd7e1ccb7a2142d032f84f

Observation e348d791-4afe-4c49-ae5b-11306a7fa47e · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.365952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.365952Z digest=sha256:a4ab246f6d57f62bfe83d3a132c81b09aa1c545ee736c85e2ec68f14a73cdb02

Observation 63dff6cf-7348-463e-be2b-411436252875 · outbound

This paper cites TransMLA: Multi-Head Latent Attention Is All You Need.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TransMLA: Multi-Head Latent Attention Is All You Need

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.508723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.508723Z digest=sha256:bff15bbe49132aed18e528de66265c68259f30ec1f5b376a387a88b929a560e6

Observation 975fa342-94ee-49e6-911b-192190c4e5a3 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.673578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.673578Z digest=sha256:f2712765f9c7fa192d44aea8e7fb3bd0534cf03530ec3639bbc3b6e44189476e

Observation 0065ab31-3123-4b57-b276-47b67095ef55 · outbound

This paper cites Llama 3 model card, 2024.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Llama 3 model card, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.854166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.854166Z digest=sha256:aba10d4a64895f760ce4c4f85eacb52fe35bf745990e136f64841654e0b67411

Observation c2c7fda1-37fe-49d8-a9a4-a5142ac201f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.985138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.985138Z digest=sha256:cf67bcaec9c9518cb0cc36703c7e7e32d9f8ac0a7a14bb6f36aac8374c160b06

Observation 3a655ef5-cf0b-481a-a0e3-b783519b25c6 · outbound

This paper cites Hardware-Efficient Attention for Fast Decoding.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hardware-Efficient Attention for Fast Decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.095113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.095113Z digest=sha256:d7daec1306a1f692ccd588a37416d6f7b43a6920e9c9b41197b39b22a890bfff

Observation bf4903d4-8734-4d1d-8c41-e98c273649d9 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.179444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.179444Z digest=sha256:685cefb164a1395c35bb45316a8c797251bba01ddf35152971e3062d3328fc04

Observation bb8e1ff1-35dd-4b7e-bb30-414854bdeb30 · outbound

This paper cites DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.261078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.261078Z digest=sha256:e84f99049f3e16c19c7feb0bec47d7bac92db87bfeb760959ef6e468113d0bf7

Observation 74ef57b3-bd3a-4e58-9d51-25808beb877d · outbound

This paper cites CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.361255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.361255Z digest=sha256:fbb055e808390e2956d2162e0fd65cebd63831c89bf6f4ac95c9dcbe26cffcde

Observation 01c6fd89-3438-408f-9755-0f5b15e2ed3f · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference SnapKV: LLM Knows What You are Looking for Before Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.477550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.477550Z digest=sha256:5e04395790a34ec5d226e3d3e69f272688ba30254d048997ebfccc5bfc3aa9ff

Observation d7ec6af8-00c3-42e1-aba7-a1a71f7b5545 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.556414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.556414Z digest=sha256:4ce65992b0536ed266c029d001ce647433ceb0435cd1f39deff2a482262a6643

Observation f9828a0e-fcb4-4bdf-be76-030d6eaa0a87 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.677974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.677974Z digest=sha256:48a5650f428c97a8c88380a1014b79a6af62b6bc60f03c7e564595010ed1daeb

Observation 53b418c4-ede0-4354-bc8b-8640b51de096 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.753827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.753827Z digest=sha256:f5e3b2bf5b1488b8f2afa2407736ee8c67a18ce2a9336cac5c6cbb7633ac097c

Observation 64f9f956-7c89-4b28-920a-cdc8de6a3b42 · outbound

This paper cites Zsmerge: Zero-shot kv cache compression for memory-efficient long-context llms.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Zsmerge: Zero-shot kv cache compression for memory-efficient long-context llms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.857090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.857090Z digest=sha256:217b7713a6f607b7fa4d24c4866784265f5121c4d14d9e6bb7a390d30151c34c

Observation f168205e-0307-4026-8c99-cef1defffce9 · outbound

This paper cites Efficient long-context llm inference via kv cache clustering.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Efficient long-context llm inference via kv cache clustering

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:55:14.366584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:08.971911Z digest=sha256:3a1d3b1df6fca8ed400ee244d9d1f43308b0661ae8f6fa97c071a5358b88d222

Observation 9d040f08-9e60-418b-ae63-6fe5022c1419 · outbound

This paper cites KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.149968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.149968Z digest=sha256:c29a9ccdb1fd3799aae60dd88d126db117515de2bccaea81b10f1fcb54218e92

Observation 0279807a-455a-42b3-9ac9-00da505464b2 · outbound

This paper cites Inference-Friendly Models With MixAttention.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Inference-Friendly Models With MixAttention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:14.064205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:09.267444Z digest=sha256:1765baa08bdca57e011519b5ead8e3445eb5f2895c452bc171b5f5f19c1b7448

Observation 60be5ed8-16ce-4a7e-941f-1cd3df3493bf · outbound

This paper cites Layer-Condensed KV Cache for Efficient Inference of Large Language Models.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.408451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.408451Z digest=sha256:2fc34c9d6429eb79b1e712cea2960a031060c5607a307436f7324f992f7263e6

Observation e62bd1eb-86d9-4cb8-8820-9fec1ed413d2 · outbound

This paper cites A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.833354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:09.555984Z digest=sha256:eace7d069e857d0d35b7d13a85b731468b3e8b449690791019d32865a9072810

Observation 6d5022ab-4df1-4ddb-b02e-098ef4e89e67 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.661574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.661574Z digest=sha256:02ccfc41469a70963020308daa430142297ce7a756b42332733fa117d3571778

Observation 0c0fb4e1-5587-4190-a412-ac69e2e71b3b · outbound

This paper cites LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.743965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.743965Z digest=sha256:53bb43c8599f1a495d9740c391e51ca2e5fbcaee5e80c75aaf309233299fcde7

Observation 2434c60e-8e4b-408c-8625-af96bc073b3c · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.824153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.824153Z digest=sha256:5de2fafc71b305681a8434ffe1063fd26b23c65de56e0091e65a660917c2a6b8

Observation 129d056b-aa76-4449-8ef0-fd01fe67112e · outbound

This paper cites Effectively Compress KV Heads for LLM.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Effectively Compress KV Heads for LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.928919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.928919Z digest=sha256:6a0255618a91e927e51b424077df181c8817f8628fb25322ae75ec1fc3678150

Observation 30df9cd3-da21-4b66-94cc-6e529e1f16c4 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.035877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.035877Z digest=sha256:671424e91dcec970890b1b06a6aebe2abc19bffac15ee0eb33e8515f583a57d8

Observation 406bcb6e-4a11-4d5f-a52a-d2c81af75543 · outbound

This paper cites MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.570436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:10.191982Z digest=sha256:e88b9ed3d11a1b5c0c98f22ff825ef175d82cafac27b8496a9c63e790077ea55

Observation 4e4f6c27-5b95-414c-8f31-40f5ea2c1fff · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.297309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.297309Z digest=sha256:d02142a984ab96e511fe56ee2dc91a4860f89cd2aa32d3040b145054a0a844d4

Observation 06792c92-3127-4413-b66c-d107532327c0 · outbound

This paper cites TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:10.372511Z digest=sha256:978096f4998a5de3e25062ad733a5b15bb254b490ac72111c02cc101d58ddae7

Observation 06455c16-75ee-45b9-a55e-80ba7532ad83 · outbound

This paper cites Large scale distributed deep networks.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Large scale distributed deep networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.478577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.478577Z digest=sha256:2ec863b29d195e58c9fafb1ae0921abdfe1aabcd6096ca4c418bf8ff9284da37

Observation 5ed51e7b-2b70-4475-a962-b76d80725d2e · outbound

This paper cites Horovod: fast and easy distributed deep learning in TensorFlow.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Horovod: fast and easy distributed deep learning in TensorFlow

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.614756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.614756Z digest=sha256:2f80767f2ef8b99c25c007bb044d7b691c860970a92ed1f22861096e5338aaa1

Observation 1925cf9e-10c1-4370-867c-5a8b2a871581 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.743554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.743554Z digest=sha256:ea4c615dba38690d15581220a5c9c118df13015ac26ebbfac08a873a2f00452d

Observation 7396cf27-1e60-4aa8-96ba-124365c1f946 · outbound

This paper cites Pipedream: Generalized pipeline parallelism for dnn training.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Pipedream: Generalized pipeline parallelism for dnn training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.818690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.818690Z digest=sha256:37e978749981a42fb12c4074f09da200125604733394ccc83fb4b555d8134ea2

Observation 69d8ba3e-4312-4cc6-a132-ffc9930d9915 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.957633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.957633Z digest=sha256:db1930df81591c3dd3dd127354cf064eb613864a6f5c179c0c40fd8fea57b7af

Observation 1771d614-49e7-4410-b557-3a488ece433a · outbound

This paper cites An Efficient 2D Method for Training Super-Large Deep Learning Models.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference An Efficient 2D Method for Training Super-Large Deep Learning Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.186945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.079164Z digest=sha256:28103fe359a328429f016b65efd43de39c7fd2d64524f57943ab4aa1b409cce8

Observation 16c73e13-a6d2-4f21-b730-3382a05f53b1 · outbound

This paper cites Maximizing Parallelism in Distributed Training for Huge Neural Networks.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Maximizing Parallelism in Distributed Training for Huge Neural Networks

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.001003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.149864Z digest=sha256:de74f47a3df62f1fb3d1f4b9ac7f4a8f0542925b71e9afce65c058dcf0e9d5eb

Observation dc945350-74c7-40d7-a7fe-22216310527c · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.234042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.234042Z digest=sha256:f04a725fdd476f9dff4675d7dd66cf962731011b6d2460f95543f3815062ddb0

Observation 70ef7060-1cc4-4fb6-a21a-afd6fa062cb5 · outbound

This paper cites NVIDIA dynamo, a low-latency distributed inference framework for scaling reasoning ai models.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference NVIDIA dynamo, a low-latency distributed inference framework for scaling reasoning ai models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:16.949194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.341384Z digest=sha256:2ff2338ee28f89aaff2b07f2b7ab96f7c43f0ec9b1713920919259d9c15d768f

Observation 0ef47a78-7883-4088-bfe2-a993c0cb07fa · outbound

This paper cites Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:12.797644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.419047Z digest=sha256:c6bd89d764f6da13cf373699e983b8dbb51cb8dc0f251f893166aac48dcfae81

Observation 5ae5a0c8-de1c-46c6-ae05-062650b6c4ca · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.497826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.497826Z digest=sha256:0ace11603d32e6457eac55268041b3e00d01d9b01224a7aa59a2436100c0b358

Observation 2bb093de-5471-4821-a6ec-228d975ac973 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.586104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.586104Z digest=sha256:f747492cebf8e74c00f994822b2531eca92d7fa83dc5968aaf8786c0819f7fb1

Observation f1dab253-a3ef-4535-9418-2b625967db27 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Kimi K2: Open Agentic Intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.656509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.656509Z digest=sha256:7a6adf2ee959ce12f513fac34f42cff9544bafefebb027709691cdc743922cfe

Observation 986acead-968f-4d81-a5cd-48dded46847d · outbound

This paper cites Mea- suring massive multitask language understanding.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Mea- suring massive multitask language understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:16.735462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.777292Z digest=sha256:04bc4e927ef53aa39a25519f9eae622b862a307af983cfb19966734d93a107cf

Observation 62c1e467-c512-4891-a383-774e9c6755cc · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.883455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.883455Z digest=sha256:203df2fc508d204de133051b4618db36018d73212d871a76ca47f5ed512e3781

Observation ce74745c-520b-4f95-a0a1-213aede0a115 · outbound

This paper cites PIQA: reasoning about physical commonsense in natural language.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference PIQA: reasoning about physical commonsense in natural language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:16.503922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:11.978403Z digest=sha256:07e37ac308d204a7751fa86e00128d1ce4f614fab53e7982927fdf63528396ed

Observation 8dfad646-0691-4224-80fa-631e4c3e4c13 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:16.307242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:12.075515Z digest=sha256:926a5f892ae99b54e9dc8a15aba010bcac915d1a904601f39d8d56a74191dc58

Observation 7a14dcce-2508-4c20-9d78-59ed1a8a7e84 · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:16.074606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:12.167299Z digest=sha256:7f5fb60ddefa348b7de09db337fbb10beb6d5097ca7bb64be0e2a1e907671066

Observation f753b1f5-1f3a-4978-ae79-f8259aa643e9 · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Winogrande: an adversarial winograd schema challenge at scale

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:15.841490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:12.240880Z digest=sha256:b06edf48b6247d24a15b72ffc0f56e970829dd186830874792234c086d4de8f5

Observation e9d15e7c-55ff-4312-b58a-8c29caef111c · outbound

This paper cites Pointer sentinel mixture models, 2016.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Pointer sentinel mixture models, 2016

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:12.333239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:12.333239Z digest=sha256:5053d89e8f8b79fbcbb8c2835409e0fec0c024066aa0d72384a25e130de47dd9

Observation 4bcaf792-9df4-4cac-a2a0-7e6b2e80661b · outbound

This paper cites Smollm-corpus.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Smollm-corpus

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:55:15.579058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:55:12.427182Z digest=sha256:1bd901e5888b28302015c26d1bb803fd9f048db466a08667178c239b7d95a5f7

Observation 150c8a9f-d5b0-4fb1-9061-5f6866b520a0 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:12.532975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:12.532975Z digest=sha256:05e6b65788350c428d3d51b5d55179c5fdcc36ce2ecab1ccf1c0e84760cfeec8

Pith citing papers

Observation bb906a33-09fe-4c7b-bf98-729123df890d · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.138975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:f50b7cf69423fb1f7b78910089a1e789881a367c0fa8eae7a4dff0aa521b9c35

Observation f560a902-6b34-41e3-84c4-5b085714bdde · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.582428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:7ae58b1fcbfbb625a8c2108a598bf455e35ef1efbd0a4dc4a4b14f797ec3af3a

Observation 6107027a-ed16-41c6-ab09-fbc260fc3daa · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.642188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.642188Z digest=sha256:75424bf80b0343142a89d064181b9d60d4f7dad5a893bd5c3dc9d21708ce2140