Lampinen, and Stephanie C

Nicolas Zucchet, Francesco d’Angelo, Andrew K · 2025 · arXiv 2505.17863

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

representative citing papers

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference

cs.LG · 2026-05-27 · unverdicted · novelty 7.0

Meta-Attention introduces per-token Bayesian routing among attention mechanisms via amortised variational inference with a Dirichlet prior, yielding lower projected FLOP cost than prior-free routing on a Tiny LM benchmark.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance

cs.AI · 2026-06-08 · unverdicted · novelty 5.0

SIFT precomputes selective attention indices via local and cross-attention invariance to speed RAG prefill 1.71x while keeping accuracy within 1% of full recompute, storing only bit vectors 24,000x smaller than KV tensors.

citing papers explorer

Showing 1 of 1 citing paper after filters.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cs.AI · 2026-06-08 · unverdicted · none · ref 42
SIFT precomputes selective attention indices via local and cross-attention invariance to speed RAG prefill 1.71x while keeping accuracy within 1% of full recompute, storing only bit vectors 24,000x smaller than KV tensors.

Lampinen, and Stephanie C

fields

years

verdicts

representative citing papers

citing papers explorer