Pith. sign in

Paper Citation Record · LEDGER

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing

As of 16 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.16187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16187 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:43:44.403049Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:59:44.970855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:19:47.322402Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3651cfa-f343-4867-bd9d-09a9e894be82 · outbound

This paper cites A Survey of Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.737452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.737452Z digest=sha256:3c2fc4f6975d09ff2728662825164838fc4e3e940517c965b173f34edc5d0864

Observation e4236455-69ec-4db3-aee9-8a24bd2391c5 · outbound

This paper cites Emergent Abilities of Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Emergent Abilities of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.743889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.743889Z digest=sha256:74a032d2bbe2611d21b312cc8225e64fedd9641f7489138279de9c5aa9f681c2

Observation 5ecfb2fd-1cef-467e-a389-84d385e54f09 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Neural Machine Translation by Jointly Learning to Align and Translate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.749297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.749297Z digest=sha256:23e1495e71ecaea3a424260e5003640fa43cfc99cb87d30d55ca37e6562ef0b3

Observation c6e600da-c089-4186-8d2e-7bc171bac888 · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Effective Approaches to Attention-based Neural Machine Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.754816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.754816Z digest=sha256:cf53b3382ef162217c68ddd047bc0e5c963f60e49c88ee5f2658091f4c09a4d5

Observation 5de0ce26-c5c2-4de4-9d7e-d75b3d7c9e92 · outbound

This paper cites Attention is all you need.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Attention is all you need

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.759942Z digest=sha256:85b128acb38edd7b3118024262230e92ece40a5957fe0baa3dda9985d85bd1e2

Observation f14c6fb9-6a5b-425e-bd20-9dd21a2ac670 · outbound

This paper cites The Llama 3 Herd of Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.764731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.764731Z digest=sha256:27a5939b7083239ed8beeabed014ace07c29016561f2a0634154d0791af4eab6

Observation a3155492-a56a-4612-9f47-99e5983436d1 · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.770122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.770122Z digest=sha256:cb05de774e4c4f92f9d915a1d55f521601c182b0c86bd7a6eb45e98960a38895

Observation 9f047785-3d64-4c5e-b6cf-04113097f87a · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.775210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.775210Z digest=sha256:be54e9c3cd66dcf8db89bc7a0770914a965671036a3f5d47a42eabfb2c5e531c

Observation fc3e1cb6-b53e-47b0-b479-de0e621e0bf0 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:45.010530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:43.780015Z digest=sha256:3f1a6d04292925061db663a848140ea6c1dd450f246ae977ad26c6d22f419ff7

Observation ed8e8fdc-114e-47f7-9766-1afe5c63fb3d · outbound

This paper cites Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.784590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.784590Z digest=sha256:5dd737d49a55c5841fa2de15301cb4f6c298c2810f6a4b0d376dfeb7aa9eb808

Observation 7fdb926b-1983-4b08-9db8-0cf51bb94c95 · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.823950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.823950Z digest=sha256:808968e86327da2c38dc4fbf49501a5725655dd16d7bdcbff966d40f412b0e45

Observation 6f9099a9-fbee-462b-849f-c7d79f79ca8b · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.867561Z digest=sha256:a65dca4a0b97154dc3e93ad63454586ce97a4e7dde4100986c9063db09a986dc

Observation 62512a98-89f9-406d-8b78-2f6820c8fcfd · outbound

This paper cites Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.980023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:43.904314Z digest=sha256:a073f9ecfcc06d09240026a1157d28e8cae2702f25f3b01dff09dd2f6e154e4e

Observation 1189bb93-5a59-4543-a9cc-6b2b413774fe · outbound

This paper cites Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.942630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.942630Z digest=sha256:f2953d18fcff8b90e521d8ff1b249f5a588db91cccfdb35ce466c21c6fb03b15

Observation eecaee27-f217-4d3e-beeb-5777a5c3c870 · outbound

This paper cites Similarity estimation techniques from rounding algorithms.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Similarity estimation techniques from rounding algorithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.953190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.005871Z digest=sha256:31e209258f2a990ef7ee72beb08a75df082f5858b81d5835cf9cf2b0433085ef

Observation efd09993-04b9-448c-ba87-7eaffe876149 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.043461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.043461Z digest=sha256:c057b63d9638ed0d30ac040793a3269fd61f6fe62b9a9704d40da14cb963e45b

Observation 2181495a-8264-4fbe-b2e9-54dbc2db228e · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.096202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.096202Z digest=sha256:6fd39ebc695d06c6bf0f3cf0f8aea4a2b40ffb11e62c6bc18a91f5564c75b492

Observation 287c502f-ddb5-470d-9d48-50824a674d0b · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.101225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.101225Z digest=sha256:c3bfe8d73185882efbf707c0e3a54fafc4b0b2ba59bc3d70da95fba56dd81dc6

Observation 299636b4-cc94-4354-bb92-d4b2ea6fd0cf · outbound

This paper cites Formal Algorithms for Transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Formal Algorithms for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.106277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.106277Z digest=sha256:8f45106e2ce3f2fe51a41f0cccd12eb8a791e6b4f2f909e6cd3d557330c0c4ed

Observation 00431976-5634-4013-bab3-96c8d29c84bf · outbound

This paper cites Approximate nearest neighbor search in high dimensions.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Approximate nearest neighbor search in high dimensions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.935968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.111826Z digest=sha256:160f6d7d3fdb1771e35f94563869130401e4af12d2b0dd11b2a93ab93ed94c37

Observation 0a464099-9a06-4753-acd4-6040b52c7b9a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.920261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.118585Z digest=sha256:b8e5628b12a822066d9a6b92d5806e49847796fc8eafe83306a022219381f3f8

Observation 73c1e8c1-98c6-4410-a6ba-076d25df1217 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.122681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.122681Z digest=sha256:996a4513782967b83a1ba5791dae45b94c06c96f7a66a626d80603a6d65eea2f

Observation 05f3dd23-7762-41ab-9af0-b38f36eee298 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.127280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.127280Z digest=sha256:01a0457cb2da816c3407965e4603f09e2d34af4030fef3bd9e8b45e1b01d980b

Observation a09a16a1-07cc-4ba0-a737-ae241f0805eb · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.131384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.131384Z digest=sha256:1b39656198f5cd00936689216a45859fe6048ee5acec9c96ac2b46f35a7fbd63

Observation 13242575-7e09-40ba-91a5-e84f8aaaf053 · outbound

This paper cites Reformer: The Efficient Transformer.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Reformer: The Efficient Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.135397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.135397Z digest=sha256:649eb5a3f1cb4f56ffa1a6759f2693c233a516302942244c08b289d9d9892323

Observation c9f6a8a4-f159-47b9-8eef-e069afea78ee · outbound

This paper cites Kdeformer: Accelerating transformers via kernel density estimation.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Kdeformer: Accelerating transformers via kernel density estimation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.905565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.139547Z digest=sha256:03251363157ea53c8aa17906008f2b5046a28e5beb2a2d5a50dfa5ca4fe67081

Observation f3fb8b69-0aea-43c6-bc60-86eeda4cfb5d · outbound

This paper cites HyperAttention: Long-context Attention in Near-Linear Time.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing HyperAttention: Long-context Attention in Near-Linear Time

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.143968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.143968Z digest=sha256:8f7a1518acdb798dbdcb882c888ccd15bfddde48aff4468f3f48b1a00d495448

Observation 4193d3ef-d509-43dd-bfaa-a13c18145727 · outbound

This paper cites QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.149261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.149261Z digest=sha256:e0e96257a7be4f700a0a3f08e3254060c6c1da5e6c6eea6d8ba4e571670e325e

Observation 5ad57403-d026-4087-add5-b4b343916b03 · outbound

This paper cites SubGen: Token Generation in Sublinear Time and Memory.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing SubGen: Token Generation in Sublinear Time and Memory

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.153690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.153690Z digest=sha256:fefd1d50b904632f65e343e26d473fb901c6222c8dfe8770f84435ed7db5d995

Observation 50d55882-1fdf-4ad4-af3e-05cd0c81f054 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Fast Transformer Decoding: One Write-Head is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.157913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.157913Z digest=sha256:f4a8770ba33007eb4b188548366e4876c8237c71ee29cbc9723e964683d1dff5

Observation 66bd7130-778b-4b9b-b5f9-bb1cb30f5782 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.162332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.162332Z digest=sha256:bf2e93010b0d141424a114f2fbc2ac93a697997cd338df6eae91a29230a0bc13

Observation 44d98f3e-24f5-472d-9ddd-3fe1da577ed0 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.891314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.166732Z digest=sha256:aa43350c63837fd26a45725d4ad7e7e57bd74addd17bd7e036729a5ea950204e

Observation e5f23e26-9437-44f2-b0a4-a624e0d6bf93 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.170735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.170735Z digest=sha256:9ec0b1cd2cc5ba21ca8847102c517c92c97ebc4f98c039cfe2362e22df6452fc

Observation 5fe69c5c-7bc4-4766-9cf6-10f572dab38b · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.201438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.201438Z digest=sha256:f11c270a0e26d4829aa9dc428a5cc300074099df22c09ff0a1908b2e88603246

Observation 4c556cb1-c18e-4442-97b0-f8adf25b3dc3 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.246679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.246679Z digest=sha256:e2fd90003911b7fab1d352755ad30b4e1d89b462ddfc501d9d4b4b54798cf71c

Observation 18bfa255-40bd-434b-9c8e-78b584cb5fb7 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.266533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.266533Z digest=sha256:a555522806a25f5e490f1c3e921cba2b5b46aa68312c5b95a190ebc9df8dae7d

Observation aab0e56d-570c-47d0-9efb-655009d2b5e4 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Generating Long Sequences with Sparse Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.304635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.304635Z digest=sha256:fb2d49ede7f8c52249f836740d3054751405640190fcf22be1d7d7d19df51952

Observation 481a80fe-1da7-4bcb-9149-bf84ce27f53a · outbound

This paper cites Longformer: The Long-Document Transformer.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Longformer: The Long-Document Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.392915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.392915Z digest=sha256:09019ec445354f9c0d4145a27191ecd9d1cc3ed574b70d84c400d1351ed6f916

Observation a80dc761-5a64-46f3-98dd-5e7ac4511eb9 · outbound

This paper cites Memory-efficient Transformers via Top-$k$ Attention.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Memory-efficient Transformers via Top-$k$ Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.403049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.403049Z digest=sha256:fc28f327cc4f3655ae5a7d1e9b065b3aad60ddf84b0c50c0ac962dc77eaf86e9

Pith citing papers

Observation fe916e22-16d0-46c3-9435-07c511f31da9 · inbound

SpotAttention: Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers cites this paper.

SpotAttention: Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.323749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T08:59:44.970855Z digest=sha256:ebef7481a6cd4d125dd7260380d8ef6d8d31561af1e5a6a4db8e42ae6936cf89