Pith. sign in

Paper Citation Record · LEDGER

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing

As of 16 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.16187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16187 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:43:44.403049Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:59:44.970855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:19:47.322402Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3651cfa-f343-4867-bd9d-09a9e894be82 · outbound

This paper cites A Survey of Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.737452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.737452Z digest=sha256:3c2fc4f6975d09ff2728662825164838fc4e3e940517c965b173f34edc5d0864

Observation e4236455-69ec-4db3-aee9-8a24bd2391c5 · outbound

This paper cites Emergent Abilities of Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Emergent Abilities of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.743889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.743889Z digest=sha256:74a032d2bbe2611d21b312cc8225e64fedd9641f7489138279de9c5aa9f681c2

Observation 5ecfb2fd-1cef-467e-a389-84d385e54f09 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Neural Machine Translation by Jointly Learning to Align and Translate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.749297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.749297Z digest=sha256:23e1495e71ecaea3a424260e5003640fa43cfc99cb87d30d55ca37e6562ef0b3

Observation c6e600da-c089-4186-8d2e-7bc171bac888 · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Effective Approaches to Attention-based Neural Machine Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.754816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.754816Z digest=sha256:cf53b3382ef162217c68ddd047bc0e5c963f60e49c88ee5f2658091f4c09a4d5

Observation 5de0ce26-c5c2-4de4-9d7e-d75b3d7c9e92 · outbound

This paper cites Attention is all you need.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Attention is all you need

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.759942Z digest=sha256:85b128acb38edd7b3118024262230e92ece40a5957fe0baa3dda9985d85bd1e2

Observation f14c6fb9-6a5b-425e-bd20-9dd21a2ac670 · outbound

This paper cites The Llama 3 Herd of Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.764731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.764731Z digest=sha256:27a5939b7083239ed8beeabed014ace07c29016561f2a0634154d0791af4eab6

Observation a3155492-a56a-4612-9f47-99e5983436d1 · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.770122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.770122Z digest=sha256:2058bd4e7b8a635e61c6e234fb0d9a0362de56c91d5d0159e3d06690da33e4a3

Observation 9f047785-3d64-4c5e-b6cf-04113097f87a · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.775210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.775210Z digest=sha256:be54e9c3cd66dcf8db89bc7a0770914a965671036a3f5d47a42eabfb2c5e531c

Observation fc3e1cb6-b53e-47b0-b479-de0e621e0bf0 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:45.010530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:43.780015Z digest=sha256:2a886d9c117f02b27906b1641d64c70b62a16d3c0739dd38421a99ac252e807c

Observation ed8e8fdc-114e-47f7-9766-1afe5c63fb3d · outbound

This paper cites Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.784590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.784590Z digest=sha256:5dd737d49a55c5841fa2de15301cb4f6c298c2810f6a4b0d376dfeb7aa9eb808

Observation 7fdb926b-1983-4b08-9db8-0cf51bb94c95 · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.823950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.823950Z digest=sha256:9c8a71e2944ba1b09d7eef7ab0087090f7e8c5e17b253cdbf97d3a13fc513559

Observation 6f9099a9-fbee-462b-849f-c7d79f79ca8b · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.867561Z digest=sha256:ae40be92b929ba08ff2dcf7660a6be60292625d61ffe5bc292cc9a66f7ccb812

Observation 62512a98-89f9-406d-8b78-2f6820c8fcfd · outbound

This paper cites Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.980023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:43.904314Z digest=sha256:bb99043ff147dda59b3302df19733889b35b75ebd020a9ca0d07e0f53200dd19

Observation 1189bb93-5a59-4543-a9cc-6b2b413774fe · outbound

This paper cites Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:43.942630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:43.942630Z digest=sha256:f2953d18fcff8b90e521d8ff1b249f5a588db91cccfdb35ce466c21c6fb03b15

Observation eecaee27-f217-4d3e-beeb-5777a5c3c870 · outbound

This paper cites Similarity estimation techniques from rounding algorithms.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Similarity estimation techniques from rounding algorithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.953190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.005871Z digest=sha256:c7d6b1d5652b3c930ed3498d631e705dc6cdbf554b12177f3b4e3d51fc1cf734

Observation efd09993-04b9-448c-ba87-7eaffe876149 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.043461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.043461Z digest=sha256:c057b63d9638ed0d30ac040793a3269fd61f6fe62b9a9704d40da14cb963e45b

Observation 2181495a-8264-4fbe-b2e9-54dbc2db228e · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.096202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.096202Z digest=sha256:6fd39ebc695d06c6bf0f3cf0f8aea4a2b40ffb11e62c6bc18a91f5564c75b492

Observation 287c502f-ddb5-470d-9d48-50824a674d0b · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.101225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.101225Z digest=sha256:c3bfe8d73185882efbf707c0e3a54fafc4b0b2ba59bc3d70da95fba56dd81dc6

Observation 299636b4-cc94-4354-bb92-d4b2ea6fd0cf · outbound

This paper cites Formal Algorithms for Transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Formal Algorithms for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.106277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.106277Z digest=sha256:8f45106e2ce3f2fe51a41f0cccd12eb8a791e6b4f2f909e6cd3d557330c0c4ed

Observation 00431976-5634-4013-bab3-96c8d29c84bf · outbound

This paper cites Approximate nearest neighbor search in high dimensions.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Approximate nearest neighbor search in high dimensions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.935968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.111826Z digest=sha256:469ef14dfe379ad964d138d4b7c3351de243f2c330ba04f9be5719d05e11a166

Observation 0a464099-9a06-4753-acd4-6040b52c7b9a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.920261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.118585Z digest=sha256:78cdad6b48f6208e6ee5111f37b1ef137f0866736a679af21be3b3b7bf9f6a21

Observation 73c1e8c1-98c6-4410-a6ba-076d25df1217 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.122681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.122681Z digest=sha256:996a4513782967b83a1ba5791dae45b94c06c96f7a66a626d80603a6d65eea2f

Observation 05f3dd23-7762-41ab-9af0-b38f36eee298 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.127280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.127280Z digest=sha256:3421c4ccd1e1cf020165fecaef3c9a1aede9c3aeb10340cafa9b00b93bdbbbd9

Observation a09a16a1-07cc-4ba0-a737-ae241f0805eb · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.131384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.131384Z digest=sha256:ad3c32e121874a13f2fc3f06d44003e640f575df578c27232f95c469a632010c

Observation 13242575-7e09-40ba-91a5-e84f8aaaf053 · outbound

This paper cites Reformer: The Efficient Transformer.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Reformer: The Efficient Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.135397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.135397Z digest=sha256:649eb5a3f1cb4f56ffa1a6759f2693c233a516302942244c08b289d9d9892323

Observation c9f6a8a4-f159-47b9-8eef-e069afea78ee · outbound

This paper cites Kdeformer: Accelerating transformers via kernel density estimation.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Kdeformer: Accelerating transformers via kernel density estimation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.905565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.139547Z digest=sha256:eac60b6b5f420d910dcc8b81514f9a3d3e3f7b0d31f870e4e06a7d034c080e34

Observation f3fb8b69-0aea-43c6-bc60-86eeda4cfb5d · outbound

This paper cites HyperAttention: Long-context Attention in Near-Linear Time.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing HyperAttention: Long-context Attention in Near-Linear Time

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.143968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.143968Z digest=sha256:18526210a9a49abdcc8489594a3d979183210cb4ee2b48378f8f2ae55ec1d142

Observation 4193d3ef-d509-43dd-bfaa-a13c18145727 · outbound

This paper cites QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.149261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.149261Z digest=sha256:b415aa68cd8acc714f20641e19790990d838e6e20567f78ceebca39d02b52eaa

Observation 5ad57403-d026-4087-add5-b4b343916b03 · outbound

This paper cites SubGen: Token Generation in Sublinear Time and Memory.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing SubGen: Token Generation in Sublinear Time and Memory

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.153690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.153690Z digest=sha256:068b506a8f7749b25c855a3cb8db9044358ed781ea45a7494d65ba1dfeaebc33

Observation 50d55882-1fdf-4ad4-af3e-05cd0c81f054 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Fast Transformer Decoding: One Write-Head is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.157913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.157913Z digest=sha256:f4a8770ba33007eb4b188548366e4876c8237c71ee29cbc9723e964683d1dff5

Observation 66bd7130-778b-4b9b-b5f9-bb1cb30f5782 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.162332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.162332Z digest=sha256:a70cd30034207846cb07e1e7c28b003f9b6b0a402ccf0960f8deb09f0f89ace7

Observation 44d98f3e-24f5-472d-9ddd-3fe1da577ed0 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:44.891314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T16:43:44.166732Z digest=sha256:89e75ae211b6313bf61ecdeb701943a9f1bfe6bc8cb56b709fa27a31858eacd8

Observation e5f23e26-9437-44f2-b0a4-a624e0d6bf93 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.170735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.170735Z digest=sha256:9ec0b1cd2cc5ba21ca8847102c517c92c97ebc4f98c039cfe2362e22df6452fc

Observation 5fe69c5c-7bc4-4766-9cf6-10f572dab38b · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.201438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.201438Z digest=sha256:c9de134d9e66ef792d92500342280b8d2ac097fc1c5e74683fca12cd0f51c9ac

Observation 4c556cb1-c18e-4442-97b0-f8adf25b3dc3 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.246679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.246679Z digest=sha256:e2fd90003911b7fab1d352755ad30b4e1d89b462ddfc501d9d4b4b54798cf71c

Observation 18bfa255-40bd-434b-9c8e-78b584cb5fb7 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.266533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.266533Z digest=sha256:a555522806a25f5e490f1c3e921cba2b5b46aa68312c5b95a190ebc9df8dae7d

Observation aab0e56d-570c-47d0-9efb-655009d2b5e4 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Generating Long Sequences with Sparse Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.304635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.304635Z digest=sha256:d7c8f4c076b1317e2e34e74239dc10f23dc19df8012b36090ba42efb4ae91e73

Observation 481a80fe-1da7-4bcb-9149-bf84ce27f53a · outbound

This paper cites Longformer: The Long-Document Transformer.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Longformer: The Long-Document Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.392915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.392915Z digest=sha256:09019ec445354f9c0d4145a27191ecd9d1cc3ed574b70d84c400d1351ed6f916

Observation a80dc761-5a64-46f3-98dd-5e7ac4511eb9 · outbound

This paper cites Memory-efficient Transformers via Top-$k$ Attention.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing Memory-efficient Transformers via Top-$k$ Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.403049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.403049Z digest=sha256:fc28f327cc4f3655ae5a7d1e9b065b3aad60ddf84b0c50c0ac962dc77eaf86e9

Pith citing papers

Observation fe916e22-16d0-46c3-9435-07c511f31da9 · inbound

SpotAttention: Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers cites this paper.

SpotAttention: Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.323749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T08:59:44.970855Z digest=sha256:97069a48cd4767ba4c4cb73b3606a0e7f744eda68b6e9da90b9f175d96b25e94