Pith. sign in

Paper Citation Record · LEDGER

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator

As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.00797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00797 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:15:08.413962Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82e96aaf-b95a-4a4a-92a5-11c2a30b477d · outbound

This paper cites GPT-4 Technical Report.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.849464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.849464Z digest=sha256:8781f502a5e5cdd89a642f37ce88d49149f7e72e0ad7a927b78f4d6fba836c29

Observation 72b6c78e-1280-462d-a626-1b509ed01e59 · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.326845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:07.859756Z digest=sha256:836aa5dedac88b42e81076501a2202d5a1214c16e09f4e8d1e04cf7d09a1c642

Observation f0a9f903-772b-4d81-ba86-1f5b0b72b145 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.260429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:07.881035Z digest=sha256:0388ab1dbe98661ee03d1e174d776cac40a47e134fe8cb71cd18f30774415786

Observation 1f6b3c51-5858-4aa2-aa87-8a6e3bdecb84 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.905855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.905855Z digest=sha256:c9d9664117c79f415a9035b2ffab1dac1a797f3a677c4b1e44e8e3267911067e

Observation af994ca5-914f-436a-ba39-997971429553 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.925788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.925788Z digest=sha256:a2a041f52c57627e9a3a90c1a36942fd3d168cdccf1bf8d8ca4bafe1c0d1014a

Observation a266e578-3faf-4682-8594-a64aaca11cd0 · outbound

This paper cites Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.936932Z digest=sha256:5ca3446bcad9a30143c4a340b122a870f55a8c147ec038bf59472ea60184e580

Observation bce1e27a-5595-499d-b96d-16313904e1ae · outbound

This paper cites Ramulator: A fast and extensible dram simulator,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Ramulator: A fast and extensible dram simulator,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.122654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:07.946905Z digest=sha256:45864cdeecce91bf9aa2b4b608d5e8882619134c869d6687106f3fa6a6ca1f5b

Observation c13cd7a9-0e80-45d4-a83b-1f4ead0f21e0 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.045540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:07.979589Z digest=sha256:a501de225001e6b35b360ade4225babf1284e45dc5bafa8591656769e1d7fe20

Observation 8c532b4d-d698-4c75-91b7-32fba6c84da2 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.952344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:07.998054Z digest=sha256:285199cf1e5504de7ec9f37fe4d05deb5122033b66ef3d6b4cedc32b9f659afb

Observation 17c0dbb8-0ca1-4db4-ac07-30d9a18fb3bc · outbound

This paper cites Online normalizer calculation for softmax.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Online normalizer calculation for softmax

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.020611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.020611Z digest=sha256:53ce30a493bdbbdd3b20bf5530b20f35b6a66e9d0ef799b6bf42b34fd927e99e

Observation 2727282a-607f-4d4a-ad12-ec698b5a3a26 · outbound

This paper cites Cacti 6.0: A tool to model large caches,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cacti 6.0: A tool to model large caches,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.845873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:08.049142Z digest=sha256:15d3f61e7e1663bc5b8130d6f013ea171cba4a3a38ed7760ff35238cf04b5c92

Observation cfb3fef2-0844-4581-9a1c-100168a196e5 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Compressive Transformers for Long-Range Sequence Modelling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.085353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.085353Z digest=sha256:5a615665a8173832ed01b3d142168abb5e14da677be1bdf15f314669d0501e4a

Observation 4613ce40-fd13-4dfe-b3a6-1785a1e4d68f · outbound

This paper cites Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.788210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:08.129401Z digest=sha256:4a985a1d65c0964baf067d6681fd1b49dc4e303c0ee6ad61f6afc22d0cc17b25

Observation 695b8885-89e2-48f5-ae99-a1b45604d5f1 · outbound

This paper cites Softermax: Hardware/software co-design of an efficient softmax for transformers,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Softermax: Hardware/software co-design of an efficient softmax for transformers,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.166263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.166263Z digest=sha256:f287c4a01cd9b5bf5e49b6ce4eceffaa1232671713c9d462412a56b7c9aa75cb

Observation 39010cac-1c04-4604-bd31-4cd13cf2be2a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.202870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.202870Z digest=sha256:0a8d3a41a140a6444e278d56577288d3656f17c74ac47071a173b370e5ca8777

Observation 06dcc06d-a9e5-492e-b322-38fc81ecfe8a · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.237183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.237183Z digest=sha256:6f62c2518e81ff112d3095f7a8892f154fa73bd5c01f0363ac4ec8fefb63dd51

Observation 77954335-5988-4f1b-bf7e-71adab422002 · outbound

This paper cites Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.695573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:08.274907Z digest=sha256:00aec28a842f70363ddc05c8b7e77d85e10bd82b4821520e974160d63342033c

Observation 6f659055-ae9e-4c17-9b50-cecc7f8b9193 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Efficient Streaming Language Models with Attention Sinks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.309519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.309519Z digest=sha256:d7c7e7d4b6de26da1716065c6ce3befdf50cfadaea3e9e53abad68df0d69ead2

Observation 7a12d836-3440-47af-b3b7-859070fee557 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Orca: A distributed serving system for {Transformer-Based} generative models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.638398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:15:08.345273Z digest=sha256:dd1026fd96cdf22cd7984889833750756581d619917649c4d3f1f2007e23e009

Observation 4e5f3e1d-e7ca-41e7-b216-ccac6a4f76db · outbound

This paper cites NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.374201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.374201Z digest=sha256:1b8cc688419a8d8e0d5c444a403c94654d0a9fd190de467c450f774104b442ec

Observation 9d67cf50-a008-43ef-8f0b-22996fae256b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.413962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.413962Z digest=sha256:bfb723c5c7977deb4be8b5ef26c2a2900cac93749eb1b6331b2321c5a201502d

Pith citing papers

No inbound Pith citation observations are available.