Pith. sign in

Paper Citation Record · LEDGER

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.02572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02572 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:02.274620Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:48.458849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:45:48.747511Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00d6d16-d450-434f-9c50-f55eeaeb2bb2 · outbound

This paper cites online" 'onlinestring :=.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.349731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.349731Z digest=sha256:e75b5ce297df183141e52aadba694ab0db913f9a477fb107333662ffcb6542f3

Observation d0264211-adc8-4dcf-9e69-f46c3df0a536 · outbound

This paper cites write newline.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.459565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.459565Z digest=sha256:999c363f708e77e7da9140277d30792c9370401987ea56e60a8c026b1fe60bf2

Observation 28ab32b5-96e1-4cf7-8a59-a4db1e863739 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.675401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.675401Z digest=sha256:69024992e72033db8e2a9c6f4a88b096ea6d9b3eb5c20309c7b264f27b83cc5c

Observation 4ef753db-df26-4f3a-ad33-345fced6e436 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.843590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.843590Z digest=sha256:e133be6de669b6b76cc049a2b12b3e0fb81da47a3026c18649876b26919fd220

Observation 5258d014-ba2d-4a4b-b4e7-f30b0cc1acb1 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:57.958618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:57.958618Z digest=sha256:6eabb2607a095f559f44a1eba52b4d53fbaa59a7475b28b5d38a7cf1322965ff

Observation 9867a5ec-f2d4-463f-bc37-ca11d9540298 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.070396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.070396Z digest=sha256:4d51f748b07ff20b904a3df00fae011a6159b0462f51d77a4e093d6de5fac78e

Observation a455538b-71a8-48d8-88b4-fa44995695e9 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.308531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.308531Z digest=sha256:75f3be4062f6d4900d203540c82a32de37fb3f6e7561f46b780b0684378b4add

Observation e84205aa-0633-4f08-9907-660047ea07d1 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.446766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.446766Z digest=sha256:052b2a7b47d497cf6e8bea1e7187fc648c2c438fe894a70eb0cd5cb923fc0f8f

Observation 0b527f51-05cf-4587-8bc5-3bf20c84f762 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.692270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.692270Z digest=sha256:7be44c78ed08979c352ca6b9fb3bb4ce2cde14062041aa66d3709eabee177941

Observation ab5eca77-3f83-46e3-8a11-8324063aa0d8 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference HashAttention: Semantic Sparsity for Faster Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:58.886175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:58.886175Z digest=sha256:214b368ced6a30c1c1c90ee647301d29c9bd2ceb7e8e9be67b52712d65e7e2b0

Observation af753a9a-dbfa-4988-8115-d17d67229f61 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:06.014550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.049157Z digest=sha256:5009a7e0703f6aace39fca70c9918e64442875c758f74617001f36c152a26467

Observation 8629c365-fbe6-4024-9823-fe469e099adb · outbound

This paper cites Memory-efficient Transformers via Top-$k$ Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Memory-efficient Transformers via Top-$k$ Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.218497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.218497Z digest=sha256:02cc70d8ff4157095f530a17996201321705dea644f664a4e21f06a458e2c12b

Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.316579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.316579Z digest=sha256:dd3d5a60a769255f0059411bc49671db2e708c4b184cf7f371746c3d87703a01

Observation 98153a66-8728-42ce-93cd-e39469ce9dac · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.386831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.386831Z digest=sha256:aba4c2477fba6c342722d6963353b05e5af7dae0a8e8ca4c16edcae7a9e1f12c

Observation 8f6999e4-445f-4533-874c-ec85146d0447 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.794332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.455430Z digest=sha256:bfd9718cc1f05801ee97e7c1210b78d4c05d58295a86e2d23c968322d38c1d98

Observation 1edc92a5-580a-4b71-85b4-741ef6a3ddd3 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.546700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.604558Z digest=sha256:10e530c5056a116d41c05d4cb2657f9610686ccdc651d9afb03af4abcba69982

Observation 02e562ac-2889-411f-b68f-268dde24cd2e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.706591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.706591Z digest=sha256:eb8206ea6a2f5e6b08aba9de407f08bbcafa5204eed760d6f905089b4be8222e

Observation 012433c7-92dc-48a1-aabe-e548db1efe5e · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.347082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:25:59.830867Z digest=sha256:acaf7cb5bfe3d881c530acbf17d06eb3319736eb5d981c414f5870c146d9d920

Observation 37306390-c055-4a70-999b-5d757f92dd6c · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.949662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.949662Z digest=sha256:fe39c194e24ea8a1bd2d71a974c6018082bfc5b922461cc0012391797c70cacf

Observation 098a8c1b-004e-4148-b148-d46686046608 · outbound

This paper cites DeepSeek-V3 Technical Report.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.038860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.038860Z digest=sha256:c7fcf3e9c56b38f35c0591448a5c407c899ff80e269746247db0f83b4ae9a622

Observation 2b7360b8-110d-47b9-a140-000b9b0cab85 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.133693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.133693Z digest=sha256:8a80bc62f46a1a0d982f468f84d32c9ebfada60a4bb499adad1754ae28d61321

Observation 9117e012-3f1c-47d6-9b79-58783242eea5 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:05.128766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.262542Z digest=sha256:c84f562285af0259b96f0c07fb48ea24c42256af8680b25e65aff54b61146ea6

Observation 4d96180c-7e34-4588-b55a-2bb8ac645924 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.945624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.355214Z digest=sha256:e53b5fc46a07ead2ddb6a03dd0dcd13da1a515eb3f65734827b1ce382f8ead90

Observation 39d47bc8-9f11-4b4f-a4d5-29fc19b45eb6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.729592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:00.450652Z digest=sha256:5c210eecf073a863c30abc4468a1cca48aa55fa6f63eef64796b26299bcf539b

Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.637951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.637951Z digest=sha256:48a6ddae57d1725ef7e1a0298489cf26f40ed731ca4d72964b09b8a6513f45e5

Observation 0e03dd71-908c-4e4f-a12d-e05dcd4523fe · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.727994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.727994Z digest=sha256:8d3ca2d16f70370d2bd76a641dc444ce360626c41a8881ada131e7d5a13a85a5

Observation db9cac96-1645-4357-b4db-31300544deb0 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Loki: Low-rank Keys for Efficient Sparse Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.797982Z digest=sha256:625721cc59b953b5fc53cfcff74f4db99fdc03cf4a8d338f59164fd7cbf9fd48

Observation 2178ee9d-6c7e-4b91-bb98-be1e9a5b5e6b · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.895653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.895653Z digest=sha256:ac6639ecfede48290b9bb6e4c5a407c8aec61068630562f5fd434c57cee800c9

Observation 1537f08e-129f-4807-ac9a-fcca9ebf512d · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.503621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.047307Z digest=sha256:7e896d8beec6fa75b0d89bba836a02ed424b347c84bfdc45989658e541775a23

Observation f8c98f77-f5ad-4674-bfef-2c7c10eb8831 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.145607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.145607Z digest=sha256:40fd3d166a0a84e4533152afc845ba6551e64bf510b0db78bd73f587f5b6fbbb

Observation 42e625ea-5c24-45e6-899b-70ef97c44005 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.299814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.299814Z digest=sha256:0e520f16db592052bd1108afafa6a6cdd637a2861a67e2dc6f080398962ff064

Observation d50ea01c-306d-4f00-ad0e-a8fda0e12902 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.278953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.430394Z digest=sha256:0e3d295eef362bf011d6f54696e06f872963cf0b57998f139bed09ff613438a0

Observation 266518e8-ac32-4b3a-aa97-26293bc11f2a · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:04.091956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.499338Z digest=sha256:c560b938540cbadf654b2464ce804b3c6faa9f199b603fc6687f1451f5c9a0aa

Observation e3fd5a6e-c4bc-4898-9132-2d7c822408d2 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.842301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.598450Z digest=sha256:82309a2a249b45c79ac0b8b49c94695aa21cadc0b0f2958a2f95d395a86f82ee

Observation 536b9b5a-dd44-432e-93ed-2e855ea981e6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.625450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.694965Z digest=sha256:705e9b56a3b7be2dfd182eb25864850db9b9674a4a113e941a8251eebd799f6d

Observation aa457d13-4aa4-42e0-b74b-cda988a847f4 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.762361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.762361Z digest=sha256:918e583bbc5a2ae89690b3c657351a7e39393689f2fdee9cfd073d9793f7614f

Observation 8438d211-1c14-4880-88be-6b85ebcf3437 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.860825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.860825Z digest=sha256:7fadc18314a17d690b0cd08d9e646bbe146a28fe472fe91d038fcdda30a59b69

Observation 823f57db-0166-4562-9bd7-c85d58cb7e38 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.351191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.933392Z digest=sha256:e9bff48aacf0ce0c5a039c18f7ebc8dabad0b79e646c8287d6fb35aaffda7ed2

Observation 99128bbc-79ff-452d-8e15-744b868cf3bf · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:03.056187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.028935Z digest=sha256:c61e1e92b2a07353e61f36703ae3ac45a4cd56d5fe8ecece585e15353d0fb95c

Observation 06830067-bf8d-4a76-ac42-edd18430f3f6 · outbound

This paper cites an unresolved cited work.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:02.786954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.149538Z digest=sha256:8cc4979ad13c4382fae777e6c5884783b3d3f8ef39264d39605a48eae92d4272

Observation baecab88-b21b-4b40-ab88-44eae0a1654e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.274620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.274620Z digest=sha256:22aea64f67500ccd5b1b8aa6dfd1d4fb4f2b32f1152cbafefdd7d361f77acdb0

Pith citing papers

Observation 8a1c847a-d352-4183-97c4-ee08663a825b · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:48.754100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:45:48.458849Z digest=sha256:34af1b28090a8c724a30cd23dd94b019111a01e4f0a01bf8b5035bafde82483e