Pith. sign in

Paper Citation Record · LEDGER

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

As of 13 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2506.15704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15704 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:41.059939Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:03:08.617257Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:04:20.888331Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e9b751b-e021-4306-9155-4dc11094e21f · outbound

This paper cites GPT-4 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:33.909659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:33.909659Z digest=sha256:cd6cba32f525c52be68ae5af6333ce827d8946e30ee4972301fc6cf19a167377

Observation 6905c1fa-97b4-4f8c-8379-e543693552b3 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:34.959797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:34.959797Z digest=sha256:afc14c28b7282321c22522ce1467d3a1533b9774f896bd5b55787d6d0920b204

Observation 4818a542-f672-4751-9d18-0e618417980e · outbound

This paper cites Loki then and now: the trickster against civilization.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Loki then and now: the trickster against civilization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.432776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:36.864775Z digest=sha256:a94093f8538b80c8f2c6a5b40d9698963abf629779d2687db9edab3acddd2e68

Observation fbc23119-1404-43fb-bcbe-a736885ebc2c · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.033773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.033773Z digest=sha256:fcb6b1fc61dcf5f32d2b9b4a54c8f5979e603ae9bda5d2e59e6d2697b9ad31fe

Observation 1b64cb96-4555-4e9f-96da-9b6ea86f7653 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.206520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.206520Z digest=sha256:808615c685e25835cb5f7bfe652067ded016f9723c64ba087323e95d894e274d

Observation 7c429e67-9fbe-4570-8558-ab2d3cd7d6ab · outbound

This paper cites Human-like episodic memory for infinite context llms.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Human-like episodic memory for infinite context llms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.366095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.366095Z digest=sha256:c8994374ca0b69f64c7578d686b3c331495cb547a7e76837e53996708c2ecdfa

Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.472671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.472671Z digest=sha256:7487f3d6294aab3b0e68295cde9f1a7253694e84165614db1e3540abb135f6e0

Observation 550258a3-3477-4a84-96ab-1b4fb5775f14 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.591170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.591170Z digest=sha256:25ebd6d025004749800d5b4cc39461e7c683861c105d75c11911bc1ea34c7f7a

Observation 8dd36822-23d0-4a69-a797-88e0c5ff4a97 · outbound

This paper cites KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.745707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.745707Z digest=sha256:4cde817b4c176798ed7d27cfb3322984154bd0cbd7aad7f3809ed170304628e8

Observation 2f11bb7e-adc0-419b-9367-506515bf257a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.911487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.911487Z digest=sha256:43d3de83e136fa9079c17a4245f597e6be9958000c27c518a3cd7160dfead300

Observation 9e46c474-80d9-4b1b-8ce7-4d9e4401fd52 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.015826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.015826Z digest=sha256:a0de9a312d93a6b9ae4cf7d226b509f8fd8a37cc4b0888c262d2b4bf923ca5bd

Observation a34eb7be-896c-4b89-b6ba-31cd8f00ef2a · outbound

This paper cites Compute Or Load KV Cache? Why Not Both?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Compute Or Load KV Cache? Why Not Both?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.180039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.180039Z digest=sha256:36d49b467e6605525c7b4237e77ff710025063c4a88854a17768efec13f057ea

Observation c1bfc8f9-aeab-4816-b3b3-d2d996cf9159 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient memory management for large language model serving with pagedattention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.354119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.354119Z digest=sha256:c22eb4c4489664cfddd5fc6c44d132408eb555f11bbe082b3aa8942f45245c1c

Observation 661a023e-a340-4058-832c-8704d5177c46 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.493837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.493837Z digest=sha256:d6818af58ae3936c801eceeebad9a04ddd05a0e5354dd981d0e65dcaed519bcb

Observation eada5741-34d4-4ef6-a9c4-bf9fa63b8c1a · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Snapkv: Llm knows what you are looking for before generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.226835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:38.606433Z digest=sha256:ca79798cd433298570e23491de5e62f75f4de017212ce6a58fc173837c187562

Observation 16d82528-4536-421b-acd9-37cd4177d06c · outbound

This paper cites DeepSeek-V3 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773218Z digest=sha256:adc5828f0f244dd892571a16b29ce26ef61ed94897daebb814ef16d560d52ed3

Observation 977bba42-95d7-49de-a137-96176e3d0470 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885050Z digest=sha256:8599cc6691fe82c04c93f40a5005ed9d68dbe5bf731983216ed830064ec4afe6

Observation e878039f-0681-4e7b-9be6-f5d943c4b78e · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010317Z digest=sha256:333358ce5f37eb914e344cfb461d0002207372abe632fb497e5be37796f5149e

Observation bc97e210-8720-4018-962e-cd4d98de5a87 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.114620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.114620Z digest=sha256:d2f4bcece8ee42fc4c6b837363f0c8d67fa78e16ef85d414b645c532dcb8672b

Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.247507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.247507Z digest=sha256:c42b4715d679acb22037f50fdab9cb4710d4bbb89d7862c4a2e84b8c3de2e59d

Observation 952fef15-d388-47c7-aa55-f96d7c1b1060 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414534Z digest=sha256:4f0ea813f63150b7cddc845e9d84ca4f2906256b5332c59b4c94c7ebd3b5c2a2

Observation cc38a9bf-8c56-45a5-abae-122c51efc83e · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.514366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.514366Z digest=sha256:5ffd94d029e23e645a678ee70c13fc24122444d313c510dbde3343e97008322d

Observation 9545cbf4-34d8-46f4-9560-9cacafa53065 · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.632271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.632271Z digest=sha256:fc01dfbd709d71bd06090691ed29d1692c3df7309b2f5f076be16919b492cccc

Observation 04f60f49-951a-40a4-a26d-2e8c88f2a8bf · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786344Z digest=sha256:7814ea1ad3de0508303e3418ef61050fdd0fc6a520e99919945192305842362b

Observation bed9b6f2-dc40-4894-92bb-2aa4d92a1ff7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.932459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.932459Z digest=sha256:4395f03714c82e57b29a0b66a51ae6b52f1bdece7c0181a8ddc1cd9ca0ec401f

Observation 6a23d7f7-75e2-45b2-9981-415acf380d00 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.064295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.064295Z digest=sha256:5fcdb96d2bb8ad48f8c2af404e0156f4e0ad71149e26fac82c7ab17cd6c7b421

Observation 9b3bc751-1820-4855-b643-3b526a4061b9 · outbound

This paper cites Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.973919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:40.195808Z digest=sha256:13a191af8d1fd97ac77b79dce22e6cacfb99855e7aec181fe86cc434afe39e4b

Observation aa115202-2d34-45d5-9d2a-faa4234d69e4 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368474Z digest=sha256:1cc7d7764d67db5f039ad7b0609609712575f59fa9505501b4cdf129fbbda17a

Observation d557360f-bcb7-4c09-9807-c01bc5effd5e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient Streaming Language Models with Attention Sinks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.524261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.524261Z digest=sha256:046d4e95f32ed2077eb2bc607c6db9c28b685054ec3cbac5ec028b83a6d5e1d0

Observation 23115589-32d9-4e42-b217-8997271304a3 · outbound

This paper cites Qwen2.5-1M Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Qwen2.5-1M Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663165Z digest=sha256:98bc034bf752c2a32f6200b9ef6f552da6492a3711a39e5614cda72682797bf7

Observation 5437d5b9-1181-4359-adaf-152295e111b3 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Orca: A distributed serving system for {Transformer-Based} generative models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.778601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.778601Z digest=sha256:fa5991d8cfdd74937f476d6f46dff643f48d8f608d20b81145ccedfb3ab34066

Observation 0a5774c3-cbdc-4a02-aab5-6c6eaea9326c · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.900312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.900312Z digest=sha256:ff864b3ff7c9b0f5ec0a2c54786e304aa03fc3ef78d10985ba7f94ef8ac5d477

Observation 0057c0b2-26c2-4e71-9083-60f8c3ffe1db · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.662361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:41.059939Z digest=sha256:896063817ba999a04beec9295946b9c8f6e03a076f3bc7fb8bc33af8a6afcdc3

Pith citing papers

Observation adc68189-911a-4ea2-bb71-b3f964fbeaf9 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.988048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:2fd95fdb6eaad692cd08c7251182c8cc110e191f69adccd87a8ba08d6eb9c56a

Observation b692aa6d-3478-4390-a79c-b57fa3cd51c3 · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:20.890041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:a19ad4a1555fd8a13d79ac747eea0b468d678e8295d407a13f98fc7ca70671d3