Pith. sign in

Paper Citation Record · LEDGER

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

As of 20 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2506.15704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15704 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:41.059939Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:03:08.617257Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:04:20.888331Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e9b751b-e021-4306-9155-4dc11094e21f · outbound

This paper cites GPT-4 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:33.909659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:33.909659Z digest=sha256:da4665c2bbf634777d71fedda5ab328a7c4a525cd6e8c11d576605655d7a0ceb

Observation 6905c1fa-97b4-4f8c-8379-e543693552b3 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:34.959797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:34.959797Z digest=sha256:e449397a123ce57afa5cd12fcfb2e66a5b84174caa56f5e69c54b36ee00506ee

Observation 4818a542-f672-4751-9d18-0e618417980e · outbound

This paper cites Loki then and now: the trickster against civilization.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Loki then and now: the trickster against civilization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.432776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:36.864775Z digest=sha256:f38df5c0225a931b6880431c3ba51574e1af33936c26c28339a6471b3cdae51c

Observation fbc23119-1404-43fb-bcbe-a736885ebc2c · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.033773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.033773Z digest=sha256:a49d558d3faffd143510e91e368d72972c33b0d5b1c83733db1a8997aa59d9cb

Observation 1b64cb96-4555-4e9f-96da-9b6ea86f7653 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.206520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.206520Z digest=sha256:d72e01b6d953ab4be1a799fb604509d8a57dabddc7236973ffa2bfbc9d6e4770

Observation 7c429e67-9fbe-4570-8558-ab2d3cd7d6ab · outbound

This paper cites Human-like episodic memory for infinite context llms.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Human-like episodic memory for infinite context llms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.366095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.366095Z digest=sha256:e56c7b2d01fedcc0b464d4079c04c6b79ea1bf255512325d14add2787b6ef864

Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.472671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.472671Z digest=sha256:1ed7cd76653bae62320091fd812263ee2268e23acd6632f4b57f1438e38f565e

Observation 550258a3-3477-4a84-96ab-1b4fb5775f14 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.591170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.591170Z digest=sha256:2b7537b3b6047b2d5ea5d4613dfdc3fb87d0e65b6e30eb02060eab77c43f43ed

Observation 8dd36822-23d0-4a69-a797-88e0c5ff4a97 · outbound

This paper cites KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.745707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.745707Z digest=sha256:89b984bbe58e6101c870a7b76f6c7794792208b105406f7d7f32af69f045bbd5

Observation 2f11bb7e-adc0-419b-9367-506515bf257a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.911487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.911487Z digest=sha256:4f2b80d6b876566c5c5a10811f95901e47739f25410be7bff2b0525a16739885

Observation 9e46c474-80d9-4b1b-8ce7-4d9e4401fd52 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.015826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.015826Z digest=sha256:86ff17c57f94ac99188c7d35e96bc417ff60e1f44cca49514156242919d9e5f7

Observation a34eb7be-896c-4b89-b6ba-31cd8f00ef2a · outbound

This paper cites Compute Or Load KV Cache? Why Not Both?.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Compute Or Load KV Cache? Why Not Both?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.180039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.180039Z digest=sha256:5f381814a52ad69a7acb41817a0ffabd62b1a792c3103445f145919889a73be6

Observation c1bfc8f9-aeab-4816-b3b3-d2d996cf9159 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient memory management for large language model serving with pagedattention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.354119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.354119Z digest=sha256:98547c18de354137e9d9260a5437b1279c62e733737833e03830e33aa8fbc642

Observation 661a023e-a340-4058-832c-8704d5177c46 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.493837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.493837Z digest=sha256:f57a16a586c5d5e360d41630b65f779ce9cb4b4ab69b4c53adf3bd7ddb1549cc

Observation eada5741-34d4-4ef6-a9c4-bf9fa63b8c1a · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Snapkv: Llm knows what you are looking for before generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:42.226835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:38.606433Z digest=sha256:db1f1dd22274a4c8b5d5c81fd9affa1fb7039113fc00a502c5a752edcfc0ef19

Observation 16d82528-4536-421b-acd9-37cd4177d06c · outbound

This paper cites DeepSeek-V3 Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773218Z digest=sha256:220ff218d7b7701c98a1f1f9096256a2d69e86f38af10682f8323c634e3845e9

Observation 977bba42-95d7-49de-a137-96176e3d0470 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885050Z digest=sha256:113088a44f60dd1000e798407c2fab0acd1a1549c98e5f8c4b22b7e06500e1db

Observation e878039f-0681-4e7b-9be6-f5d943c4b78e · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010317Z digest=sha256:51ed9393945edf6bc42ce4a1e5ffa346efb655d50db46d4ff7697f685dcde500

Observation bc97e210-8720-4018-962e-cd4d98de5a87 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.114620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.114620Z digest=sha256:ae3165bc57d8b010d61f9846ab8a1f636d87f92897a471db2002d331af903c4a

Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.247507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.247507Z digest=sha256:4a4b3772aed5e925c37fb32d2f0540aba9bec733875f45dd755e7fcbf943ec83

Observation 952fef15-d388-47c7-aa55-f96d7c1b1060 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414534Z digest=sha256:838a3be3ff0638244b2e57e343463cc9b83bfc6bd46d22a960170b6060acbda5

Observation cc38a9bf-8c56-45a5-abae-122c51efc83e · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.514366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.514366Z digest=sha256:ce24a089567420ee96d001f36532345e7759ee81311eae7fcf5c6c3c64e18d85

Observation 9545cbf4-34d8-46f4-9560-9cacafa53065 · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.632271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.632271Z digest=sha256:67963f0cdaa20fccb8e773ac771a5c8a1f3867e9c30c6b91f31b6d9fa5d8a984

Observation 04f60f49-951a-40a4-a26d-2e8c88f2a8bf · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786344Z digest=sha256:45804bb6e4d43d0a86e2a0f1b9b3117f2f27c69b5e4e6a5cd80f173df3b984bf

Observation bed9b6f2-dc40-4894-92bb-2aa4d92a1ff7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.932459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.932459Z digest=sha256:69992024e3c3850f7bb4f969e05449c4ead4796448655c6a858026ec2ed666dd

Observation 6a23d7f7-75e2-45b2-9981-415acf380d00 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.064295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.064295Z digest=sha256:625198affcc25455e243a0559fcedefb3dcda07699c6c48dd055781f3fb37187

Observation 9b3bc751-1820-4855-b643-3b526a4061b9 · outbound

This paper cites Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.973919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:40.195808Z digest=sha256:10e1c09976ff9bd948b414f2f238d4eaae1bbac5500e6c4328307057adcca269

Observation aa115202-2d34-45d5-9d2a-faa4234d69e4 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368474Z digest=sha256:efa049332460a8573d583bae92447acc1b610aeb6b782b65b5b867943123db53

Observation d557360f-bcb7-4c09-9807-c01bc5effd5e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient Streaming Language Models with Attention Sinks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.524261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.524261Z digest=sha256:e77a9fea5769592c6c5496be80c8023423c53048b176193df524b64eef926e90

Observation 23115589-32d9-4e42-b217-8997271304a3 · outbound

This paper cites Qwen2.5-1M Technical Report.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Qwen2.5-1M Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663165Z digest=sha256:2a4f1f54d11ea080852afb3ed705317db14ac16bdf091b6c26c82b4963a04246

Observation 5437d5b9-1181-4359-adaf-152295e111b3 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Orca: A distributed serving system for {Transformer-Based} generative models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.778601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.778601Z digest=sha256:49dc071486ca9bb93c7be2b085beeb15de4ef8fc7ac56766699b531b3d21fe4b

Observation 0a5774c3-cbdc-4a02-aab5-6c6eaea9326c · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.900312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.900312Z digest=sha256:77989ef179409dc788cdc97ca15d87b14b58e5b6d142cc731c2c771ee5362fdf

Observation 0057c0b2-26c2-4e71-9083-60f8c3ffe1db · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:41.662361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:41.059939Z digest=sha256:bc9d600e5f49f4cd4cbb4f555ccfb5a7ca50d36b6a89f860245962efbb0b76df

Pith citing papers

Observation adc68189-911a-4ea2-bb71-b3f964fbeaf9 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.988048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:85d0e1211f256d108dc1d82e2ec8633af33f9389f0deb83a8a7e45063a4172dd

Observation b692aa6d-3478-4390-a79c-b57fa3cd51c3 · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:20.890041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:736cd51dc462224309436a329f09a5011f4341801902e2b836ef6be9647d329e