Pith. sign in

Paper Citation Record · LEDGER

SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2307.02628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02628 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.503529Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5df63f8b-9d33-44e0-94ca-b8818c6cd662 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.379404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:223b22e503423c4f196c5c8211b3b39380e15debc25961ae3cc917dc6782ecb6

Observation a11ad5dc-7073-471b-a5e2-d670ee99f0ba · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:49.952631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:08b0c71bd65f38447b3fe3281030db2409cad0854fa3231d3e9b47b98f5df613

Observation e26b19c1-73ca-4a02-901d-1d9b158496c8 · inbound

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference cites this paper.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.434433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.434433Z digest=sha256:7a7eaf2834e062e8da449606f82c154725af391b35f34b3292bad9ac4d3c8368

Observation 9ab3af2b-2f8d-4b4c-bcae-ff369ea093b4 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.308313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.308313Z digest=sha256:73dd83e4ab5be75564ff77911488004084dfa68638b8176c398e45da58966cf2

Observation 76a8c7e2-caa3-4204-8c28-77bb60f2d59f · inbound

DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies cites this paper.

DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:51.501090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:51.501090Z digest=sha256:c52a0dadb1f17e31d315b4b5130e3831ae67c8289a9bbe854c53078c8f5028b7

Observation ac8afd31-eb27-4983-b3d3-4b83630c0fbf · inbound

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts cites this paper.

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:37.871305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:37.871305Z digest=sha256:e8330a22710b6676a304c539c6a2347bcb8716e600a2eccbd1050a4082778d9b

Observation a28ec175-a345-4258-aa1f-c00f2cc48c57 · inbound

CLaSp: In-Context Layer Skip for Self-Speculative Decoding cites this paper.

CLaSp: In-Context Layer Skip for Self-Speculative Decoding SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:10.411437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:10.411437Z digest=sha256:2abed0e56b695ea66fffdaf2b65b0b48632c186a0a081335acaa36940884d8a7

Observation 8387e030-7012-406e-b94c-0a712bcafefc · inbound

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism cites this paper.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.187286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.187286Z digest=sha256:a657cfbbd45ec995d036f1fa0e635772d98b1effa228c2f2a97f87afba85230a

Observation 1e81d5a2-1bf9-4671-b9ad-37fb719fdde5 · inbound

SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling cites this paper.

SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:51.465417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:51:51.465417Z digest=sha256:d15ccf64d1b55997815cb63d1f86a107cf94595cba4b4f4841c7ecd802e0de1f

Observation f4905ae9-ab05-421f-bcd0-23469d277d8a · inbound

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts cites this paper.

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:14.591057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:14.591057Z digest=sha256:225fd6cda57757ac5636c1f11fd5e11d58f3a152996ad50dd8cf05cd0946d425

Observation 8fe2c9d0-f948-465b-942a-3381ee8bc030 · inbound

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference cites this paper.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.884228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.884228Z digest=sha256:713fd2aa3d9da4d7056427fdb1979558dff90868683e8ecd309a694dba5313d7

Observation 67315da6-4d1b-44f4-ab48-8590e629b0a2 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.599010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.599010Z digest=sha256:a1c5ed3f953d0a50b765f767d13626cc4dbf44832adee33794d36226ccb25507

Observation fa2da869-017f-4f5f-8bc1-2d4616a095ca · inbound

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping cites this paper.

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:35:17.569538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:34:14.371914Z digest=sha256:bab4920e5a7555f03c55d1fc8eb3dda65a1c2418d95af12f2a51fb727e7b94f4

Observation f9d686ef-61f8-41bd-9c6d-53bf6b1067af · inbound

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs cites this paper.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.784748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.784748Z digest=sha256:655e5e6ab4abd64c52640b0085eaddc4920d2fe32f6fe4acd1bf46126ee4875c

Observation 1dfb2969-6b3f-41d5-bc01-e6d110b2d784 · inbound

Networking-Aware Energy Efficiency in Agentic AI Inference: A Survey cites this paper.

Networking-Aware Energy Efficiency in Agentic AI Inference: A Survey SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.530651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:39:02.634334Z digest=sha256:bbe0b64ebe27bf92b18fefd7a18f55aa90007d522b80403e08082b8b442f702c

Observation 346aa319-838f-44e8-b307-fa6bad1c7577 · inbound

When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth cites this paper.

When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.637794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T09:28:29.238615Z digest=sha256:c2037c2b358a691dd267ebef30485de4518ac32164a45e090fca8863c8871130

Observation c95ffadd-1837-459b-af2e-d5db23303a91 · inbound

Depth Adaptive Efficient Visual Autoregressive Modeling cites this paper.

Depth Adaptive Efficient Visual Autoregressive Modeling SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.627925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:22:55.035749Z digest=sha256:36cd604334a4eae3752e9c6ab223b4343141551dd7579d24750a340dc44276c2

Observation 7267a922-ce82-484f-b1ea-449b60807c5a · inbound

River-LLM: Large Language Model Seamless Exit Based on KV Share cites this paper.

River-LLM: Large Language Model Seamless Exit Based on KV Share SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:10:23.273492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:36:46.851450Z digest=sha256:891651512e009f9a11f961347242608d918c7fada61be6241e3a0a77b0eb4b81

Observation de247a6c-11ba-4492-ac32-750217c5860b · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.393173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:4fce618a9162a542c8c15965c5d8dbc7ddb3e93c9ad41aa45627073741642b36

Observation 765d1458-1b04-4ec3-bec6-1fb1f8c05cd9 · inbound

N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation cites this paper.

N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.243965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:46:16.391521Z digest=sha256:67d31efb88fe2df60b77df02f89d5593fa9f8f2619824caef85b4ac948a8931a

Observation 28dd328d-0752-4925-a4a8-742b323b8eca · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.168571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.168571Z digest=sha256:caa047922a6013fb8cfc7ae92fb3c040ed2130ff5ecc52ea26baca38a1cce1bb

Observation bf4ba506-0967-4a2d-ac28-8682fd074c09 · inbound

Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers cites this paper.

Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:11:14.454774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:11:14.454774Z digest=sha256:800c929a45d9af0fe8b0198c02ca42626843af5659cba75ccffb6a3f7e9a7f04

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:a88da83973ac35fea411453a9635eb6c751c2d90767c92ac6a419b740a1d53bc