Pith. sign in

Paper Citation Record · LEDGER

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2506.03700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03700 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:02:39.419910Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:36:59.761975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:44:27.295523Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bbd273e-ae5b-44cd-b242-8d013d4217c0 · outbound

This paper cites write newline.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.111163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.111163Z digest=sha256:79d3e904c83b4248626dc938e77f2671dad449ea8371c1f270940df3d38e4eb3

Observation e876ddf6-366d-4099-a62e-a32c355c1ba1 · outbound

This paper cites GPT-4 Technical Report.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.117325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.117325Z digest=sha256:fcf2aa7eacdc0e396974ab3fba83094eedc93f045e30f5c8fb2f56db24efd1e1

Observation fa0354e5-809f-4963-96b7-b4e489d0721f · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Hydra: Sequentially-dependent draft heads for medusa decoding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.542572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.122576Z digest=sha256:c76c1a57892dd32ba0f7b47ac98ecb1c1a62242237ad1fee156dc71c7f50d175

Observation cc87aa4a-3646-4859-b130-69d357c602d1 · outbound

This paper cites Anthropic: Introducing claude 3.5 sonnet, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Anthropic: Introducing claude 3.5 sonnet, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.526458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.127205Z digest=sha256:59a5a87aa250acb8bf92d53551a79e67c599ed742cb7093222ec16c5ab43a7b3

Observation 19c0759c-be96-4050-8989-a97eb1fd32cd · outbound

This paper cites Program Synthesis with Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.131952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.131952Z digest=sha256:a3e174a179cf7f16db7be098fe85417de676b92a801972e5c41e4864e9b9c174

Observation 19d12a47-2d5d-44d4-8aea-3aee6179fd5d · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.136964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.136964Z digest=sha256:47436600447a4e6ac19be01a33d81a0e9508ef855836731db8125f1683c9a215

Observation b8914505-6fde-4fe7-8f1d-0292bec2367d · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.141519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.141519Z digest=sha256:0d60071861b62d0ca28515aaa3d4a0cf35f7ea491a3b30347b0663ea5ad3c83c

Observation dad9f962-f222-425e-8389-f2f82e441797 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.146799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.146799Z digest=sha256:17da2c6f4d9fe3a98599d61ff6e3ef53da586fe83bb6b47ec5cdfe6e376bfd4b

Observation a6d50787-59f9-4d9b-9fde-cbfb998908a9 · outbound

This paper cites D., Chen, D., and Dao, T.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism D., Chen, D., and Dao, T

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.499957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.151371Z digest=sha256:2e436af64010841407569b94509c495785711e0f3c9f3cd4379c6da080f802f8

Observation d6f76329-635d-4bdd-8141-deb9541e6a9e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating Large Language Model Decoding with Speculative Sampling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.155905Z digest=sha256:ac0e9e01aafdc455a99aab9d3a5cd7146f0ddcad26d685173141809b8b6c7e36

Observation 517a3181-d6a9-43b6-b6a7-e03263455086 · outbound

This paper cites WAPITI: A Watermark for Finetuned Open-Source LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism WAPITI: A Watermark for Finetuned Open-Source LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.160378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.160378Z digest=sha256:7bc674282e50ea95c0491a2621b386dd79cd4ba2f95f996cfbdc8840649420e3

Observation 2995c851-694e-414b-8043-b2cf4dc76ec0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.164777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.164777Z digest=sha256:3f08e27fa7db0d8e572e8ef34291229cc4bb734fd64f341f9bc11c49ff67860e

Observation 06865bfe-f5cb-4986-bc51-0c98ec6ac986 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.169280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.169280Z digest=sha256:190ddb36f9c76e5c493ee774c6ed84ea765cb2e3d76c67634488d661ecfd7389

Observation b5ab56a2-a86e-48d8-aeee-ffb48ab8ecca · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.173568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.173568Z digest=sha256:e87c782f66e04509dd9d122aacd9a11c06a5c6c97329d1b1de8dfbe1c44838c1

Observation c5a98939-fdc4-4245-b86f-46aa7c54c832 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.177936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.177936Z digest=sha256:f16b80168cf530b582d40fbb4792a7336d177dfffbc8bd115e1ddf7483d93cb9

Observation 18e49189-a87c-4c21-9a4e-6006e3a73a7f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.182650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.182650Z digest=sha256:1ff8488e6418526c2ed4f870190c230d7e0ebb769dbcfd906ce37b992b25b315

Observation 8387e030-7012-406e-b94c-0a712bcafefc · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.187286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.187286Z digest=sha256:c3360d3fcc38d1253427bde66f5e0aafa6b1a1863a92e5b587ec41fde54974b7

Observation 38732595-4509-47ae-922b-34fc6e2a3a39 · outbound

This paper cites QLoRA : Efficient finetuning of quantized LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism QLoRA : Efficient finetuning of quantized LLMs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.474491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.193356Z digest=sha256:734488b331e2e03c4d67793170f00ea18f56af8c92eb19cb46fb01062791c8a4

Observation 7b6071e4-831a-43f8-8200-da6fc141262c · outbound

This paper cites Jump to Conclusions: Short-Cutting Transformers With Linear Transformations.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Jump to Conclusions: Short-Cutting Transformers With Linear Transformations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.197592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.197592Z digest=sha256:3ac4bc30a837a89542ab54942cbeab8ddbe88a3b3ddddaa472c1d3a7d627824a

Observation 8ab443d8-4c37-4ed2-881b-a094f71fe195 · outbound

This paper cites Glide with a cape: A low-hassle method to accelerate speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Glide with a cape: A low-hassle method to accelerate speculative decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.459979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.202238Z digest=sha256:2462e7982f450856ee7d7ab1181d56e520453fcd8762737c5a278dbe44985309

Observation 15e11d1e-e8f8-4514-815f-f1ae43ebe166 · outbound

This paper cites The Llama 3 Herd of Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.206389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.206389Z digest=sha256:59c69cb71f282bdf37c4060f194afdc1950dcf08e9c61b304a791217020016a1

Observation e86cd2b5-5a9a-440d-be4c-034879542cfd · outbound

This paper cites Depth-adaptive transformer.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Depth-adaptive transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.445245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.210467Z digest=sha256:0864bf0e70224a329ed7d45d40db643670e5fbf5b8b6c31f12e411c442c69a86

Observation a80e998c-4443-4ea8-adbe-f0323b900831 · outbound

This paper cites L ayer S kip: Enabling early exit inference and self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L ayer S kip: Enabling early exit inference and self-speculative decoding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.430525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.214544Z digest=sha256:fd250bb0932560a5d38e696221ddc2277d857f34ee136bc01993ec0d5a9dd84c

Observation 201dc01e-7564-4e6e-a3fa-9b6a8a11b585 · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Break the sequential dependency of LLM inference using lookahead decoding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.415153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.218593Z digest=sha256:02fdec9ccf2852d02613b1f9f153d31462310d4c02aa0a3bf7386e5fd7369fe2

Observation bb14de1f-0ab4-4afb-8f6a-818d935ffdca · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.399611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.222519Z digest=sha256:79e05425a6051a6bcfde2e2acacfb84faa09ee13163a17970ccbadeb7ca93570

Observation 94f56415-7df1-433c-b465-91a8350e4ce4 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.226807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.226807Z digest=sha256:783a78a0f2f3dac2eb961460cc70ac5685307fd1c199fb49935fe7809eafe637

Observation 5bed347e-7dc9-447c-a58b-b38a6354fe65 · outbound

This paper cites REST : Retrieval-based speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism REST : Retrieval-based speculative decoding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.383163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.231129Z digest=sha256:59dec24eefa184cec78b3796111b75396acf8e92a3dd6c938eba5944f45a349d

Observation a12cfdc8-b27b-4192-a2d0-6e48d50e6eb5 · outbound

This paper cites Training Compute-Optimal Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Compute-Optimal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.235113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.235113Z digest=sha256:4bc98a348625bc4795dfb87ec2257ea8f1fee4831d34227c34430d913d21150b

Observation daff1876-b4f7-4876-b4fd-120090b0c010 · outbound

This paper cites The curious case of neural text degeneration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The curious case of neural text degeneration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.366794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.239313Z digest=sha256:db296909468fc9d5426018ad4fe0659b80e71c72706e40b1cdd81c1ebcf56f87

Observation 5b54c427-bdf2-4722-a5b0-6f787aff1049 · outbound

This paper cites SPEED: Speculative Pipelined Execution for Efficient Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SPEED: Speculative Pipelined Execution for Efficient Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.243484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.243484Z digest=sha256:cffa90c3e012ecc29b69a59ffe22af79720d014e9e648a260d1b44594bb02a87

Observation 8031e3f1-13b2-4eed-8dad-9b94ecfa7908 · outbound

This paper cites E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.351586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.247841Z digest=sha256:51284a0575a94cf485a1e44bab3fc55125da55522c114363f3aca291ac9c8acf

Observation 92faf0a8-b416-4acc-8d08-63636b1f2689 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Efficient Test-Time Scaling via Self-Calibration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.251956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.251956Z digest=sha256:d4cd6a7b6fdc1763c5a3097c123132a00b44f7f0bd168183b870de8f1e65cd25

Observation 0b37d04a-8585-4b7f-9845-86a22812cabb · outbound

This paper cites Multi-scale dense networks for resource efficient image classification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Multi-scale dense networks for resource efficient image classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.336787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.256248Z digest=sha256:252d7b7e629bb6f17df1d2f6247166187eee2cca32def249536a4a52d8e2c69e

Observation c6ba6a95-54c5-4eb3-b122-6a9b5d7193e7 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.260580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.260580Z digest=sha256:798bed297c07789723a4b4f87361d3906fc9c2ae31b0d9cb303d41ebee663ae1

Observation 757b7c6c-b126-4ce0-84d0-6d6d96164eeb · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.321949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.264904Z digest=sha256:ad138b2e516343638562bc8fdb544eb377036f70f4c8428d800daeb2982d908c

Observation 411e6a88-6391-4fd9-b649-592e0f0e39da · outbound

This paper cites Mixtral of Experts.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixtral of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.268868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.268868Z digest=sha256:d856aa9152286de863d6a993f668a4325ed2bc1972e139a90f45ccfc88510ee4

Observation 8ef25ad9-1f64-4dba-a51a-44173e48128e · outbound

This paper cites Sigsoftmax: Reanalysis of the softmax bottleneck.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sigsoftmax: Reanalysis of the softmax bottleneck

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.307036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.273111Z digest=sha256:81f3aa4f1cce7ab5bb28c83d92a7863cb1b30813923f6c7b22395f033c299ea0

Observation 1b63bef3-2853-4021-818a-d4ccc8d5b4e0 · outbound

This paper cites Scaling Laws for Neural Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling Laws for Neural Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.277280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.277280Z digest=sha256:274575fe5ad88942fdac1c5fbe4619b81489e300f718e2eb5f2df71230cdb7b4

Observation 471210c7-f1bc-4815-8654-554b09d948a9 · outbound

This paper cites A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.281273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.281273Z digest=sha256:4a109e9d4f20478bb528735e3922a809504ecfb18f98e89e7fa39e6c03e298da

Observation 2beb81a9-c3c8-4764-a4ce-583a06234c48 · outbound

This paper cites W., Gholami, A., and Keutzer, K.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism W., Gholami, A., and Keutzer, K

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.285634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.285634Z digest=sha256:e2ccb3e98ca66c1e4024c8018c0143359232f8f282e016bdc9acd55b20af4847

Observation bb7f8bf5-bb5c-4c31-981c-8a06f41d05cf · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.289715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.289715Z digest=sha256:a47282bfac902aa11dda6f0baca8208b64f795cd9672cf2ec9162f74b015f947

Observation 3ef55932-148c-41f3-9398-670a7691f026 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Adam: A Method for Stochastic Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.294461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.294461Z digest=sha256:3693f494e8533a4c1727cfb2123d408b3667384e6a557dbfe33d5693fd345142

Observation 9866707e-911f-403d-8b81-2558603b726e · outbound

This paper cites Fast inference from transformers via speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.298736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.298736Z digest=sha256:9cf6a3a6b312a18f5331f25980531c4ba410aaa1df83b3b4b63cda44eb0df779

Observation 94acec46-634f-44f3-b2e3-ff75c53c7aaa · outbound

This paper cites EAGLE : Speculative sampling requires rethinking feature uncertainty.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE : Speculative sampling requires rethinking feature uncertainty

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.268093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.302878Z digest=sha256:5518103cd8f96c9899bfc17b0a8cd08bec40c681664269472545fc4ddf3107df

Observation 8c182b05-b6c8-49b4-9489-3bab0d4c6b74 · outbound

This paper cites EAGLE -2: Faster inference of language models with dynamic draft trees.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE -2: Faster inference of language models with dynamic draft trees

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.306771Z digest=sha256:67e1689dc23748ea3b26e29873c4491ba26ecac11dddbccb795b9bd2b6dd7a40

Observation 4a3730a1-1a9a-448b-b099-e8041ad0a0d7 · outbound

This paper cites Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.310757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.310757Z digest=sha256:4464264c07f5507a6e819a1c43e11d75612566e8775b8d43013f0eea8b50ad44

Observation c3e6d8d3-1059-4a40-b3e0-6c887e8ac97c · outbound

This paper cites Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.236346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.315020Z digest=sha256:df7dbcc75ebcb280150a9371593afee9605c9bf35d6b0696912802d4b77604d4

Observation 392f706e-b5d3-4f78-8388-c4441d6cd354 · outbound

This paper cites Online Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Online Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.319649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.319649Z digest=sha256:d3b48acffd63591a9b5836c7ee6a54a5bddd1dd63915980075441438fe3a7a41

Observation d65e517c-e34b-44ef-8000-15ee09bcc21e · outbound

This paper cites SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.324065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.324065Z digest=sha256:a9a33ddca9c4c7b75497e2c611099a14ff1c63d677ffb76750a637206e6c4e42

Observation 13b92ac1-0853-49ca-b29a-7b46567ace67 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.328762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.328762Z digest=sha256:78592317ffe47938ff2107aa7490253c674729e46f792f36d478ef3b05d03596

Observation 524c374e-19b6-41cb-94cf-db6ad6372692 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.332990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.332990Z digest=sha256:5f468bfdc11472677f77fe6dca7606a56d6599119731e26256c1ac572e043489

Observation fbeac526-b188-40c9-92ca-a53d9fd3eff2 · outbound

This paper cites B., and Lapata, M.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism B., and Lapata, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.220011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.337310Z digest=sha256:0530db2507f863d9c39490968f10cd4a5c06862a9cb55fe342bef93d02d4999b

Observation 6dbb81cc-5103-41ad-b311-0f9705e9b60f · outbound

This paper cites Introducing OpenAI o1: Learning to reason with large language models, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Introducing OpenAI o1: Learning to reason with large language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.205186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.341535Z digest=sha256:3ddc5b25abb270b6812d95aeb319d413f1c2412fe28dd5e9f4abfce7738807f8

Observation b5b52e0b-3b03-412f-8a18-be2363e6336d · outbound

This paper cites Suri: Multi-constraint Instruction Following for Long-form Text Generation.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Suri: Multi-constraint Instruction Following for Long-form Text Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.345559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.345559Z digest=sha256:1c487c854e859866ae6d29e661510feffe7d492a2842a302d8488ea9ff50f0d8

Observation 137b5722-76e6-45f9-93f9-7dd36980ffda · outbound

This paper cites Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.349604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.349604Z digest=sha256:399054d08ab9cdc13fc8076565d8a0b589e4bcbd95143ffcde3ec16d36510476

Observation c7d2b5b6-ba6f-4ae0-bd91-d6335a15b7d7 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Zero: Memory optimizations toward training trillion parameter models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.354357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.354357Z digest=sha256:56cb00b5a04c59ff0d6a30493951ac29d723f50a88b7bbdda99a5465ea7880fc

Observation 00f1598f-a45e-45c9-8378-1ca46fde6337 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.358502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.358502Z digest=sha256:552905e7e82b424cf52758c1857e12d384822f797a35276c7212b46c03d5afe9

Observation c5b092e6-bafb-4015-b99a-ff921311b22b · outbound

This paper cites Confident adaptive language modeling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Confident adaptive language modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.363381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.363381Z digest=sha256:9752c6483ff21393641cab761e41939ec7af772ee3c88abb55d6c456aa6d13a2

Observation 1a47a6e7-9004-4fb9-af90-21d5e8a57ec0 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.367516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.367516Z digest=sha256:c3d62f7198e4dbcdddc7e1f08caea8944a923bacd29e09df3f135b8b7f1d7287

Observation 80b9a603-75a1-4c5a-9dc0-d52cb414d087 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Blockwise parallel decoding for deep autoregressive models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.371960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.371960Z digest=sha256:bed92e245a58fa5b0fb6fac42dd98b5bd0e72a1be87f9a96ac48408e8c658e04

Observation f90a9b21-851c-4a9f-957b-736288df6d89 · outbound

This paper cites Branchynet: Fast inference via early exiting from deep neural networks.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Branchynet: Fast inference via early exiting from deep neural networks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.376015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.376015Z digest=sha256:d270d0a54cad210265caaf3eeba7b00f795869e6c2787a1e95f2813de7d5b2e4

Observation 91e438b8-eb7b-48b9-9186-9aea743e3c0b · outbound

This paper cites Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.380095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.380095Z digest=sha256:35b44a8e6bce3ad7cd774de708ee9b93d717db0d9a72ed2faf11792cab253bba

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:697bdab0fb9fac1f30fd2c29707a8e32789a7a5088e4112783b64f4e6cf9d29b

Observation 6a266fe9-7a37-46dd-ac7d-0d2d051c1307 · outbound

This paper cites SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.146961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.389151Z digest=sha256:187e43f58af145378e4e0dcaa3ccf5494d8e9eab775b1b0a43945d673896efe1

Observation 3c447bb5-cafd-47be-b11f-5fc3ac132133 · outbound

This paper cites Sheared LLaMA : Accelerating language model pre-training via structured pruning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sheared LLaMA : Accelerating language model pre-training via structured pruning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.131589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.393390Z digest=sha256:8d9a1330f774a8092fd49d9f166fa46d5b6f6a5a4956a06ce6b6eb0eef400991

Observation 1b1fd645-bab9-4770-a385-c6b917c22b62 · outbound

This paper cites Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.114487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.397839Z digest=sha256:00a77e0dbf3856e519c575720aad65cd9cf55875ce7d61c2b1c1b6caea9c4b28

Observation 21a3dae6-973a-4857-968f-03fbbd74d4bf · outbound

This paper cites an unresolved cited work.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:02:40.097652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.402276Z digest=sha256:3d03994c7d508a0554bfc7ef4fe7d399346c9327fb9edca25c4e7d41c0e14915

Observation 99ec4d90-e5a2-4bc7-9b00-edef53252766 · outbound

This paper cites C o S afe: Evaluating large language model safety in multi-turn dialogue coreference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism C o S afe: Evaluating large language model safety in multi-turn dialogue coreference

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.082567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.406543Z digest=sha256:e7da2a29c2fbaf18dad398d40cee4707e8c682fa7cacb0b6ad43ae08f84a3844

Observation 1ab76d82-5c3e-4c69-b777-6e309c525fb3 · outbound

This paper cites L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.066657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.410964Z digest=sha256:893a046678f5504a189ca0a913bfcc917555204cf906402cfa7996583e1235dd

Observation b43b1424-d65b-4f94-b3e4-c5ea393af016 · outbound

This paper cites Draft & verify: Lossless large language model acceleration via self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Draft & verify: Lossless large language model acceleration via self-speculative decoding

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.051401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.415285Z digest=sha256:49bd33e867b1d498e65a0399bad2724e6e8c7d5854fcb4603df7ae005f25c96e

Observation 98a08fef-5fd2-4b5e-a4e8-127f583d7fcc · outbound

This paper cites SafetyBench : Evaluating the safety of large language models with multiple choice questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SafetyBench : Evaluating the safety of large language models with multiple choice questions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.035083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.419910Z digest=sha256:d37e54d5039a9e54ac5454741a9991c8b519dda1af3fd7003d02b58bc1377643

Pith citing papers

Observation f6daaf60-f3ba-4e44-b348-d33c26c2b93a · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.182708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.182708Z digest=sha256:42f41b0b305ded27a1a7ba9e6903de1009449a08567a09e9705f1194168adb72

Observation 475e5a31-0525-44da-9544-e3508eba031d · inbound

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration cites this paper.

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:58.697801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:22:05.049712Z digest=sha256:e381a4170861a4bf45264832cd6d1d507e9bdd49a256db9f60f8d414b64ee664

Observation 642cdc18-96fd-4579-97ce-916b42d78cd2 · inbound

Depth Exploration for LLM Decoding cites this paper.

Depth Exploration for LLM Decoding AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:44:27.297075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T08:43:47.472682Z digest=sha256:04eaed610170c763fb27bffccaccd73def429594716e5e6b989aa333ea7bd79e

Observation ca429a9b-ae99-4be2-b3ba-6fe0e5517d08 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.552840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.552840Z digest=sha256:b27a660962baea2b4a0ca37f7bbb0a87cfcf4a0759df2e25eacbedda70e92b0e

Observation e8654665-5847-4dfb-a35b-5caf3e3d6f74 · inbound

LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization cites this paper.

LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:59.761975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:59.761975Z digest=sha256:03185c127c650ca3658770426989035e00211e43ba6dbed1c3b25a2423772bc2