Pith. sign in

Paper Citation Record · LEDGER

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

As of 8 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 4 inbound Pith citation observations for arXiv:2506.03700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03700 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:02:39.419910Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:16:08.552840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:44:27.295523Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bbd273e-ae5b-44cd-b242-8d013d4217c0 · outbound

This paper cites write newline.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.111163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.111163Z digest=sha256:1d521af6637ee3ee440d839e53926827eb8a04cd6b14a7fa82e41dc569b64e68

Observation e876ddf6-366d-4099-a62e-a32c355c1ba1 · outbound

This paper cites GPT-4 Technical Report.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.117325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.117325Z digest=sha256:c56e83bd3482773022186f9e1afd8d4bb6abed9902b59d48e7df07868fec404c

Observation fa0354e5-809f-4963-96b7-b4e489d0721f · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Hydra: Sequentially-dependent draft heads for medusa decoding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.542572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.122576Z digest=sha256:4a3a5554ebef96c3f18bac91e2b304fd684427cecc82596014b372b4fceb4340

Observation cc87aa4a-3646-4859-b130-69d357c602d1 · outbound

This paper cites Anthropic: Introducing claude 3.5 sonnet, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Anthropic: Introducing claude 3.5 sonnet, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.526458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.127205Z digest=sha256:e5e68383e188268800a46c2cf073aa2a1e47af08d6ead3a20d64033853914314

Observation 19c0759c-be96-4050-8989-a97eb1fd32cd · outbound

This paper cites Program Synthesis with Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.131952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.131952Z digest=sha256:bbdf46fecd01552d4b04eaea857de1b8b33ff0ac98fd21a288745bcb28e9349c

Observation 19d12a47-2d5d-44d4-8aea-3aee6179fd5d · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.136964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.136964Z digest=sha256:4dd4a1652b5deb8cc0f0fce7b224ae48088659f5cd1d3f99ed918801ebd5aba7

Observation b8914505-6fde-4fe7-8f1d-0292bec2367d · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.141519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.141519Z digest=sha256:cc953e1de07150088b2e27b0c170a679056060ef13a3a75d8f851db043022ab9

Observation dad9f962-f222-425e-8389-f2f82e441797 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.146799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.146799Z digest=sha256:25566460f720ac023f2a105b570ff0f2c185dabc37b84484c93bd483effba887

Observation a6d50787-59f9-4d9b-9fde-cbfb998908a9 · outbound

This paper cites D., Chen, D., and Dao, T.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism D., Chen, D., and Dao, T

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.499957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.151371Z digest=sha256:8f410c1c31cce4a7d703768c74cd88e2023266523a61d3117da527c000c31d94

Observation d6f76329-635d-4bdd-8141-deb9541e6a9e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating Large Language Model Decoding with Speculative Sampling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.155905Z digest=sha256:1f150f482e6e32438d4b4b5b15a55a71bccf8803b18e2941b5653f153fbf5b4d

Observation 517a3181-d6a9-43b6-b6a7-e03263455086 · outbound

This paper cites WAPITI: A Watermark for Finetuned Open-Source LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism WAPITI: A Watermark for Finetuned Open-Source LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.160378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.160378Z digest=sha256:6bc861f6c9aa8fbd82ee0bc1c3159bc99ae335e1c03a0793cf02fef9eb4996d2

Observation 2995c851-694e-414b-8043-b2cf4dc76ec0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.164777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.164777Z digest=sha256:f4b47847e3f87cfc34563b77ec70d8f4de64729e880f8fee700d02a86f8b233e

Observation 06865bfe-f5cb-4986-bc51-0c98ec6ac986 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.169280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.169280Z digest=sha256:a04cee9ff20adad5f563800ff57fb38606e318e7eab651bfede807e0e91cf927

Observation b5ab56a2-a86e-48d8-aeee-ffb48ab8ecca · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.173568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.173568Z digest=sha256:7f9edb57f7ca4cb877b133adf2f742843c9d219edff9cbb6f5028755b0ddeda2

Observation c5a98939-fdc4-4245-b86f-46aa7c54c832 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.177936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.177936Z digest=sha256:6cc98b3609270f53780be5a520c31d59eba5fbcc3757c7bcdc7a52a15dc7ee48

Observation 18e49189-a87c-4c21-9a4e-6006e3a73a7f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.182650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.182650Z digest=sha256:2970446775230a07423ca570f3c5e7df69756b7a6cd315b537d1338c70e84fe8

Observation 8387e030-7012-406e-b94c-0a712bcafefc · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.187286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.187286Z digest=sha256:c662fcfb761827c0da181af40c8a3847aab9a8c5099cf3f93ba02d555ba6d9d9

Observation 38732595-4509-47ae-922b-34fc6e2a3a39 · outbound

This paper cites QLoRA : Efficient finetuning of quantized LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism QLoRA : Efficient finetuning of quantized LLMs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.474491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.193356Z digest=sha256:2430ea72fe468614f45288a28d69f7e8672376fa540ca5977e610d30e9a50efe

Observation 7b6071e4-831a-43f8-8200-da6fc141262c · outbound

This paper cites Jump to Conclusions: Short-Cutting Transformers With Linear Transformations.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Jump to Conclusions: Short-Cutting Transformers With Linear Transformations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.197592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.197592Z digest=sha256:bea848803ad80aba26ce478f65d0763a77b718bc2dc31bd9d5f4499faf980466

Observation 8ab443d8-4c37-4ed2-881b-a094f71fe195 · outbound

This paper cites Glide with a cape: A low-hassle method to accelerate speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Glide with a cape: A low-hassle method to accelerate speculative decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.459979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.202238Z digest=sha256:039cbbc2389139ee95f1f342ee6912561aa73a0465d0cef5490568153be21b00

Observation 15e11d1e-e8f8-4514-815f-f1ae43ebe166 · outbound

This paper cites The Llama 3 Herd of Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.206389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.206389Z digest=sha256:118233489b280817d44037e09286ade6893dc85b1c3dd3ff10cbe510ffbf6e4e

Observation e86cd2b5-5a9a-440d-be4c-034879542cfd · outbound

This paper cites Depth-adaptive transformer.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Depth-adaptive transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.445245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.210467Z digest=sha256:2a1d9aceb858b1cc745d062a28592b3c84099854ddcaf0d076879e1a14a8a8fb

Observation a80e998c-4443-4ea8-adbe-f0323b900831 · outbound

This paper cites L ayer S kip: Enabling early exit inference and self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L ayer S kip: Enabling early exit inference and self-speculative decoding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.430525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.214544Z digest=sha256:e6d889b7a2b470937759ff4c25c15f2e16beee03ad7f0565f13ae2a5f10b86dd

Observation 201dc01e-7564-4e6e-a3fa-9b6a8a11b585 · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Break the sequential dependency of LLM inference using lookahead decoding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.415153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.218593Z digest=sha256:e2f4a7081016c65dd4bc36c0158f68e164704533c88005607cd7417f228f6282

Observation bb14de1f-0ab4-4afb-8f6a-818d935ffdca · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.399611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.222519Z digest=sha256:9d9d0d9b663b34fe57c9fbb6b5f0b851a1d5b707074716705eb2b3de5480f2e6

Observation 94f56415-7df1-433c-b465-91a8350e4ce4 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.226807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.226807Z digest=sha256:bd48cb8d2ad9242aff2c456b13e00ab5ab5ab4f85efb122b37b8ab6dcdd2c6fd

Observation 5bed347e-7dc9-447c-a58b-b38a6354fe65 · outbound

This paper cites REST : Retrieval-based speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism REST : Retrieval-based speculative decoding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.383163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.231129Z digest=sha256:ec173e3290c259af32218cd3c279090f286aef0c76cd12a645bad7f73865e9f6

Observation a12cfdc8-b27b-4192-a2d0-6e48d50e6eb5 · outbound

This paper cites Training Compute-Optimal Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Compute-Optimal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.235113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.235113Z digest=sha256:cc9c7b274de763a813eee831eb22ac057831d0c90c5f27df92250dc059f7c315

Observation daff1876-b4f7-4876-b4fd-120090b0c010 · outbound

This paper cites The curious case of neural text degeneration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The curious case of neural text degeneration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.366794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.239313Z digest=sha256:5dc0fc018ed3c294f1949eaa853083e1631ec06e014c59befb1a8ec8c5f2505b

Observation 5b54c427-bdf2-4722-a5b0-6f787aff1049 · outbound

This paper cites SPEED: Speculative Pipelined Execution for Efficient Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SPEED: Speculative Pipelined Execution for Efficient Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.243484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.243484Z digest=sha256:e3fb88ec43c5c8414992d85a8ecd6e82c092f5c71a76c2999859b3415d3b4dbb

Observation 8031e3f1-13b2-4eed-8dad-9b94ecfa7908 · outbound

This paper cites E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.351586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.247841Z digest=sha256:c070e7aa82739c6e18e80fe836ba7191449b147cbe476e42b6744b9ceb6c0dd1

Observation 92faf0a8-b416-4acc-8d08-63636b1f2689 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Efficient Test-Time Scaling via Self-Calibration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.251956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.251956Z digest=sha256:0b218c6c61a582604623c8af8ec36cbb7837aed1a8daa9626d342ff5cd4a92ad

Observation 0b37d04a-8585-4b7f-9845-86a22812cabb · outbound

This paper cites Multi-scale dense networks for resource efficient image classification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Multi-scale dense networks for resource efficient image classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.336787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.256248Z digest=sha256:620e84d8ef865b80e2dcaff8861d9f88e91dd26dad53b519311e296e8b2f7222

Observation c6ba6a95-54c5-4eb3-b122-6a9b5d7193e7 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.260580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.260580Z digest=sha256:8b0fb7e492d21be0a1998b5f57fa35acf6b84fe1593c8763c6654f8fbbe0c3ba

Observation 757b7c6c-b126-4ce0-84d0-6d6d96164eeb · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.321949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.264904Z digest=sha256:76beef4468e1c087577ed4daa3e257c3675cd95c6dcc63b6d67647b9adc57810

Observation 411e6a88-6391-4fd9-b649-592e0f0e39da · outbound

This paper cites Mixtral of Experts.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixtral of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.268868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.268868Z digest=sha256:dde34381384f349b7676083c990183ebdf93cab389145930376fb2a7b65f9492

Observation 8ef25ad9-1f64-4dba-a51a-44173e48128e · outbound

This paper cites Sigsoftmax: Reanalysis of the softmax bottleneck.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sigsoftmax: Reanalysis of the softmax bottleneck

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.307036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.273111Z digest=sha256:0d61cbd56c959e267309221205a048a5a411595cd0dfe8a1dc75156e20783db9

Observation 1b63bef3-2853-4021-818a-d4ccc8d5b4e0 · outbound

This paper cites Scaling Laws for Neural Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling Laws for Neural Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.277280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.277280Z digest=sha256:7dc1ae4fdfeba88f6f24ee3e790851e1a800000a10397d0cbf195cf072ff4cb3

Observation 471210c7-f1bc-4815-8654-554b09d948a9 · outbound

This paper cites A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.281273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.281273Z digest=sha256:78b68e78dc74fc6dadd741471f9a16a7f8520666aa56b7bd0589fccb34c35dc6

Observation 2beb81a9-c3c8-4764-a4ce-583a06234c48 · outbound

This paper cites W., Gholami, A., and Keutzer, K.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism W., Gholami, A., and Keutzer, K

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.285634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.285634Z digest=sha256:9cbec45656a0cb24439245516663905a702ae2d0ba52ca2327e026f2554a88db

Observation bb7f8bf5-bb5c-4c31-981c-8a06f41d05cf · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.289715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.289715Z digest=sha256:8aefb7bb64d3179e82b88e63166e1cbc3b45165e519ff7ecf5fd8690ad276d73

Observation 3ef55932-148c-41f3-9398-670a7691f026 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Adam: A Method for Stochastic Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.294461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.294461Z digest=sha256:de1ab04750903d4fd7de5b6e909a6abc415110b0889409f85d4b27f6a46bf669

Observation 9866707e-911f-403d-8b81-2558603b726e · outbound

This paper cites Fast inference from transformers via speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.298736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.298736Z digest=sha256:2af1b7f9e7ea0248f703e4152deb3c45cc78cf232f0005054b1859e8b3b05b01

Observation 94acec46-634f-44f3-b2e3-ff75c53c7aaa · outbound

This paper cites EAGLE : Speculative sampling requires rethinking feature uncertainty.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE : Speculative sampling requires rethinking feature uncertainty

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.268093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.302878Z digest=sha256:e49d462c673e51b69b79913e9531d01c73db445784d70d4385a00411bd648b30

Observation 8c182b05-b6c8-49b4-9489-3bab0d4c6b74 · outbound

This paper cites EAGLE -2: Faster inference of language models with dynamic draft trees.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE -2: Faster inference of language models with dynamic draft trees

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.306771Z digest=sha256:147c3ff6ebd431dfa54bb6e414894ab5be55c26777f8eda19ee348a203f1d373

Observation 4a3730a1-1a9a-448b-b099-e8041ad0a0d7 · outbound

This paper cites Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.310757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.310757Z digest=sha256:9f850e53da63002853c15ec8a88886cf574ff58e139b326ea7d609d15023f047

Observation c3e6d8d3-1059-4a40-b3e0-6c887e8ac97c · outbound

This paper cites Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.236346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.315020Z digest=sha256:c7b1ef7677dd6b23b1ad4b0f8837dc65e1b4cd2198e1034c89cef398f8a91d05

Observation 392f706e-b5d3-4f78-8388-c4441d6cd354 · outbound

This paper cites Online Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Online Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.319649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.319649Z digest=sha256:e1471cfcc1e00bd8d144b44db372a8552ec789afa713ed08d047efd05d160b14

Observation d65e517c-e34b-44ef-8000-15ee09bcc21e · outbound

This paper cites SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.324065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.324065Z digest=sha256:c4214709e24e39c880c94f428dc2be09205f0cc5f40d1f1f7d34457bd597f185

Observation 13b92ac1-0853-49ca-b29a-7b46567ace67 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.328762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.328762Z digest=sha256:10c08de3bea6bcbc48d50dfb63f0fa4f27a14229c9b6da589c2f650a6dd37da3

Observation 524c374e-19b6-41cb-94cf-db6ad6372692 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.332990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.332990Z digest=sha256:34fa384bfc5f4f2fa7a74b32cb7e7c74e823f00069e0941ad16365a360e358fa

Observation fbeac526-b188-40c9-92ca-a53d9fd3eff2 · outbound

This paper cites B., and Lapata, M.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism B., and Lapata, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.220011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.337310Z digest=sha256:42a0d9177f4cb120594aface4f78960bce545618c95763d4bb6f5ed881eb838e

Observation 6dbb81cc-5103-41ad-b311-0f9705e9b60f · outbound

This paper cites Introducing OpenAI o1: Learning to reason with large language models, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Introducing OpenAI o1: Learning to reason with large language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.205186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.341535Z digest=sha256:01895414ffbebc48009d6f075b1ad863c46312bb85a5bc5dc6e21a9f53372d95

Observation b5b52e0b-3b03-412f-8a18-be2363e6336d · outbound

This paper cites Suri: Multi-constraint Instruction Following for Long-form Text Generation.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Suri: Multi-constraint Instruction Following for Long-form Text Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.345559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.345559Z digest=sha256:68c4a699a91dccf21ebbf4d8c171d271da7a88d2a18bfc224597323e63d8fe22

Observation 137b5722-76e6-45f9-93f9-7dd36980ffda · outbound

This paper cites Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.349604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.349604Z digest=sha256:dc4e660bc7a02d9ce6585a4485102b6611ca52bf3631dd0a6e8ff85932e8ad6d

Observation c7d2b5b6-ba6f-4ae0-bd91-d6335a15b7d7 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Zero: Memory optimizations toward training trillion parameter models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.354357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.354357Z digest=sha256:46700b84d612896cbc97163e4f107bf481193595992d9f7b988b329c10d809b4

Observation 00f1598f-a45e-45c9-8378-1ca46fde6337 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.358502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.358502Z digest=sha256:b655052595b2ad19002fe98f9ffcacabe345f8dce27596c996675ae7e58f0d1f

Observation c5b092e6-bafb-4015-b99a-ff921311b22b · outbound

This paper cites Confident adaptive language modeling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Confident adaptive language modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.363381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.363381Z digest=sha256:3720894c9dd2ba2659dbb248ffe60a01e34d71f7021c36813bad0323adef8899

Observation 1a47a6e7-9004-4fb9-af90-21d5e8a57ec0 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.367516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.367516Z digest=sha256:4290cd5a7ea8df3677df0a672000d02548dd45ffdfc870ea6d5088669120f1c0

Observation 80b9a603-75a1-4c5a-9dc0-d52cb414d087 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Blockwise parallel decoding for deep autoregressive models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.371960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.371960Z digest=sha256:1c21b8f2149f5f8a07ea6d4b16cfb76be5fc943c0f9b3b1250d8b032f0f52441

Observation f90a9b21-851c-4a9f-957b-736288df6d89 · outbound

This paper cites Branchynet: Fast inference via early exiting from deep neural networks.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Branchynet: Fast inference via early exiting from deep neural networks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.376015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.376015Z digest=sha256:7b4f24b77e73012e7393ef9dd5f2fe6bf2b35b4d9b7d632dc1984c4ff0ddbeee

Observation 91e438b8-eb7b-48b9-9186-9aea743e3c0b · outbound

This paper cites Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.380095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.380095Z digest=sha256:28b1cdec642788b0ce442c30e4eb204203ed9b3d4c7bd8c5eb5ab4a86191d213

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:1b2da0d983384704651370fa56abf9a8e8c9e15569d0d4d14a6d6b92fb4b7c74

Observation 6a266fe9-7a37-46dd-ac7d-0d2d051c1307 · outbound

This paper cites SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.146961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.389151Z digest=sha256:e0aad37851195b69197d7210d696e7acf12c5fd765a10230cf84453bbd3d9cca

Observation 3c447bb5-cafd-47be-b11f-5fc3ac132133 · outbound

This paper cites Sheared LLaMA : Accelerating language model pre-training via structured pruning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sheared LLaMA : Accelerating language model pre-training via structured pruning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.131589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.393390Z digest=sha256:a96a05ae767b96f19862565b091ae1a3287cbdd9c701bf912f2e1fe664eeb44c

Observation 1b1fd645-bab9-4770-a385-c6b917c22b62 · outbound

This paper cites Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.114487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.397839Z digest=sha256:d682ae54c99acaf9c72da42be035ff5692a7b89972f800bcbd5061d5ba27e8da

Observation 21a3dae6-973a-4857-968f-03fbbd74d4bf · outbound

This paper cites an unresolved cited work.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:02:40.097652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.402276Z digest=sha256:f3ec9be84839308ec3b2610a5f6c61f4e1b0a5a5e8075a58d0d76fedc8dfb344

Observation 99ec4d90-e5a2-4bc7-9b00-edef53252766 · outbound

This paper cites C o S afe: Evaluating large language model safety in multi-turn dialogue coreference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism C o S afe: Evaluating large language model safety in multi-turn dialogue coreference

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.082567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.406543Z digest=sha256:985d2f8951826af98084dfccfc353f57b6547074c9c354271e16e8a0537d237c

Observation 1ab76d82-5c3e-4c69-b777-6e309c525fb3 · outbound

This paper cites L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.066657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.410964Z digest=sha256:099f7ba67d057e8c69674ca8557d17b0faf9ae8d795621e26e65bb2cbe3ead85

Observation b43b1424-d65b-4f94-b3e4-c5ea393af016 · outbound

This paper cites Draft & verify: Lossless large language model acceleration via self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Draft & verify: Lossless large language model acceleration via self-speculative decoding

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.051401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.415285Z digest=sha256:e99badaf3d76ef2025bcef2816a1862d21b246f2ef3f0477a9d19eb689172b15

Observation 98a08fef-5fd2-4b5e-a4e8-127f583d7fcc · outbound

This paper cites SafetyBench : Evaluating the safety of large language models with multiple choice questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SafetyBench : Evaluating the safety of large language models with multiple choice questions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.035083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.419910Z digest=sha256:8032b2de0da8fc1fa646d3e9564122780633652733a8ea5b80ad3df17090e774

Pith citing papers

Observation f6daaf60-f3ba-4e44-b348-d33c26c2b93a · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.182708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.182708Z digest=sha256:48e51f30bb23e86b0f821e81cfaa7864c2fd24c18ca06f3a4706b37c5466920f

Observation 475e5a31-0525-44da-9544-e3508eba031d · inbound

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration cites this paper.

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:58.697801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:22:05.049712Z digest=sha256:380455f75c05bb3445f5af4ea44a4036d75c66fd39d3cce28118218094c83fb5

Observation 642cdc18-96fd-4579-97ce-916b42d78cd2 · inbound

Depth Exploration for LLM Decoding cites this paper.

Depth Exploration for LLM Decoding AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:44:27.297075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T08:43:47.472682Z digest=sha256:f1f66bd4878f6b3633d3e35cf4cdf5deea0fde1364f9a3b8bee88fe2bf9a0b46

Observation ca429a9b-ae99-4be2-b3ba-6fe0e5517d08 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.552840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.552840Z digest=sha256:af2a0ecba21160c2641182b9e00123e9b0c86f5716a5754cc15d91ed67549e60