Pith. sign in

Paper Citation Record · LEDGER

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree

As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2412.12639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12639 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:56:39.808556Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:08:14.247342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:26:24.739905Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 088995a7-78ae-4a33-96d8-086b3ba695a1 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.544245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.544245Z digest=sha256:0b716da74485977e41f3ebfae690478f8e62b4a2378db46686273e71f3975dd8

Observation 1095de3a-02e5-48b5-8e39-cdeb02d13c4d · outbound

This paper cites write newline.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.549623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.549623Z digest=sha256:096496627d9c2ab12450ba3c264d7f45c6f4672fe997adeb7325b267e5c647c4

Observation cd21b300-3a5a-4f28-bccd-710ef8584d9c · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.560090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.555810Z digest=sha256:3ee0dcad046ca6e4532bc483936857426b205cd82bf52f3ac75def88cf3ca83d

Observation 1f4542ba-b36b-4454-b281-de72137be92f · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.561574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.561574Z digest=sha256:a646b1c1fa780890b61563815452068ce57d08543e59dc5a1801a966dc35585c

Observation c8fa522f-0984-42ad-99c2-8f573c815683 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating Large Language Model Decoding with Speculative Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.568070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.568070Z digest=sha256:ed04084302ff0afb00a48a84606cd1bc55bd9bbfeb8f75485cea7a6023eab840

Observation b102a94c-ac3e-486c-b195-872c3f3f1a09 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.573168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.573168Z digest=sha256:3f4f8522ef6648b2100d577703f4b8e6481c777605847b5a0511e4e884553885

Observation 85d3b22d-8c13-4fad-9bc1-0a8201207dca · outbound

This paper cites Cascade Speculative Drafting for Even Faster LLM Inference.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Cascade Speculative Drafting for Even Faster LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.580987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.580987Z digest=sha256:bfce334893f142ac52e4ea5a763fa336ed7b283139cce0d0388ad9038372a44c

Observation 49babc3b-b03a-4f88-9090-b3237ccf8188 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.586758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.586758Z digest=sha256:a1d747bbd35c1ce7728d01f5b6f1a66dff8b295d64202353cfe9d8cc78df91a4

Observation 330b6d1d-7df0-4e21-8124-61eeafdf9cd6 · outbound

This paper cites F.; Tao, D.; and Tu, Z.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree F.; Tao, D.; and Tu, Z

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:56:40.547022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.592117Z digest=sha256:2cf914606ef2b4d06687fa88e07a0f6d0f0f19b39f79a65375ba9b7aed427c03

Observation dfc0bfd4-2231-4a0d-b1f1-898a3cd10303 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.596093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.596093Z digest=sha256:aa88482f1730bea6e8801a29af5e3c4de6dd007541d337f972cdd51c9c716ed1

Observation 2ada1f50-1235-4c8d-a9bc-f04c24a8fe48 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.534753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.600341Z digest=sha256:c1936b58168cadb8603aa75381d525bedb5808b35bffc24b039d137b893240e2

Observation 6f84257a-97a4-414e-9fef-66e472a28cbf · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.519615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.603874Z digest=sha256:7a2a8e49936a6bbff5261ac1b263becc2508039815b96bc3445ab58f3d49b484

Observation 175d18a3-778f-45f0-a3f4-316ff51a0d06 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Better & Faster Large Language Models via Multi-token Prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.607525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.607525Z digest=sha256:806695969ae47d07595f750adcbb92325624ea191dc45107e7adeb3115f2c85a

Observation af66018e-a323-43ff-a041-e8aa6ef34276 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.612401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.612401Z digest=sha256:aa8f5ef44d4ca492df67fb766051098aa636726104ba2b7388748cbb6d5a89dc

Observation 8ab821b5-56ec-4d84-af51-00a24bbfdb33 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.493299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.616664Z digest=sha256:8e8d81659ec5ce26db93aa5a3efcc0f39a5dfcd1d7c9e688460c9b38e66095ab

Observation 8fa911fd-5676-4554-b832-7d0412bbc4be · outbound

This paper cites Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:40.085206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.621807Z digest=sha256:a140047fa5268ee7777fe0633e2ef94f55893165b0b27fac9efff061f6201402

Observation a288e5bd-b530-420d-939f-c51e55f907c5 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.477145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.627068Z digest=sha256:69da80fca0af1bd77602a93e3c62ec8ae46a50a435f474253fa1ead2efc33ba6

Observation 4482b961-7191-436d-a1f9-4a51fd5f0c59 · outbound

This paper cites Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:40.061510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.631862Z digest=sha256:55dc74f38cdfe98773402155d4b90f3e4635ec9a2069ea9a95d91f29ceaf8d39

Observation e903e635-f4ba-40a0-a7b8-19bfd2aff34f · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.638681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.638681Z digest=sha256:d7f884db53fe8f73f771bf40ac67b93914dee6e34b448247a4842f91d260bbea

Observation 47ab6896-f1db-4b66-b3b6-ac9762015293 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.642790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.642790Z digest=sha256:2d196af1b2ad0195ca417434fac7664831a7e7e24ac576f5625c79cae46890d8

Observation 176e4f2f-8909-4647-a87e-36107980e473 · outbound

This paper cites W.; Gholami, A.; and Keutzer, K.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree W.; Gholami, A.; and Keutzer, K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:56:40.446329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.647897Z digest=sha256:445c70790bd77584d87cbb3528dc92f543cd6b71d01a5d64234a326b01fb1cb3

Observation ba1f2e9a-4969-45f3-9bfd-3d67295ac2cc · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fast Inference from Transformers via Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.653053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.653053Z digest=sha256:360a6f9cb0cd2c3b678571f9e7317e43ca0198f0d8d3d63ce8b6faf874ff50d6

Observation 018023f6-a511-40a7-abf7-378e706924c6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.663532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.663532Z digest=sha256:c7933a9e4bcf6c31d36c12503afde720db2efdf308b296aab168fc08c4fbcb91

Observation fd32a44b-b5b5-45b3-80be-770b28e89738 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.426314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.670302Z digest=sha256:50e09ba17e1f21824ab36cf39c2d0da99f60f6a83cc4735d37d4f321c73ebd58

Observation 9c9e0826-8656-4623-89a3-6b8b1145ee52 · outbound

This paper cites PaSS: Parallel Speculative Sampling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree PaSS: Parallel Speculative Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.674571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.674571Z digest=sha256:59b65e18b2aba6d63d9a4fc3bd59c0d7a25d3485d759f574a433b407df99ad1f

Observation dc02a5c5-cb1a-44f5-b382-3ce5f730c8e1 · outbound

This paper cites In-context Learning and Induction Heads.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree In-context Learning and Induction Heads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.679327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.679327Z digest=sha256:2def9e3f027883575d1f5ceb0a2ba767f8bd0d0ed2548ceaa55bd34de4610521

Observation 85628b68-1705-4171-bacb-ad0d145d294f · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.408816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.688977Z digest=sha256:473868451d7458cc920f4727660b80b8ec045f08b6a0788eee8a0c28d5e786c0

Observation 9885abe7-7969-4c1a-8386-d0338e15da0a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.388929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.693837Z digest=sha256:360d0ab27cc8ed41644025bf0cd384464efeb74dc5a29450498e9163e8c5e9dc

Observation d0236f5d-c9bd-4bdf-8d7c-04b16e9ac3d0 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.369184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.699417Z digest=sha256:47b5e234178135a7b975ff317567ac9d31bc170a86742252461336f13289c1d0

Observation 7ff372d0-6ce5-423d-9cf6-8092dfad27ca · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.348487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.707155Z digest=sha256:69796fc5961bdb5f02b81dfe4789d928416740ff0c34178ae040d5bc18bda2e9

Observation dd547bb5-cfd0-4fcf-9315-2da5ce487123 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.329362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.713796Z digest=sha256:942c022fbe00ef3047a72a76e6ad4e49a92e0886323c82eb486668311bb1eaa7

Observation 975475f5-45a8-4cae-b9be-ff813764f66e · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating LLM Inference with Staged Speculative Decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.717832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.717832Z digest=sha256:c17c7f5b27a4de56e7499107893c3ff5c6af3af5a259a2071d9b2f0db94949cb

Observation 209fa9cc-c15f-45fa-a7bc-36baa9148d3a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.313913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.723725Z digest=sha256:2c3d69369a2a6715927acc79323d80b4022305469711faf46308e416f566fe57

Observation 62d5b023-e51c-4dfe-a105-c319e0e4f94e · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.298005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.729611Z digest=sha256:515cf63fd8a43fe7aa977176a746019520ee30781c255f8abc2c95c6e3bb36c4

Observation 15dee101-6225-40c4-b916-5ceafb872b5e · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.734235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.734235Z digest=sha256:fd387d5de74022c2fc014e2641223e3586f974e44f3827544f02f6136d4d82d4

Observation 329f5adc-5855-4742-a60b-bdb427d69611 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.280552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.739625Z digest=sha256:3c8c8e0afb18eac1d4dc9ff2f58587c8cec2cc9cbb97a73b6a23438690469c19

Observation 162cc3db-c0d0-45ed-b514-07c4426e776d · outbound

This paper cites Efficient Large Language Models: A Survey.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Efficient Large Language Models: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.744899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.744899Z digest=sha256:c4f6b71800afe4e6fd44b3a683a34db404cec554655c3eab4c487b20ff0c01ce

Observation 9f89cce7-c135-44d4-aa75-2e5694ae6d62 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.749541Z digest=sha256:153b7f2c6c7ebf73664d27aa9c62ce30e43a931f16eea1c7d74615a1d5c8efa9

Observation 560487e9-aabd-4d9c-9d73-6e1b71cf50c8 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.242349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.756167Z digest=sha256:5c105fa7b1942dd34d9011ff49b7f418e917c6b5678cf5a5653f37be1bbf6fc8

Observation 399bc475-9fcd-49fc-b503-f18391603f7e · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.760297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.760297Z digest=sha256:40e8b8c8a52a78fd888b76594c6a6ddc3580361a7461e51629e22f8393b50e2c

Observation 9e19fcaf-326a-4171-89ea-ff63a7fc60b1 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.217318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.765552Z digest=sha256:5fe31d38e840051098df6644ed3eb864bdc1975e059fe77e86570ba245f8702e

Observation 6a36a5c1-318b-4044-b772-3322b34f97d5 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.772067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.772067Z digest=sha256:6b465005151cf0df30b2f41f425259ee02c2a1590fac52420d2b236e84361c3b

Observation 355893e3-66d0-4fdb-9c4b-b7356dafd000 · outbound

This paper cites A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.777354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.777354Z digest=sha256:7841a7ae0d9367d7cdcd329ac8aa0ab42c8af262d1c056f5ef8ffb8dea75359e

Observation eecc80ce-032a-4d7f-a8c0-6b2e2fff821b · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.202894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.782758Z digest=sha256:5c60b3eed8c52c75b32ef0d6d38c3bc7be1b3100dbe6d41485f15e98583cc960

Observation 368bf258-9b1c-4e01-a035-843ea2b7ba09 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.788351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.788351Z digest=sha256:6acb36c138b12586a3ee67c1b566e5753fc052dac4f3bdb6081eef12bb5ab7d2

Observation 871eb455-5dcf-4a0b-8ac6-99956511008e · outbound

This paper cites Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.793794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.793794Z digest=sha256:d51f890be58ad6568fc9e5600511c25b9621d89e7a559fa0a27a60f8effd129f

Observation 1320dda8-e617-4b36-adfa-b0a7a0f8b86a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.798834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.798834Z digest=sha256:c2ae0488ed023c23cf43c5be5d4459b93582e63cda4894ee346cbce18463443a

Observation b884096a-6e81-40bd-ac00-1248f41f1541 · outbound

This paper cites A Survey on Model Compression for Large Language Models.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree A Survey on Model Compression for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.803614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.803614Z digest=sha256:57c1598dd295625d414f228be9be4b3c57ccfc27e4088873a09b5aa257578236

Observation 6908f2f6-4ebb-4d34-8076-e355979e08aa · outbound

This paper cites Fast Decoding in Sequence Models using Discrete Latent Variables.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fast Decoding in Sequence Models using Discrete Latent Variables

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.808556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.808556Z digest=sha256:36adcbe27621844d8b152929d58e478afaa8df12c9cd6ca9c2121f3c3395971a

Pith citing papers

Observation 4e8c7bb5-17af-47b6-a400-6099c7817c98 · inbound

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding cites this paper.

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.741992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:08:14.247342Z digest=sha256:19749afca2c47a370cc5a64f55c8e42e71ec48d67853e0854c340c99c4e3f00f