Pith. sign in

Paper Citation Record · LEDGER

Speculative Decoding Reimagined for Multimodal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.14260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14260 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:56.294364Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:44:37.884895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:11:02.504431Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dd385fe-9b0d-445d-b8f1-cc68bc70f85e · outbound

This paper cites Gpt-4 technical report, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Gpt-4 technical report, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.607827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.002187Z digest=sha256:d519d76ce55b69ad7368a2a316f13665f74c710aaf2996baf4c03cf284d8c0b3

Observation 7f67eed2-59f5-48e5-991c-cd46b061895c · outbound

This paper cites Sharegpt.

Speculative Decoding Reimagined for Multimodal Large Language Models Sharegpt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.438602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.078322Z digest=sha256:f591c1e952cef055ecb53df8e8e63f7d0dedcc1234b28422f869ccc8a017e723

Observation e6927f43-8e81-4274-8bf4-258444046e7f · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.149539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.179436Z digest=sha256:764fe8777c5e79c64c7023acba07c70683f754643a4b3e83f3a49ca6b34d9be5

Observation 965a96a4-3d00-4de6-ab69-719b76398cba · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Speculative Decoding Reimagined for Multimodal Large Language Models Qwen2.5-vl technical report, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.929426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.248582Z digest=sha256:fb671c4824028f220ca4b921b0eef39a648e02cbad4df80a6409764833317723

Observation 8a5d4887-01d3-4a29-a57d-195b85dab7ac · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Speculative Decoding Reimagined for Multimodal Large Language Models Lee, Deming Chen, and Tri Dao

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.679671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.378487Z digest=sha256:5a665f9bf9b0c57ea35ec5ccb7ed048ebce28834e503572d433132e306b63c05

Observation eb9916d6-7160-437e-ab3b-039bc1341d38 · outbound

This paper cites Madtp: Multimodal alignment-guided dynamic token pruning for accelerating vision-language transformer.

Speculative Decoding Reimagined for Multimodal Large Language Models Madtp: Multimodal alignment-guided dynamic token pruning for accelerating vision-language transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.486357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.477020Z digest=sha256:2729dc2155ca6257cf69948068d3d4dfa949bb6137b693c6a96194c53b676e96

Observation 2f28a232-067d-40cb-b824-62a0df9bdbc0 · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:51.603894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:51.603894Z digest=sha256:75894e43d22634f26621accfb6d7b2ed90dcd40b6e078a9654177eb0bfa5cefe

Observation d2cb13a1-3cfb-4125-9508-c2cb4afd6868 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Speculative Decoding Reimagined for Multimodal Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.308827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.758861Z digest=sha256:bbde1961d8fc44fcaf3a9bc7fd819453df5401d8b6d3328be4487f36d376eb93

Observation 8472b815-8778-4719-b53a-e657fa0b0a88 · outbound

This paper cites Diffrate : Differentiable compression rate for efficient vision transformers.

Speculative Decoding Reimagined for Multimodal Large Language Models Diffrate : Differentiable compression rate for efficient vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:04.090731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:51.856526Z digest=sha256:76cb5175275847a249915dff08937e05ef8fef6335c3c3c0bdaa420c018222e5

Observation 5971ae9b-2c20-4fff-85ce-4694e2f964cf · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Speculative Decoding Reimagined for Multimodal Large Language Models Evaluating Large Language Models Trained on Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:51.998002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:51.998002Z digest=sha256:45770f847de09493f6e0cb787d1e51f53e5ec1cc208d21392a3d53ee4b232080

Observation cc6f647b-2343-4021-b403-87dbb174dd76 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Speculative Decoding Reimagined for Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.140080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.140080Z digest=sha256:dae90cd758fb4acd51fada669062310c5e7de78bee7d56d90ce3c75ad641b136

Observation a1a6d5c4-a8c2-4b6b-a810-159e414f5e6d · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Speculative Decoding Reimagined for Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.269586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.269586Z digest=sha256:c02524827511213f5976fd023ecc8a7e18594ca907950143489061cd41cd6b47

Observation 35311bcf-9417-48ea-8de5-2a63cfc3d0a1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Speculative Decoding Reimagined for Multimodal Large Language Models Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.384747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.384747Z digest=sha256:bf8bb719c54b9edfc0006a70cbe5328cdb40575e47b6ec893d3b79035bbaed0d

Observation 607f0ff6-625e-4e48-b6a0-4c393f4a9d1d · outbound

This paper cites Flashattention-2: Faster attention with better paral- lelism and work partitioning.

Speculative Decoding Reimagined for Multimodal Large Language Models Flashattention-2: Faster attention with better paral- lelism and work partitioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.902246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:52.520459Z digest=sha256:0ad8d13e718b18c5a30ea14b28f73dbd331505a36e42ead6d0a87e165b58eb4b

Observation 2cca68c7-34f7-4400-a9b6-cb1ffa537e3d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Speculative Decoding Reimagined for Multimodal Large Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.688863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:52.654998Z digest=sha256:5b2aaeee7cdab13a256aea4f47fcf030b7184c93f0dcff52f991c09b06dc8e2b

Observation 46f9dfcc-1f41-458f-8b67-45ad1b2ab7f0 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.454182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:52.766877Z digest=sha256:757e2aef97d8518f71bcb31805d02031f3da873aff64de3d37bea7ae09c22ccd

Observation 40c1840c-6422-42a9-9344-4dacfb33782b · outbound

This paper cites Break the Sequential Dependency of LLM Inference Using Lookahead Decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.964039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.964039Z digest=sha256:ceccf4ed65f9ef3c4224f9a00e48e33b093de78f885a3851cf1ed046b711c34b

Observation db0aa4eb-b18e-4dcd-9db6-1983100197a2 · outbound

This paper cites On speculative de- coding for multimodal large language models, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models On speculative de- coding for multimodal large language models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.284745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.096692Z digest=sha256:d2786d4a43dc75ac8d63c9b210c828d83bdb62b2aeceaf44c161e6cc167df627

Observation e8600f12-c8e8-4e6d-bf18-324907bf2c49 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Speculative Decoding Reimagined for Multimodal Large Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:03.107501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.184615Z digest=sha256:70bb620e2b92cee50ce1e2f4c108accf06903671b08de4e734e16a87f2b89877

Observation e755fa42-dd18-4317-a2b3-9d9c1240456e · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Speculative Decoding Reimagined for Multimodal Large Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.878228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.273506Z digest=sha256:2b0cfeb38cd4593e149d5c91b4b989bc80a647ef1da5724a3b4452b6cb3bdd24

Observation 00150653-8d2e-4abe-93b2-8d56c5ad849c · outbound

This paper cites A diagram is worth a dozen images.

Speculative Decoding Reimagined for Multimodal Large Language Models A diagram is worth a dozen images

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.615171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.386578Z digest=sha256:ddf9cdfafa5f89378a7e5d55cc355f626cd1b6875a5b7244b959554e3576ce47

Observation e581eb29-484a-4097-97be-101792a256b0 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Speculative Decoding Reimagined for Multimodal Large Language Models Efficient memory management for large language model serving with pagedattention

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.306653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.502385Z digest=sha256:b4411b4cb6430d776d33aba97133c39500d0eaa4215eb8961ea7b17654e5b764

Observation df04763c-b5f0-4827-aee5-b89b7354b870 · outbound

This paper cites Fast in- ference from transformers via speculative decoding, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Fast in- ference from transformers via speculative decoding, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:02.079645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.602320Z digest=sha256:f7511adf7518b16a30f5929d1d1ef1fda6d4216c9be483309db7fde79d95be11

Observation a7c93abb-242d-4d10-bde5-3d0554f5ef1d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Speculative Decoding Reimagined for Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.815404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.666871Z digest=sha256:464c74f04ef1422ca8fca6393c4001415d9fe9ef64fa8b8d9a0757570e331479

Observation 21f14983-b513-4013-aeb5-2932a222da6a · outbound

This paper cites Mbq: Modality-balanced quantization for large vision-language models, 2025.

Speculative Decoding Reimagined for Multimodal Large Language Models Mbq: Modality-balanced quantization for large vision-language models, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.617127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.769305Z digest=sha256:b271add93182a0550335a05886f6ff42fe62ce36a5f9756624a1c322de09499d

Observation 5240c826-d93f-4cdd-824b-a258b39e2219 · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Tokenpacker: Efficient visual projector for multimodal llm, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.352493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:53.976045Z digest=sha256:c4c577949795db933914fbd7b0c5174f5de36c269776631ff4c7c0d099ff2525

Observation 67046b93-e169-4e1e-96d0-0e3727a4473c · outbound

This paper cites EAGLE-2: Faster inference of language models with dy- namic draft trees.

Speculative Decoding Reimagined for Multimodal Large Language Models EAGLE-2: Faster inference of language models with dy- namic draft trees

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:01.132549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.072661Z digest=sha256:68572648fa1267e76383407e6762f29cb54f897c5c5a5d2e167694cde4a178df

Observation 44eecb0f-c787-40ae-a6bb-4606fbfc0afe · outbound

This paper cites EAGLE: Speculative sampling requires rethinking feature uncertainty.

Speculative Decoding Reimagined for Multimodal Large Language Models EAGLE: Speculative sampling requires rethinking feature uncertainty

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.916324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.172704Z digest=sha256:e2915e04a9e75ff2d8c93f5f57435ff42495d7d51148a566f1f83b358d578598

Observation e58f0841-c3dd-4168-914c-81caba70ba05 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Speculative Decoding Reimagined for Multimodal Large Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:54.275560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:54.275560Z digest=sha256:d3f42d37897b3fd8ecbbdbf60984fad30236522ecbaae7ebb69b5c77dee69f3a

Observation fb07f3c9-5560-4d84-ae63-8a5668afcedc · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference.

Speculative Decoding Reimagined for Multimodal Large Language Models Boosting multimodal large language models with visual to- kens withdrawal for rapid inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.756646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.379762Z digest=sha256:91ddfd36b157f2fa673f6ea8d24781fe5e78e5c3392ab2073c4953097d00369e

Observation 7329108f-1ae6-449f-965e-d2f725a69bb8 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Improved baselines with visual instruction tuning, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.538357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.468070Z digest=sha256:de527da1cd21e3a8ff6e5c64bd37a53930842fe52a281f7e37b841b7a82bcb7c

Observation 35094642-eb29-49a9-8ad5-976e6191ee16 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.419555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.618384Z digest=sha256:163ef0b474932b3f49fe773b9b724837d4382fd944deaed702a46d72b9da7078

Observation 6220cd67-de78-4ac0-81a0-8e3d457447ef · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player?, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:00.118135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.717220Z digest=sha256:809f9a6fa928001d4c694269d4a21ff252c3f0b949ebd88a915fb6687b372b97

Observation 986fedbd-0afd-4107-8494-be3dbdc3aad6 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Speculative Decoding Reimagined for Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.857099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.806172Z digest=sha256:3d800eee57436924572e599f5e29222be534a57eff12f7931456c5cb3e699b11

Observation 207bbc7e-f974-421d-8697-829ca8da4c0a · outbound

This paper cites ChartQA: A benchmark for question an- swering about charts with visual and logical reasoning.

Speculative Decoding Reimagined for Multimodal Large Language Models ChartQA: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.680608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:54.900842Z digest=sha256:fb89ac3c03cf2c374ea2b263be2aa0cb3b920cdd4c7d5601e84cf25f048ec966

Observation 53475137-9217-4efa-b2a9-87ad353e4469 · outbound

This paper cites Mm1: Methods, analysis and in- sights from multimodal llm pre-training.

Speculative Decoding Reimagined for Multimodal Large Language Models Mm1: Methods, analysis and in- sights from multimodal llm pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.426790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.018726Z digest=sha256:9bc4641c16a033c730def14b368625412f94246c0bf8bc16cfdc94852a91ddb4

Observation e5507623-85a9-4fa1-848a-c3647f6cdb67 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Speculative Decoding Reimagined for Multimodal Large Language Models Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:55.084527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:55.084527Z digest=sha256:e4e19a850a96dc5d79ec30fd0bedca7a43052f563ea67dfb2f66ef8265bde6f5

Observation b7bc08e0-ec7a-4491-a807-8c5a8b5bdafb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Speculative Decoding Reimagined for Multimodal Large Language Models Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:59.052147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.175992Z digest=sha256:05a0ec828ef54d1c91254d0d74380c0f17f5bee8922f0086ec8fde553df951c7

Observation 19239ca9-9e8a-4aa9-9123-15b289b32672 · outbound

This paper cites Crossget: cross-guided ensem- ble of tokens for accelerating vision-language transformers.

Speculative Decoding Reimagined for Multimodal Large Language Models Crossget: cross-guided ensem- ble of tokens for accelerating vision-language transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.783792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.378336Z digest=sha256:a556aa07102bd8c6dadd59bafc6770c72d630cd2c3842d6381b1ed8e82bff739

Observation f628e393-4091-47fc-9b2e-6495cd4a7187 · outbound

This paper cites Towards vqa models that can read.

Speculative Decoding Reimagined for Multimodal Large Language Models Towards vqa models that can read

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.567691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.483714Z digest=sha256:4c6c554211d64601b8dde32a790623a6b7078da7e6245ce6146b029a369fa4b4

Observation 174f5246-7735-4a63-9e22-aa1561093d9a · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Gemini: A family of highly capable multimodal models, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:58.189988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.579960Z digest=sha256:a4f2e4b75edb35d4df9981ac940e6bc9c8d381f9dafac91ea893977b7e1bf942

Observation f5c3e04f-90a1-43b1-a083-a396502bbb87 · outbound

This paper cites Q-VLM: Post-training Quantization for Large Vision-Language Models.

Speculative Decoding Reimagined for Multimodal Large Language Models Q-VLM: Post-training Quantization for Large Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:55.672801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:55.672801Z digest=sha256:c73d956b5f0300afbbff9da750da0f659fa0a693c89b703c62f7cf5089fe87e9

Observation 625280b5-3ea4-4d1a-a6bc-a95f5059fee9 · outbound

This paper cites Large multimodal model compression via iterative efficient prun- ing and distillation.

Speculative Decoding Reimagined for Multimodal Large Language Models Large multimodal model compression via iterative efficient prun- ing and distillation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.935490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.730004Z digest=sha256:be77aa1992e335e2f1d17f2cd2d9f460996b276656b8af46e054356f039ba87c

Observation 92345a1f-c526-408b-bc67-86134a9b672a · outbound

This paper cites Ppt: Token pruning and pooling for efficient vision transformers, 2023.

Speculative Decoding Reimagined for Multimodal Large Language Models Ppt: Token pruning and pooling for efficient vision transformers, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.686430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.802858Z digest=sha256:76c9f18007f4c952798f34b93218d2b03c7992c23daf5e299b32fb636a6c8c53

Observation 3af64a77-fdc8-47e3-9376-1d68bc731d99 · outbound

This paper cites mplug-owl: Modularization empowers large language mod- els with multimodality, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models mplug-owl: Modularization empowers large language mod- els with multimodality, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.404553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.863908Z digest=sha256:838f36649b12918709e6cba65a6abdf70e1e629ef3ecf8c52756b105bfb48510

Observation be0b5308-a3d3-45b4-bce2-be9e6adb24a8 · outbound

This paper cites Generation meets verification: Accel- erating large language model inference with smart parallel auto-correct decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Generation meets verification: Accel- erating large language model inference with smart parallel auto-correct decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:57.136593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:55.955642Z digest=sha256:4e0d2f5fd669e48c432a88a1ba5d504bbd14b6911b3e7660fd9de9c70c6db9bd

Observation a7e61b52-b387-400e-abea-ad29daf30a8e · outbound

This paper cites Draft& verify: Lossless large language model acceleration via self-speculative decoding.

Speculative Decoding Reimagined for Multimodal Large Language Models Draft& verify: Lossless large language model acceleration via self-speculative decoding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.950279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:56.015580Z digest=sha256:74e69fbdac5c1afa0859b1d77f37c43440d6c06ae93223ac76e176599e4cad0a

Observation da12a79a-721c-4b19-83ec-cd7f8a09eda0 · outbound

This paper cites Learning harmonized representations for speculative sampling, 2024.

Speculative Decoding Reimagined for Multimodal Large Language Models Learning harmonized representations for speculative sampling, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.802116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:56.084348Z digest=sha256:31b21f83933ed45ee4508fbac6650dc19fdd4fc36f59cef0165a90006cb7cca8

Observation 5bee170d-f6b2-4047-acff-1e972a7594c2 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Speculative Decoding Reimagined for Multimodal Large Language Models Xing, Hao Zhang, Joseph E

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.679378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:56.141651Z digest=sha256:abd6cb737d5ff91913f6bd947a757c98f9a00fec6b2b7abe1a422e9655bb9838

Observation 04cc1c5d-e3fa-4e99-b933-ece44387040a · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Speculative Decoding Reimagined for Multimodal Large Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.206898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.206898Z digest=sha256:f7e5e3aee27efa5f7ebb7504b733be54d41162ab5eb66fc8ff7b36cf4cadfca0

Observation 10f9c9d9-75e0-4ad6-bb35-d5d11f6e01c2 · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

Speculative Decoding Reimagined for Multimodal Large Language Models Llava-phi: Efficient multi-modal assistant with small language model

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:56.539136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:42:56.294364Z digest=sha256:7492366b95e47eb47c1e63c018e58484f146341b01f4fdde69dfd8cad5401310

Pith citing papers

Observation d950a3a3-52f9-4710-9137-a53eeb6e7b30 · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:37.884895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:37.884895Z digest=sha256:255370b0f733e0e87f9e5bd9dfe1aa7f73d6d8f34f7d96aa85e9bde320b9eb90

Observation 8ef8e9bd-1a19-4d4c-a9c4-16e926bbd613 · inbound

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning cites this paper.

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T19:34:58.789459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:34:58.789459Z digest=sha256:c261ca5dce2c3710077a173ef144a531c7ae37bc959094de5732a1d9262d1ff2

Observation 2212655e-0e7c-4a7f-b6d1-3393eb50a814 · inbound

SMART: When is it Actually Worth Expanding a Speculative Tree? cites this paper.

SMART: When is it Actually Worth Expanding a Speculative Tree? Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:02.508599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:47:57.421156Z digest=sha256:ba7ea26ea5439a4a54dce8294adc1f907bb6a815057a18372b23f476e0bdcede

Observation a35dae9a-87f5-42d7-8f24-55aeac61e1b2 · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation Speculative Decoding Reimagined for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:52.479619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:52.479619Z digest=sha256:a7e405a4d704731f32b2ba99a5e3455d4d04ad0c94ff3ef4a256437d0c0ec945