Pith. sign in

Paper Citation Record · LEDGER

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

As of 18 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 8 inbound Pith citation observations for arXiv:2412.03324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03324 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:35:24.414685Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.887036Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.832612Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b80bf627-b4f5-4c8f-92dd-4d45c5596364 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.131649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.131649Z digest=sha256:9388a589032c7d89e7a0a990938b203d6b3269982993d6f75cd53cbb1fba1f66

Observation dffbe916-ce9c-43fd-814d-26e8436859e9 · outbound

This paper cites Token Merging: Your ViT But Faster.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.136453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.136453Z digest=sha256:da7d9e4fffacd5e1f0001ef7e0ab832c47cad36ee8e698fdc60a020c9c3c9b75

Observation 597db8eb-6cfd-4ddb-b0ea-a777332ef748 · outbound

This paper cites Language Models are Few-Shot Learners.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.141236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.141236Z digest=sha256:3f9e8c0f00d5b36bcffdadec4a473933a22575e3966457465d8562354af145e4

Observation 8bfbba58-3684-481e-96d2-33a833a80c18 · outbound

This paper cites InternLM2 Technical Report.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs InternLM2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.145088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.145088Z digest=sha256:44458b46ea5cccf2eb1fe785bc4752f121370753b81a3fb8451d46dee07a553c

Observation 7a680d99-9396-4a95-a4fa-dd313db57ba8 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.150592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.150592Z digest=sha256:b6507bcf94393ed31029b4476819ebfcfb0acd935cbb5912f626c779b20d717a

Observation cf64be47-647d-4e7a-b025-8120b4953f05 · outbound

This paper cites Diffrate: Differentiable compression rate for efficient vision transformers.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Diffrate: Differentiable compression rate for efficient vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.598509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.155420Z digest=sha256:1cca8c6c3a5ed9b14bac39c8cddc4238505e70cd9e46e6479cdb394cddd7269b

Observation 87d881bb-902c-427c-ad94-3bd277521ab6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.159958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.159958Z digest=sha256:f66d36a618d57a1824830aa9f336d83f5434d317a44fc078415b9084a989295b

Observation 7bb2ad88-8998-40fc-80cb-a5cbe7647c19 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.587132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.164533Z digest=sha256:a47533aa6cf47da13cc6ef7bf709c406970b977ef0863a22ceb12962781719b0

Observation 59c7c511-e49d-478d-b72d-f71617a54223 · outbound

This paper cites Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.575168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.168457Z digest=sha256:7840138a53f9d2d3c294cb444d1598d31708ccd937ec51f04e04dabdf3a864ec

Observation 21089d75-128e-4dee-8e7e-65fe7f6a4461 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.172408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.172408Z digest=sha256:ad84d15857fa0146d7c07b610ab771ec580fcf593567a0af8420c1b3fa911849

Observation 9a348b81-a753-410b-8de3-cc47f50e1c0e · outbound

This paper cites Variational Inference and Bayesian CNNs for Uncertainty Estimation in Multi-Factorial Bone Age Prediction.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Variational Inference and Bayesian CNNs for Uncertainty Estimation in Multi-Factorial Bone Age Prediction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:25.017858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.176457Z digest=sha256:2f1276960e3cd3c6aed492988b6e9df24b85f688d891cb783671fdb3b41c2184

Observation 9a0bda54-eb31-49e8-a12c-c55887089ee2 · outbound

This paper cites LM-Polygraph: Uncertainty Estimation for Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LM-Polygraph: Uncertainty Estimation for Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.180459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.180459Z digest=sha256:50dd2c2ae4cabb810be2cc6df508dc3407a5221e935e2637403cf6f0d0a9c81f

Observation 26ae7411-ae43-4cc6-aad5-088bb7696a03 · outbound

This paper cites Unsupervised quality estimation for neural machine translation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Unsupervised quality estimation for neural machine translation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.563403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.184754Z digest=sha256:f9f0da4578827ebc1a0e542dcc06cf61437e37c6a78e40269c5e8e996cc79787

Observation 162d0beb-03bf-4071-9ad6-6e14f8f408a9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.188429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.188429Z digest=sha256:c5faeb234974cde386b1e24cedd03541ecef1683fc5ac90b33e8b891d442ff93

Observation f0ba39a9-4c5f-4249-9075-fa8b7cc898de · outbound

This paper cites A survey of uncertainty in deep neural networks.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A survey of uncertainty in deep neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.551305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.192530Z digest=sha256:d6539b9cb181e81e0fabac8abb9f06cb75637e3b6a1bb6361c8eb590c48237f9

Observation 8fe0ac03-4f69-4af7-ba62-330c607f4be2 · outbound

This paper cites A survey of confidence estima- tion and calibration in large language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A survey of confidence estima- tion and calibration in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.538211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.196476Z digest=sha256:e53c4f8d2325a3474d10d44f343b3cf319d5f875a575b95baaba81bfa546d9f4

Observation 42afce4d-dc94-4a5b-945d-9aa562181af1 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Model Cascades: Token-level uncertainty and beyond

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.200147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.200147Z digest=sha256:528cc5251070dcdb1a31d619238a15e7de7c517d27bdd9f217b36536b7425f12

Observation 22c5d3d7-51f2-4bda-a0fe-adffab19e0b3 · outbound

This paper cites Dynamic neural networks: A survey.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic neural networks: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.525821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.204411Z digest=sha256:7275b7779237545be59395163993d2a9b7cfe283d62ace66b600d2a679b3252a

Observation f530a7f7-9811-4456-a029-46dba21dbd53 · outbound

This paper cites Learning to weight samples for dynamic early-exiting net- works.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Learning to weight samples for dynamic early-exiting net- works

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.512291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.208236Z digest=sha256:2b9176aaee4204a5f62752b037e222417f0b566a501b52e6383c081034a5be2c

Observation ae31749d-d47f-44f2-bf9f-42032d5546f8 · outbound

This paper cites Dynamic perceiver for efficient visual recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic perceiver for efficient visual recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.499337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.212183Z digest=sha256:8bdd3077c22694d76d9c8b381e3808675948a08df4dda08e79fcec3a2c6149eb

Observation 280d79c8-7b66-44b4-9532-ce7da5ca1a32 · outbound

This paper cites Latency-aware unified dynamic networks for efficient image recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Latency-aware unified dynamic networks for efficient image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.486571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.216022Z digest=sha256:ccdc8695dc38cb9134d93034faedee77f3accb70c9de8222f7e7775663bf6c08

Observation 547c8ff2-a1fe-4320-9230-035534c038af · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.224741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.224741Z digest=sha256:f7f835e6c4622f81fb5b0d69ffdd8c33bac9de2465d73f66050433dffda6eb94

Observation e5a79a57-2da3-41b8-84e2-d768c6fb0dd0 · outbound

This paper cites Mistral 7B.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.228819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.228819Z digest=sha256:9cb59648d362fabbbbced758a7928811a0faf363e90307a79941719673fb12d1

Observation ab0e875e-af33-4d1a-aa56-0b50ec8d0b3d · outbound

This paper cites Language Models (Mostly) Know What They Know.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Language Models (Mostly) Know What They Know

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.232497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.232497Z digest=sha256:6bd870a9606c57da66f4d5e12a0d66c2c2a09f5ef271d06bbde3a548e12fd8b9

Observation c19f64a2-204a-4c4f-b76b-6669c9e6c06d · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.236297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.236297Z digest=sha256:f4fb06082fabe0170eab74980f90a7da4efd2871fc3a04a0a8f047225bc89ccd

Observation fabfd2d6-482e-4881-9f79-0f3000040416 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.239987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.239987Z digest=sha256:5ee96966c645bfaff58908dc5feaf813b44e54570a5887a9614d0e2527fd530c

Observation a729fb64-efa4-479d-8806-f2c3748c4c1f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.244102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.244102Z digest=sha256:b524914c6b9c9625063cecc6da91e7923d6add671b6417b720d291f79f0c52a8

Observation f9ace29f-b2c3-4c1a-9975-52a0a6b3eaf0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llama-vid: An image is worth 2 tokens in large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.457111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.248345Z digest=sha256:b9c4f113e194cd197ad2f21afe3f3c64710b4bedd9ddea35f8d182d3b2e6705a

Observation 37753420-09e9-4d16-b5a4-4b970e057cb1 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.251928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.251928Z digest=sha256:4ef2a3551aaa610023e44c2d23117b55469766dd38d5f284f4b7219d70078d79

Observation eacf2643-c1c8-431d-add7-0b837a19ad86 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.255857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.255857Z digest=sha256:beaa146914f7de413fff347284d511c937e5de9623ce6bd5d5f9fbd008eda256

Observation 3c0c87d7-f737-454c-acb4-ab55915478e9 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Improved baselines with visual instruction tuning, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.444753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.260042Z digest=sha256:1ffc8278f02925bbfbb2afe9ab794c54cad395a78cc4b2e987a499b984ef59bf

Observation 7bc46c4d-9efe-4d0d-84c4-536956c22cb7 · outbound

This paper cites Visual instruction tuning, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Visual instruction tuning, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.418833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.263882Z digest=sha256:5c08d566195e31a83964ad003e41dc2cf3f9e69a043924b644379415c53b3894

Observation ee878242-9dd9-4d55-ab3b-0d4c9dedd592 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.407046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.267743Z digest=sha256:2827e6083c067e3e661f9e819826171f731994f59d5338b7c84ed2ab0e638958

Observation 3b44333d-2a1c-4329-9d71-99a573bfff67 · outbound

This paper cites A general framework for uncertainty estimation in deep learn- ing.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs A general framework for uncertainty estimation in deep learn- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.393865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.271700Z digest=sha256:4bf2ad32dda7121fcbe62065f9502c5cacadfc66f2c4fc27e87c3cdd56702c38

Observation fb6c33aa-7d86-47d8-b5c7-34e380f5c08c · outbound

This paper cites Visual Perception by Large Language Model's Weights.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Visual Perception by Large Language Model's Weights

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.275563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.275563Z digest=sha256:5aceb0c2dccf6a9b3c7bf3e965405af5b49bf400721edc1cfe6a0b65c55593d8

Observation 9d0b0378-4832-4928-8740-31b2c9434eaa · outbound

This paper cites Uncertainty Estimation in Autoregressive Structured Prediction.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Uncertainty Estimation in Autoregressive Structured Prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.279509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.279509Z digest=sha256:a5eebc9ecc824774e83bbd50e09137140d913e67eca9899706155b0aafee8f4c

Observation 256de611-863d-431a-a553-3b9a3ad11307 · outbound

This paper cites Generation and com- prehension of unambiguous object descriptions.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Generation and com- prehension of unambiguous object descriptions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.379430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.283433Z digest=sha256:f2de3f209b5f6ca4d8421e00f00991d62ef4263ad045274195399d3ca908707b

Observation 80bb03c8-50fa-4f9b-a971-e9f869ebe14a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.287112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.287112Z digest=sha256:449984d5440b8b0e7983bb99ad28d2fd63427ea999b22070339848b585ef8795

Observation b65d3200-f144-4f79-ab5c-9f1837cedf35 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Docvqa: A dataset for vqa on document images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.365988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.290804Z digest=sha256:3661eb93ce4586edb2f3306883202cdc33ed5ba3bff02383d90234e3e69b44cf

Observation 6daa8012-bc72-46fc-bb99-f3ad33ae1f73 · outbound

This paper cites Adavit: Adaptive vision transformers for efficient image recognition.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Adavit: Adaptive vision transformers for efficient image recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.353525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.294063Z digest=sha256:2bfe31bd95349d6e1ded464ee931afec8771391bdde291335eb2ca1bf7b8df06

Observation 750d9100-fad5-47cb-ad7d-a4d9a3e72df4 · outbound

This paper cites Correcting Length Bias in Neural Machine Translation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Correcting Length Bias in Neural Machine Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.297510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.297510Z digest=sha256:1b901feb9325d098553ad4d4e7b3a3a84fdb54bc3bbce32f7532d6dde1afa0ec

Observation b315dadf-151e-41d1-b558-947a3f938775 · outbound

This paper cites Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.341116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.301667Z digest=sha256:d7678d5ea54439d824b9fead7fbd0837cf8ddbb2cbb220d97af91d64d9d9a05a

Observation 54226b09-ad0e-4586-91f6-30610c57c2c4 · outbound

This paper cites Hermes-2-theta-llama-3-70b.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Hermes-2-theta-llama-3-70b

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.330393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.305277Z digest=sha256:ba0396676c787b22d8dab854412d06d211d496bbc66290d76bd6446d64eed3f6

Observation 1f7da94c-1103-4914-828a-614fdb2de761 · outbound

This paper cites Nous-hermes-2-yi-34b.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Nous-hermes-2-yi-34b

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.319561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.309019Z digest=sha256:0758fd79985808c60e7d671d8935f4de1319e4ca77b9d32f983c8382bd32f466

Observation eaa61918-7a26-4005-9505-0f7e8f136789 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Learning transferable visual models from natural language supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.313519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.313519Z digest=sha256:8ded3e43c334224237c39140d783cde2ecfead23d9917d0fd7f0241ad3f2e7ec

Observation 10c028e6-4ee5-485e-877b-a40dd3b1076b · outbound

This paper cites Dynamicvit: Efficient vision trans- formers with dynamic token sparsification.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamicvit: Efficient vision trans- formers with dynamic token sparsification

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.302495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.319170Z digest=sha256:ea01f20372c68cedbf0aadf78e95aca1956b3195bd7c610c4ddd8095baedff4d

Observation 8048377b-bdad-479d-959c-b7b1e2d4f570 · outbound

This paper cites Out-of- distribution detection and selective generation for conditional language models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Out-of- distribution detection and selective generation for conditional language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.288698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.323178Z digest=sha256:be1d63f6a10f13e8c8ad8bca8f91a0a0f2792426ec24223f77d904293d33e630

Observation 5cc36cde-1b41-442f-848e-46e7bdff52f8 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.326720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.326720Z digest=sha256:abc4263b3ce93cead30da4802f3067e7ea306d6ef4747dea7d804959087b4aef

Observation 4eee219d-cf0a-4a03-a1a4-1cc5f006f489 · outbound

This paper cites TransFusion: Generating Long, High Fidelity Time Series using Diffusion Models with Transformers.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs TransFusion: Generating Long, High Fidelity Time Series using Diffusion Models with Transformers

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:24.709346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.330001Z digest=sha256:2f00b70c30d946c81911e04b54610b4f12477d31c8e27d2574ba51c4d1450320

Observation f1d2419a-4c06-42a1-ace2-857c86002f0c · outbound

This paper cites Towards vqa models that can read.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Towards vqa models that can read

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.264339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.333889Z digest=sha256:96c0cd546bbe0e7d5f0ce8d5ed4350e3aba1c0303b03cb7da88a5bc53cded670

Observation a3f622d8-a03e-4e49-a129-718206b33d4e · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.337472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.337472Z digest=sha256:e4ab90c7a591aac3c731a4f7c7866f6c98299a3f2e05addc1204ffbdf16bff4a

Observation f410656b-ceb0-47da-8a38-6305159b2b8e · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.341013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.341013Z digest=sha256:3e2b3b20ae376b5cd86df9fc7684042479ed88343352caf4ea43bd1456920c6b

Observation 06c1507d-9f7e-4a92-99dc-209c79fc2735 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs LLaMA: Open and Efficient Foundation Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.344055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.344055Z digest=sha256:79cc55ccf041c8888378b063e882db649f1a226bb8a493ffa814475b64380b20

Observation efbd56ec-f1a6-43eb-828f-672efa1fa9e8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.347240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.347240Z digest=sha256:a5a46860b58a9b0ad376354785ee757ef6b7039a3772301c1eabd236d9c7516c

Observation 8de10cbe-8d0d-4f19-8629-9b34d56b30f4 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Emu3: Next-Token Prediction is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.350738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.350738Z digest=sha256:6ff4f0cffd58f9c607e21fac96f990fcf4e02e5723a6394082f887153a69a3ad

Observation aa046629-c8b7-45d6-a284-e4906f8f7665 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.353965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.353965Z digest=sha256:7d2ee6560cf28724e4f4bb535d82bf176c7a400888ac387f6a5a112359eea723

Observation 068bc7e2-f04b-44eb-8bbd-a7b133f09212 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.357686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.357686Z digest=sha256:6a4268a79f45be8100ea3e2a62c267711a59830502fcbae6883b647edbff84d1

Observation 0a008773-a7cf-4d40-8821-176c5e918590 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.362275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.362275Z digest=sha256:022334e74b9f1475f38cf3f690834339aab1604c1c49c1e64533065f4e1a4668

Observation 24e3a296-f2d8-4f6a-a149-160a987d8658 · outbound

This paper cites Qwen2 Technical Report.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Qwen2 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.366370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.366370Z digest=sha256:f7fb81dbc89b30002ecd2a97adc7706458671e5f0d92f81cf045a090e438fb90

Observation 94e33742-a456-418b-8ce9-4b18cac4efd5 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.370458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.370458Z digest=sha256:145f8fdb180e521c51fb11152ee406af9f97a791b021637ac4478ef2fada931a

Observation 0ee9a5c0-53d0-4abd-9d99-71082969652e · outbound

This paper cites Detection of Word Adversarial Examples in Text Classification: Benchmark and Baseline via Robust Density Estimation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Detection of Word Adversarial Examples in Text Classification: Benchmark and Baseline via Robust Density Estimation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:35:24.543601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.375218Z digest=sha256:b66c10a41ec1d08c72698bc4026d08eab331e7c92ceeb7abbbb2dc7f8acad6f8

Observation 412e5bb1-1918-49e7-83e3-a5b88d4495d0 · outbound

This paper cites Modeling context in referring expres- sions.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Modeling context in referring expres- sions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.235115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.379188Z digest=sha256:f50d45625d3bd50db2e296bda418a4e284df51ecb14380fd3ae468dae2e0a4e6

Observation 136acb78-8d3d-47d9-8e9b-927cd930b3ba · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.382882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.382882Z digest=sha256:30f1961770370648a2a37446887342003f3cd0ee88db129a368d785f794d8b0e

Observation 67df8694-328f-484f-bca6-b88b44094052 · outbound

This paper cites Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.386895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.386895Z digest=sha256:fc32ecb44bd2c66e041b35dd3c4d81805cc72b9eaa3832de9dbca4f91b30a2ce

Observation 185df56a-4a07-48c2-9de0-536b5bd57a45 · outbound

This paper cites Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.205492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.391458Z digest=sha256:adb2b83b0dc520679489ae20ae54f111a70a2c10f4dbb859a43709d969656d76

Observation 126c116e-66f3-4582-89ce-82949449cbe2 · outbound

This paper cites Sigmoid loss for language image pre-training.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.131544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.395578Z digest=sha256:8935e2fa4d102d055991971a9cf5e1529b8e435b41596872546d527058ddc5c8

Observation cf0e0db3-b9e4-457f-9d14-79a47e836c7a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.399461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.399461Z digest=sha256:e61cfc890afb6af0413ea3d61f949f05725997e138fc73868f16ace7d11cf83c

Observation af0d7fdc-e02d-4373-92f8-76bd17dcb453 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Llava- next: A strong zero-shot video understanding model, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.403563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.403563Z digest=sha256:15a19a2ff7d9f6c211a73fb8023f7e96b90476817d3767b8265ab596cff08d73

Observation 202604d5-e842-4a4a-a66c-965ad95557eb · outbound

This paper cites Dynamic Diffusion Transformer.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic Diffusion Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.407029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.407029Z digest=sha256:3a7f6cf008c45aae9f89f52ec4ad552c18e4ec5c31efd0450606d1aed668609b

Observation 45205573-571c-4b89-b3e7-df6f90924a63 · outbound

This paper cites Dynamic tuning towards parameter and inference efficiency for vit adaptation.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs Dynamic tuning towards parameter and inference efficiency for vit adaptation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:35:25.111180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:35:24.410825Z digest=sha256:60d2fbfa7c486f90d6caed735498d184dc0b5a79150497f6044322820e3f1975

Observation 9476b23e-cb47-4439-82a8-3a88b722b234 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.414685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.414685Z digest=sha256:57654d639f14538df0495e78a39ebb164f84d0a069a2f91458e863165a37c8cd

Pith citing papers

Observation f9dca469-3ca7-4d88-a195-6e300ad78f29 · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.908074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.908074Z digest=sha256:e3481f1661b40ea87c4904aac0211d0b73241bf1c7e8e216d09168aaed3bbba1

Observation dd486c63-78d5-459a-be7e-d5ee4d9da191 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.571714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.571714Z digest=sha256:dc7db7fd2e3b1eaccb9e259fd23ff9b3f2a3f3f1e5f96acb7117b8cda174580b

Observation a5b0afe7-b960-49c2-926c-4f6ae91eaa2c · inbound

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection cites this paper.

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:50:08.595852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:50:08.595852Z digest=sha256:6ece17bcb71c41356a2a979e927b6b0f6c5aa96fb8575bbbd77a00fb8deb7431

Observation 21413577-52f0-4762-ac3a-89d7f774f25d · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.835408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:2cb0ce36f6be298290df3305ff2703dc310a0c528c81704dc489a51f18386b76

Observation d1e56ae1-92f8-45ce-a098-69b64aae0a9d · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.887036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.887036Z digest=sha256:e717997c427907feb44886fc692ab01df9237792565c556c53ad1741f2b405b4

Observation bded14ac-604b-402d-8a64-1dfa3143d598 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.547382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.547382Z digest=sha256:e05c1e3cf0219e25c5554e0cc0cca60e696859f512c21ddd0bd01eda6541234a

Observation 40c8c8cd-cadc-47ae-af50-4b537eae3337 · inbound

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding cites this paper.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.639839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.639839Z digest=sha256:0697ce0de6b01eb9138129013a4494401724949f716c51a3357c6bb42b8c7140

Observation 69f62ef9-cb4b-4adc-950e-d49c7a52aed6 · inbound

CARES: Context-Aware Resolution Selector for VLMs cites this paper.

CARES: Context-Aware Resolution Selector for VLMs A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:44:42.492230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:44:42.492230Z digest=sha256:e3e90e0fae30e714a742244bfdfb9b66115bb811a370ca2a4f152a355cbd6d24