Pith. sign in

Paper Citation Record · LEDGER

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving

As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2508.12603.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12603 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:25:15.023517Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:40:53.581229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:29:38.015297Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation edb5b6ac-4e90-4cd9-b9a1-093d7367b839 · outbound

This paper cites 4" FUNCTION default.is.dash.repeated.names #1 FUNCTION default.name.format.string.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving 4" FUNCTION default.is.dash.repeated.names #1 FUNCTION default.name.format.string

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:25:15.815014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.849051Z digest=sha256:0cf5ee3e9cd824ee42b889f81776621e2284d072492f94dc8168700ca7faba54

Observation 6cf8f7e5-1640-4567-8129-72be0c90c85b · outbound

This paper cites write newline.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.854891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.854891Z digest=sha256:99d40c2e16c53c4908dece1d0562fd9a1be357c78d4a67355e5536e32c0629dd

Observation 1ce5662c-b6f7-4aea-b749-dbbb3b648906 · outbound

This paper cites A Survey on Multimodal Large Language Models for Autonomous Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving A Survey on Multimodal Large Language Models for Autonomous Driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.859646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.859646Z digest=sha256:1f973ac0c0aa722df2faf97bfb4a5d7e6986b30724d4c2608395c137967b0f92

Observation a172a0f8-e6d3-4e55-be87-27e4fa23e01b · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:25:15.785501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.865273Z digest=sha256:f45a9b0f02b800d9da03c7ad4a33f0be5c22a0d1979d2bf1e9714c091d026153

Observation 070e87f5-0122-46b2-822c-48dbc553f2b4 · outbound

This paper cites DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.869948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.869948Z digest=sha256:96f9e583ccb40a237c3d342f734ee053716de18e14b8abe74d57c57e62c0fb03

Observation c62a4c4b-40ec-4fc5-8dda-8ba953c1cd86 · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:25:15.642600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.874882Z digest=sha256:294c6bc7ccfb61f2bf8f7a48c82c36016708201e23cdf3dbcd15af2890679ab1

Observation 4c7a3a3a-4d1d-484a-a985-b2cf339231e5 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.879443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.879443Z digest=sha256:7c933c3f1d5f14653917124715d75c4623c4e004574dc0f57289f56f2cab35e2

Observation 14c41f62-c223-45d7-ab8c-45a5256bd57b · outbound

This paper cites LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.884679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.884679Z digest=sha256:22a8113fdbe61ac45597c9187264f5657649ca49fdfed21e7cc9bf0c55991833

Observation 5950d4bd-3052-484e-aa3c-fb32e645840d · outbound

This paper cites DriveGPT: Scaling Autoregressive Behavior Models for Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving DriveGPT: Scaling Autoregressive Behavior Models for Driving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.889326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.889326Z digest=sha256:d8794125e68836896faf03463b00dee11a3aa8196601455f0ac938cbf23cf0cb

Observation 5ca50948-78f0-49f6-ad20-0aa98cdaeb49 · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:25:15.769315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.894087Z digest=sha256:26dfc7c6b9c0acbc8062afe91e4f1012987bf1507426e2756083ee0512845b88

Observation 7ae2e67f-303f-4c45-812e-e9f7ae9ed0ba · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.898535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.898535Z digest=sha256:654859d94904a245f1b7bdb990ac72e4b4835deee969f4b9cead98775f652c5d

Observation a7e28159-6215-4e3e-8d7a-9da8172801e5 · outbound

This paper cites DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.903303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.903303Z digest=sha256:c4cb3b51cc78698160c2be5bb74a2a606d6b860b053866cefe1c774a87faa204

Observation 0988a1b5-94dd-4335-88b3-4ee97704786e · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.907908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.907908Z digest=sha256:c5ca99861520aa8afa2d2a063006ae84deffe9a11f359b8b71338215c1593782

Observation a85f98ac-2f67-4feb-be41-89234914192c · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:25:15.753288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.912630Z digest=sha256:72bc6cd9dbd2b2d80b8b2c7c7b3bceec9ebf33e8f200a3a28830603bd4dd31e1

Observation 3ed956ca-8230-4dcf-8c6b-2c156ae2dba1 · outbound

This paper cites Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.917027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.917027Z digest=sha256:085a957f3cbe2c3887b4ff2997f79b012d67b34612786f5f2e4d6fb9d7157be6

Observation a0deb9c3-430d-40b6-af0a-8bed90c44161 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.921543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.921543Z digest=sha256:a6649aec612daa4a341709ec1afb370b49875add2b8e1ba6f11ec6502db21bdf

Observation c1fae45a-ce62-424f-98e6-53329e7ea804 · outbound

This paper cites VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.926265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.926265Z digest=sha256:a1bbc70ddabbe0391385fcce5e65e514f124352e4c41eea8a7d26606f3ed7bed

Observation 6135445c-b0c4-4944-8518-97b151d6176e · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.930877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.930877Z digest=sha256:3d7003c98f98ac720e0926ef7e94785a2505a4146a74212112020d3bc45196c0

Observation be4665de-631c-4724-af0e-127ef806586b · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.935617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.935617Z digest=sha256:e4fa300d5911b62f361df046cbefce7a61ac7dff30052510411a8fe49956bd6d

Observation 20996183-76b3-4cae-963b-667ef82e9cb0 · outbound

This paper cites Sohl-Dickstein, E.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Sohl-Dickstein, E

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:25:15.736691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.940559Z digest=sha256:16fb732573a507e36157014f34530d473c1b2c4f8e46df891dd89c7229cb3c01

Observation 2622ac05-761d-4208-b187-f9261418cd7b · outbound

This paper cites an unresolved cited work.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.944929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.944929Z digest=sha256:df4cced8207566ec1fd9c1c93462d2c541e58270e6ce8c458bf8aadc71631fcd

Observation db56dc24-9beb-4f3e-8fe7-12861a24e098 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Score-Based Generative Modeling through Stochastic Differential Equations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.949427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.949427Z digest=sha256:d7424e6b650d8fdbff4a86c406f73f3d36e72bd7a39f6cc5b1f6a78f9999e643

Observation fd6a6fcc-515d-4094-ad35-f8cc59fa2967 · outbound

This paper cites Large Language Diffusion Models.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Large Language Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.953864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.953864Z digest=sha256:b167fcf49155fd465a9d61c487a3fc78639639b6e7e9a507d4ed297d3ccc81f5

Observation 481c74a2-3e3f-4758-8fa1-a0f83f09630f · outbound

This paper cites Austin, D.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Austin, D

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:25:15.708776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.958468Z digest=sha256:d44e613c52d94488704986bd14ddfc7191b203aed69795770c98dde99553a40f

Observation 01a247c3-4c09-4a01-9db2-45e6febd95ee · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.962924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.962924Z digest=sha256:6309729cab9c3db6200d96a191b23b3d9e8601ef71cc9db9ddf3665660836cdd

Observation a578e6ae-d826-4f4b-9f70-c28694dd7da9 · outbound

This paper cites Scaling up Masked Diffusion Models on Text.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Scaling up Masked Diffusion Models on Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.967767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.967767Z digest=sha256:90acb2810626081b133fd2ea613c9c0017db6c6c883351ef90f9c8d76412d2cb

Observation 4a2132a4-071c-4a55-bb50-ab758f0796f7 · outbound

This paper cites Scaling Diffusion Language Models via Adaptation from Autoregressive Models.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Scaling Diffusion Language Models via Adaptation from Autoregressive Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.972325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.972325Z digest=sha256:a5a77e9fb35de18d18d1f566ba43901ad0748cc14c3eb1ef9f03bac4c3c6a50e

Observation 14f56680-3ca2-492c-b40b-a6e36f27482e · outbound

This paper cites On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.977204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.977204Z digest=sha256:98cb5cec07c7376705846ed001f976f92cb531a59ac6cc6166311d36cf0d385c

Observation d975345e-3c88-4a71-8d50-2035972636ff · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.981768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.981768Z digest=sha256:bbe2c9ac9afc5d15d5207a2017817d4abe58c653f0a365bedb827e2345a48f15

Observation eb204c35-728d-494d-968b-3cd52d07d108 · outbound

This paper cites Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.986125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.986125Z digest=sha256:6daa651abab6cc45724992e78f2a3487f75e1b26fad76332e41d30affd88e132

Observation d4cc1fb2-6a36-4525-a826-bea642becbdc · outbound

This paper cites Caesar, V.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Caesar, V

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:25:15.692781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T17:25:14.990875Z digest=sha256:d2c60c5e4d2ffdc63543123ffd783de81bd83ef3a76a046e8eca99fcaeaf5cf4

Observation fe38b8b8-4940-4816-b660-99bf82b67ee6 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.995230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.995230Z digest=sha256:eab22d7928d4c338393f1237652e1f7ef7c9a4b9923c656ad915f3b590d81f24

Observation 33399631-4d76-41e3-a37c-d413ec697a3c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:15.004897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:15.004897Z digest=sha256:762529ea6d3b5a308215cfdc68b6ba176d0fc60a535914f3c10b074ae41e1554

Observation f68ea2e7-3502-47f2-a83b-a8fd80be93bf · outbound

This paper cites dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:15.009446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:15.009446Z digest=sha256:b34bd955f9101a4fb1212195b46586c869998e0a6d0240bf1a414e233934ca26

Observation 3535b9a6-f48e-47db-b830-6a25ab853418 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:15.014532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:15.014532Z digest=sha256:d8132a7922c63848f9cce78d558e33af6412176de8604108109a03f65306a511

Observation 1acfb941-ee84-4fba-8875-8de455978eb5 · outbound

This paper cites GPT-4 Technical Report.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:15.019067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:15.019067Z digest=sha256:de247115c38e6cf942715728f411b694477eca532a512baa5405f03da42d35f0

Observation d2c02d31-849a-4dde-922d-ba786a720962 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Gemini: A Family of Highly Capable Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:15.023517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:15.023517Z digest=sha256:fa805a806ce9af62bd762137ec8c13c9dcf14f2d9e8087e3a83f89d06f05ce47

Pith citing papers

Observation 21b30d6e-e7a1-4faa-946f-b180a8dd5508 · inbound

OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving cites this paper.

OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:38.017036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T14:24:07.496389Z digest=sha256:c6edd6bd829eb482f03813b19d0ac09471b6ea64b9f344aa2414c3d9e0f6a017

Observation c0c61636-e4ab-4a06-808e-921822e59ef5 · inbound

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation cites this paper.

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:28.595180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:28.595180Z digest=sha256:d4de68a65d7884fc86efb4308f7b47719b15bb775122f9b1d067976ff1e29764

Observation 4bd99c2f-8510-436b-84a9-38936027b707 · inbound

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA cites this paper.

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:53.581229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:40:53.581229Z digest=sha256:28415fac16b8b8c11183820dc7fbd97e0f22ae60d5605c1ce761b228330dab15