Pith. sign in

Paper Citation Record · LEDGER

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

As of 18 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 6 inbound Pith citation observations for arXiv:2505.08725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08725 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:53:10.078556Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:26.525113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T17:10:25.159616Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact2
  • verified fuzzy46
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af0e2db5-97df-48e6-ba34-eef2b8fb4744 · outbound

This paper cites Visual instruction tuning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.606344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.606344Z digest=sha256:5525765caea1555d632d8442e4e87c38380ced0c9a061fafa4cdf3fa4d00fb7b

Observation 02b1dff0-eea1-4800-95eb-7dd44a64bac4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.610547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.610547Z digest=sha256:32e3ad73b8066f4f41d68209420332fba8864c4dfc1411b7fb37f9efaaa06f8c

Observation a87bd67b-a921-4263-8acb-e453cf170b50 · outbound

This paper cites GPT-4 Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.614593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.614593Z digest=sha256:60d94f0c9851908c9c01a581745aaf2f219fc9ca8a998c28ce015070246a0ec2

Observation 15eb341e-5305-44cb-8388-04eba5b584f1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.618465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.618465Z digest=sha256:8e6b48ce860a117a9826334ee4412ab5341a8e9a44a6b0dea419ebd161faa7a6

Observation 5ef1f138-a931-4f50-93bd-25247ad7119c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.622381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.622381Z digest=sha256:5101526296a1d0b59ba3568cc50b069fa72bc35139ad51c1df8480aecf86254d

Observation b3f29d88-3a74-43d1-bdc2-cf4ea0137144 · outbound

This paper cites Qwen2.5-VL Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.626041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.626041Z digest=sha256:4a5bb3951d95b958491cd103983cf69af220d0d7660e6ff4b09b9da6d4853436

Observation a200733a-a8c2-4235-be19-f139feb09e1d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.630196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.630196Z digest=sha256:01d16087ec897963c9fc293d5b640c7e1a5261c2d471146720bd2628d397e1e5

Observation 0ffcde34-759e-40dc-b65a-8b9352e44612 · outbound

This paper cites Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.633985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.633985Z digest=sha256:cc8b90304e701003f22e3a40902d4570b0e730c7bd59b694d20e89c4b5fb36e8

Observation 241fb771-5cb1-4263-a7f6-dce7ab92e8d5 · outbound

This paper cites InternLM2 Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.637512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.637512Z digest=sha256:fab418eb215840a3f60631c18c1fc145892f6093441c6649903a105e665ab6ff

Observation fc2a629a-9792-4350-b954-26c79c47f028 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.741034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.741034Z digest=sha256:9a61ef92955e1772f366774a8a7324f0c60dee8a9639ac14fd250ac49dbcfbc2

Observation b4472e26-9493-49ce-8ba0-a4559cdb6f64 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.744895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.744895Z digest=sha256:ed996d67e8d44f77dd4b4847e4cc03012c8e7b815f4ff60fb898e93cd1a1770e

Observation 150e5978-133c-46b3-901e-938a75c08f7c · outbound

This paper cites Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.748705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.748705Z digest=sha256:62412396c664f29cda39cfc86962fc82a7205a6167cf8c760a9b0660cf9f0789

Observation da94fea8-9b4a-4abd-a4ed-6ac2c33d19fd · outbound

This paper cites Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.752104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.752104Z digest=sha256:695cd9e0eb6a7d715c5c668efefb9a0575f0a076bd33a31bfe9b6f0f9ac31fb8

Observation 13a885fc-63c3-4161-8387-c7ea46159ffd · outbound

This paper cites Embodied understanding of driving scenarios,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Embodied understanding of driving scenarios,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.755779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.755779Z digest=sha256:ddb4108f10cf78659d1d481f9b83919b7b15ae1e9de046a967b5f7343e23406a

Observation 72c513d5-8c01-4cc3-b493-8013c377c597 · outbound

This paper cites Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.759471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.759471Z digest=sha256:261616f37bd802fcbc34e49644128f2237b98050f0dbca0af9321be194a55dc5

Observation 0c2f36b6-edc5-45b5-8277-056f7ea8f29c · outbound

This paper cites Drivevlm: The convergence of autonomous driving and large vision-language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivevlm: The convergence of autonomous driving and large vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.763107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.763107Z digest=sha256:eab4934d7090e81056b3e9479a2c5526ed0389edfb1e42d85dcb7397ef27ee32

Observation e0c4f03a-4367-47b4-bd00-9f2f6084ae81 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.766656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.766656Z digest=sha256:128ccdef89249ec3517b1b1e1f89bd40870b81d85593f01af1d8a05c7e334136

Observation b1266064-9caa-47b3-af39-d24e24de19d3 · outbound

This paper cites WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.771166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.771166Z digest=sha256:0b6beb8edfe07f425977f9ff6d478fef63c033cc3b205b87e4977bfd31e9eb4e

Observation 80804ced-713e-4aeb-bfe1-5f89af80d114 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.775204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.775204Z digest=sha256:951c73e17d97fae1f1a408128b90c8222840c11cf1de6da1148eecd34b2347d8

Observation 057b03da-0b8b-4083-97b3-9f1a6283553d · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lingoqa: Visual question answering for autonomous driving,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.779230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.779230Z digest=sha256:726d2a87cc5affcc7d510f79b93251faa5e3b4debdd0b0e16c16606f269a6772

Observation c0a35f8a-2bbb-4e2e-acbd-042353accd74 · outbound

This paper cites Automated evaluation of large vision-language models on self-driving corner cases,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Automated evaluation of large vision-language models on self-driving corner cases,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.782949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.782949Z digest=sha256:d0a8d30af1a5cdd378f450dae83842e1e75286533ee868004e8c7d2f525a9ac0

Observation 1a66bee5-b1e7-4988-905e-38537054d902 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.786674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.786674Z digest=sha256:94b0c4ed16fa05c72f8f36844e39f24e62a9ee570489d0561e0344261e8cc826

Observation 7e553ca0-123d-4a98-a893-c797c35e4a75 · outbound

This paper cites Drivelm: Driving with graph visual question answering,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivelm: Driving with graph visual question answering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.790238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.790238Z digest=sha256:e8863adb75a7fe9f137b1b213905f52f2c63288ee76580749bd1790d3111e748

Observation edc6e766-13cd-4d68-81b5-1b6ea6c962a3 · outbound

This paper cites Language prompt for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language prompt for autonomous driving,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.793816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.793816Z digest=sha256:3f961f9ed4ef630188be59357d7cd4556f470f02d152f6c758735fea672d4f52

Observation 35ee3faa-f307-44e3-9b6c-d28be498676b · outbound

This paper cites Language-Image Models with 3D Understanding.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language-Image Models with 3D Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.797343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.797343Z digest=sha256:81f2427b184a62bb5b72dbd0ce3301b329f4392becf153f71439ec8719dd47f9

Observation 9d731add-91c8-40da-af82-0540c830deeb · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.801297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.801297Z digest=sha256:2bb5e71607cbd2f8414394d799273373ae3196a55e0fa0b6f9770764ad1f2e31

Observation c816d584-4c43-467a-938d-9a22b791c108 · outbound

This paper cites Talk2car: Taking control of your self-driving car,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2car: Taking control of your self-driving car,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.804694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.804694Z digest=sha256:cb268f3de522c8e1b9c4992fac91fa789f519929dc1b5553733e7b91f77a142e

Observation cc639249-fb3a-406e-bdc0-a2413c90e37a · outbound

This paper cites Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.808274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.808274Z digest=sha256:baa6c964e7f552eb7275a4e271761433137267dea5ad50d0bb235ec8c97f9c8f

Observation 7ab54154-d3f1-4af2-8559-f7a4be381719 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Improved baselines with visual instruction tuning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.811720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.811720Z digest=sha256:1724f1d559c530c3a5da00f11730c2018df9662a6714a7b68b8bff759a502a16

Observation a236f099-dca0-40ab-ade2-e099f95af1b5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.815225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.815225Z digest=sha256:2167752be204a76508e9980778649ca38e395b066ac0ea69e449fb275f30514b

Observation b77ab0f0-04c0-49da-9d91-68a022e90d8e · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.818802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.818802Z digest=sha256:4da2a6c2735b264db7c7f7ace275b7c5ae214a86b25768d2793e34507f177754

Observation 80c44b41-cc05-47fd-8529-3f803263c0e4 · outbound

This paper cites Pix2seq: A language modeling framework for object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Pix2seq: A language modeling framework for object detection,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.822281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.822281Z digest=sha256:3357e028b5c6c5e2693c9cd513ede4c5419e5531ed12bfbd5a72240fc4d40b51

Observation 509210da-73bd-4da1-97e3-0c23bb99be10 · outbound

This paper cites Petr: Position embedding transformation for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petr: Position embedding transformation for multi-view 3d object detection,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.825816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.825816Z digest=sha256:d8fd9cdb8d97cecffb6e93f217170cc6cd6637acda8bc11fe1ce57e50009118e

Observation 05ce9192-cab8-4dea-84d2-aec9baa3765b · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.829420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.829420Z digest=sha256:79b88d56a7dafe6ac04657d24dc1dcd450bc2340e786448d6590d7dc7726f2a6

Observation 69ec23d1-b54e-4a2f-a360-7286e37d526f · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.833021Z digest=sha256:3ff8fc6b43a1870a021541deef938c38f5a38a46642aaa511f545bc3c0c0b226

Observation f18875db-46df-4a97-8b72-06eda01de60c · outbound

This paper cites Grounding human-to-vehicle advice for self-driving vehicles,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounding human-to-vehicle advice for self-driving vehicles,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.836649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.836649Z digest=sha256:399437480b5a13e1ef5c995e36da99cdb4f6b579bba18e2de34c36547fa4e1b7

Observation 9b1f4411-3818-4d74-b3eb-0d99dd8081e6 · outbound

This paper cites HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.840222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.840222Z digest=sha256:b1463f95f241a1678c513ec1b2d18911d92d80d545fa57f3a09dfd8616374c52

Observation 81c503aa-8639-4c2a-a326-8e878e824e3e · outbound

This paper cites Drama: Joint risk localization and captioning in driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drama: Joint risk localization and captioning in driving,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.844453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.844453Z digest=sha256:d2112c039b9de4959d5fb3ebd16303877f4b6d191a2adae1acf0989c44df368a

Observation fa4099c1-778f-4a97-bf85-ad9358af885d · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.848095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.848095Z digest=sha256:b984b41d9f7bcd1a19701aae9406dd1c5e69875c40fbe53a7c8fcc19fbbabd9d

Observation 076f3704-9b96-4ee8-ad67-a00a728329a6 · outbound

This paper cites Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.083936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.851562Z digest=sha256:34150cc0e7ba39596f4f72d2fea660d5e97274e251fc0bb63b51e889c9d39e79

Observation 8c525c83-1dde-4239-b171-898e93d02cdc · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.855200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.855200Z digest=sha256:57bdfd0861de1062af44e2185f1f72e6004758a6823032e203ed1d78aa929d19

Observation b0351992-9a9b-40fc-a141-ae809611e947 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.063220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.858638Z digest=sha256:6cab0253a5d2f44e3a3649022ecc1a7fafaf25e1357ab0aeffd0f44c19081565

Observation 72584531-c02e-4537-af12-46682891f22e · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llava- next: Improved reasoning, ocr, and world knowledge,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.049928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.862308Z digest=sha256:21b0bda230576b9f96a85b4e00658589fbd5eaf9fa543b7aeb0efa8f25ade5c9

Observation 15b41a1e-7211-4ac7-856c-dfebccaf96f4 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Learning transferable visual models from natural language supervision,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.037147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.865801Z digest=sha256:2e64584a14d7876f9421395d956fdfee5dab0620743353fa8640875c470002c2

Observation 62406d58-a7b0-40aa-b1d7-0805bdf5e0b1 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Sigmoid loss for language image pre-training,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.025373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.869255Z digest=sha256:892d68cbecb9736ef5022b727cddb1a4b0e09c30705106e15cece4d60b44aa1a

Observation 569009a9-a579-4d5d-99f5-95ba72d36488 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Dinov2: Learning robust visual features without supervision,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.013827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.872768Z digest=sha256:5d6e44838985955e1fbecae34bad4e75e87845199ffc818e9da17314842a78a5

Observation 2e91f92a-2861-4361-a337-d7a9656543d7 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vision-language models for vision tasks: A survey,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.876303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.876303Z digest=sha256:cdc110ba59ae268817ff5382c7dae47ae9e1fbc9a2888d4f9c557b4408dbbae0

Observation ecec2b5d-db96-4866-afa2-18be6461c90a · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.995209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.880018Z digest=sha256:59c77f2316de4cdd83d1b3533a7274f82a9021da5c521375611ab79bd8bad88a

Observation 9e14063f-7254-4a27-a907-eb4952f5b386 · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Uni-moe: Scaling unified multimodal llms with mixture of experts,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.983833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.884186Z digest=sha256:8684990a4286d934e51da7791b97e32683dc0a109b93a80ae960b03b98c3d215

Observation d7a730da-cb98-432f-a98d-07c5897e15e7 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.972612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.888040Z digest=sha256:6797126f7d768a6b8547d6de094d546a47c0ad78ac2d60dab151865ac4a8fce8

Observation 9f16af07-e6cc-4f25-b1e9-18f1ae6e4f7d · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret: Refer and ground anything anywhere at any granularity,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.961906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.891792Z digest=sha256:38ea139ce97ef17fc599a3cb429c18d69dff60572ca4c79e0acfac44b153d032

Observation 96470717-e76d-424b-bbc2-d0e1043c5001 · outbound

This paper cites Ferret-v2: An improved baseline for referring and grounding with large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret-v2: An improved baseline for referring and grounding with large language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.895836Z digest=sha256:2bb9ac6bcea473a9b2eaecacd2cf8b6b2f772e3542545544fa19d7e2cb75acae

Observation daad5a10-c313-49e9-a894-0087c9786e26 · outbound

This paper cites Relationlmm: Large multimodal model as open and versatile visual relationship generalist,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Relationlmm: Large multimodal model as open and versatile visual relationship generalist,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.940250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.899785Z digest=sha256:70cc750507dd918ed145d93bcc381aa4c388c7e62bf040065768172093ded1d1

Observation 0cf9eba9-e938-443c-9dbd-99c564403b34 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Groma: Localized visual tokenization for grounding multimodal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.929883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.903100Z digest=sha256:ef09ef4fcf00ade318909380a976957ba7895ec63a20d154c16aa8c4d098f447

Observation 40a5d163-4b92-494d-9f53-009a63553b0c · outbound

This paper cites Segment anything,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Segment anything,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.919161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.906547Z digest=sha256:63cda0a8429c84eee956aa6a050f3a9b0a1d4e176dc97cdb9660129e6974b72c

Observation 513d95c0-5368-4b54-845a-297bb16acb97 · outbound

This paper cites Masked-attention mask transformer for universal image segmenta- tion,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Masked-attention mask transformer for universal image segmenta- tion,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.909006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.910064Z digest=sha256:590e121eaeebbf8d3bc3cee7a95ddaacedfb5456a2cf704b110b6ca8e87fa651

Observation 8890a925-f7ab-44b7-8f6b-fe1c213b118e · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lisa: Reasoning segmentation via large language model,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.898061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.913619Z digest=sha256:aad583959e62016b6c6f9bc96094c77af3bb80f917294f68e020624942a77ed4

Observation 7293f22d-464a-45d6-8aa3-e02dc3fedd46 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Psalm: Pixelwise segmentation with large multi-modal model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.886398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.917950Z digest=sha256:e6d887ecf328fbd5c13c0a1978f54ad7baccf369d3e568d1279f9538503c6408

Observation 95e58821-a84a-4ab3-a11c-13d0928aae29 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.874930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.921672Z digest=sha256:3936853ee5d6221c9d12d21e3d360edc633bdcabeb35b3f45254b374f1ac5f06

Observation 355d12dc-ed99-4b68-89bb-ebbfffa8cd80 · outbound

This paper cites Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.863558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.925108Z digest=sha256:ddb40a0a76c13f38d407ee1db44aaced593177697fbd90449a6661e476b26e76

Observation 08b527f2-7726-4b34-bb2f-561516633621 · outbound

This paper cites Tod3cap: Towards 3d dense captioning in outdoor scenes,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Tod3cap: Towards 3d dense captioning in outdoor scenes,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.852934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.928355Z digest=sha256:580548d7031ce9df36be09d0f53c0349b5d7daf6b7a10da0a941dfaa43462b8e

Observation 56bc0e69-faaa-47f5-92dc-7ef0993400a4 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.931808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.931808Z digest=sha256:0e6e25904436f57cf7b3edf11b9e504fd01130d25a6ec8e6c06106ffe1e5ae67

Observation eeeb4bff-6e2e-4bb1-a37f-74f3fd7af46d · outbound

This paper cites Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.935135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.935135Z digest=sha256:755d17d035929c6d13fbd169a22ecf290b5805889370c7ef8ebcab342873faf4

Observation a968a4f6-892f-47ca-8685-b032c9d935de · outbound

This paper cites Gpt-driver: Learning to drive with gpt,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gpt-driver: Learning to drive with gpt,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.841943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.938984Z digest=sha256:04521927cafa8da257bc663f30f7350136f6d7905e634f925dd0ecefd06a8c60

Observation badf7126-f77f-42d7-9abf-e50ca4687240 · outbound

This paper cites Making large language models better planners with reasoning-decision alignment,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Making large language models better planners with reasoning-decision alignment,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.830737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.942801Z digest=sha256:8a01f57d837bed311b3996c081b9e110e192d1d849f350505bebd6243ea86b62

Observation 120ab713-3f78-47f2-ab94-13d325e1a21d · outbound

This paper cites Distilling multi- modal large language models for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Distilling multi- modal large language models for autonomous driving,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.819866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.946579Z digest=sha256:407a37744ff80e589ed4e239e353f5153a476f2c9fa6f1752f99574de946d48f

Observation c75dc518-574d-49aa-b789-f44dc378b1e9 · outbound

This paper cites Generative plan- ning with 3d-vision language pre-training for end-to-end autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Generative plan- ning with 3d-vision language pre-training for end-to-end autonomous driving,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.808574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.950415Z digest=sha256:8df57758aee440775657c75b07f6e28c293e63f2bb1f36461e9a0e684333dc0c

Observation 78706104-26be-41cf-a195-1bdadec92af9 · outbound

This paper cites World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:53:10.287648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.955286Z digest=sha256:d1d76ade836cb622852cb56406bd2ae199b341ce788f5cfcd18126a76908b303

Observation 70d8969c-ee11-425b-a5c5-65bb897228a4 · outbound

This paper cites LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:53:10.269257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.960199Z digest=sha256:88c9ae8019f370f5cd92b383b561bc015917e71be9e1cfc1c5a0d71c763c9b36

Observation 451ac5ed-dfca-4ea8-828a-f01a0dca9c0d · outbound

This paper cites Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.797193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.963841Z digest=sha256:178c93bc879fd01a003b1a77085450f85168a3a8ac01aaa0fa87ab9ad8adf1af

Observation 34407945-3071-440e-b026-cf31d84494f1 · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.967292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.967292Z digest=sha256:31f3091ada585ddae607509194c6c809d12f6f220a3aafa740755e073323ae78

Observation b970df43-224a-4f28-a776-41ddd56c168c · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lmdrive: Closed-loop end-to-end driving with large language models,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.785869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.970639Z digest=sha256:31e87a355d91c6041cbef5335d266248cc75ebd8396e5e4e738a06e4903723fc

Observation 337b915b-67f3-424f-92a9-99247d4290c6 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.974080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.974080Z digest=sha256:58b593bf11f528812d6b3a5b2072df92f64bc99578dcfac85d200bb3fe82b26c

Observation 4663e554-d863-45b0-9644-94e37527db6d · outbound

This paper cites Graph-detr4d: Spatio-temporal graph modeling for multi- view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Graph-detr4d: Spatio-temporal graph modeling for multi- view 3d object detection,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.978067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.978067Z digest=sha256:8d8277d5f00a5fed83f3f0965ac89652a73b2aa6db5a736f7c1f5904b8cc3210

Observation d988ba15-8d25-4f99-ac2c-019b819fbe97 · outbound

This paper cites Physically realizable adversarial creating attack against vision-based bev space 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Physically realizable adversarial creating attack against vision-based bev space 3d object detection,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.768257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.981826Z digest=sha256:902f65ffc9d9b8a94e6c4ca07bfb60d50c79c2f1405bfa890f7a4e6ca3e34a1e

Observation af0a7518-2426-4606-a6e0-fcd1212d1553 · outbound

This paper cites Notice of violation of ieee publication principles: Recent advances in 3d object detection in the era of deep neural networks: A survey,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Notice of violation of ieee publication principles: Recent advances in 3d object detection in the era of deep neural networks: A survey,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.756585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.985380Z digest=sha256:a7648f116987b63c808c0b2fe2cf669574316cef5959669aa2673c9b0c1bb108

Observation 36daa477-5ef5-4f81-9a69-6658308eb72a · outbound

This paper cites Obmo: One bounding box multiple objects for monocular 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Obmo: One bounding box multiple objects for monocular 3d object detection,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.744729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.988903Z digest=sha256:b044f228ed668a863881d7ec656413331b9f68ccb4a97544fd77e50fe1772ff1

Observation a91342ba-febe-4693-80fa-49e53d5a7757 · outbound

This paper cites Stereoscopic vision recalling memory for monocular 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Stereoscopic vision recalling memory for monocular 3d object detection,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.733383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.992940Z digest=sha256:96f6b30a6fe9a1fd5e89ee4842ffe18cd4b180b97d8e4b2a3285bfd023b9f4ec

Observation 643cd81b-df84-4ca9-aa99-79e395cff86f · outbound

This paper cites X-view: Non-egocentric multi-view 3d object detector,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving X-view: Non-egocentric multi-view 3d object detector,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.722634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:09.996622Z digest=sha256:4d83299789bdefb3522a2fd0a93a23c5fc726d7f0e1e194568e379c993a51cc0

Observation fcfde415-e9c7-4027-abf6-73bcda875309 · outbound

This paper cites Do-sa&r: Distant object augmented set abstraction and regression for point-based 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Do-sa&r: Distant object augmented set abstraction and regression for point-based 3d object detection,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.712088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.000788Z digest=sha256:147f3a9b68d183993b70147ee27321ce634b7a54dc8dbdf16e5fe7521cc2edb1

Observation 2ed55f55-0093-499b-b0f8-c1f22cb69acb · outbound

This paper cites 3d cascade rcnn: High quality object detection in point clouds,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving 3d cascade rcnn: High quality object detection in point clouds,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.701126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.004869Z digest=sha256:aa90d5258eeb5cb599aba7de343f8415debab4e32f29faad37f42344dd7d8275

Observation 14f4dbbe-42fb-4eba-906d-4882482e3aab · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.690077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.009010Z digest=sha256:f37e1322c29f46d407dcb6319570429b425cdf8385ef91312d1318a3c8f97728

Observation 90da7222-d46c-4883-bc41-1523eb343adc · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.013003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.013003Z digest=sha256:3ea3137200bae852cca9e02485fff14594d12c0cb36d8db500ae0aa589bac019

Observation c9199523-fb97-4b64-b608-9838b3aad4fa · outbound

This paper cites Petrv2: A unified framework for 3d perception from multi-camera images,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petrv2: A unified framework for 3d perception from multi-camera images,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.679295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.017715Z digest=sha256:fa3ad2880e01b8e086f99f6eecd82dc6706712ac7c288caf9c7f4cfb18b3bae8

Observation fb7f6392-75cd-4790-b24d-2cc0897914e9 · outbound

This paper cites Cape: Camera view position embedding for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cape: Camera view position embedding for multi-view 3d object detection,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.667963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.021479Z digest=sha256:01ed15fc3d4fd0c36642e0a0ba5dc60a8a72703d65011d9f38a5a399ef83235b

Observation 503998d9-5ce9-4875-a161-f574d5ed3826 · outbound

This paper cites Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.656426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.025262Z digest=sha256:26fc32ce6a49481db0ad55d58335c611ab38ba279b2410d060253dc61a8004d6

Observation 9f196b30-cefa-46da-8530-3ee3c469f82b · outbound

This paper cites Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.645704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.028772Z digest=sha256:03cdc45eda0ff840f5573d0bd5558bac26e3994e6802cbf2245a65a1ea113c8a

Observation a44a9a7c-1fa1-4556-948d-3fd3fe9ba1a9 · outbound

This paper cites Open: Object-wise position embedding for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Open: Object-wise position embedding for multi-view 3d object detection,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.633351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.032996Z digest=sha256:3c99213296778add1367f146817c8a7040afb511b1e511c07c79886af7abba8e

Observation 586804eb-24b7-4830-8eea-790c4d954b48 · outbound

This paper cites Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.621918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.037169Z digest=sha256:f3801958bb47966e036e6d5cc972d84f946b88e353be5fed801ad346d401a81a

Observation 4d4d5945-1cba-42e6-a12a-42b464cfeb3b · outbound

This paper cites Far3d: Expanding the horizon for surround-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Far3d: Expanding the horizon for surround-view 3d object detection,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.610574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.041250Z digest=sha256:b981bba6ab9602db0c263d331d570d470683bc7dca28a69ba70cfed5c89c621e

Observation 31c094b9-5d34-429f-8a0c-d46f01ac88d0 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.044871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.044871Z digest=sha256:184b090086ebbba7bf8f906c8dd02937b5366be37467fa7e20952a9a259deaa8

Observation eaa0a421-08be-4a19-b4e9-0b8ef187121e · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving nuscenes: A multimodal dataset for autonomous driving,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.048887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.048887Z digest=sha256:e7aa9ef6b37c8c6a85965ce33602a0c66c02c35e2fbd5c3044e55ff51f2c6ad1

Observation 6e332cea-28fc-470f-9fc2-fc05ce598759 · outbound

This paper cites Grit: Faster and better image captioning transformer using dual visual features,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grit: Faster and better image captioning transformer using dual visual features,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.591964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.052761Z digest=sha256:f65e8b504ec1e817bd87b6a02706258213d99b164785651ff376a8a0ba00d523

Observation 7038bcaa-0acd-43ba-8388-8ab3c01b7303 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.056914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.056914Z digest=sha256:58763512e0639396ab230b6a3a3a033d37f8dd0689b261969ca44606eddff865

Observation a5ce548a-ed9b-412f-9b67-fa8a857000bb · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Glamm: Pixel grounding large multimodal model,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.578397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.060557Z digest=sha256:7d2aae7986cb00245fc8c06268221aff58ce2a53c68deac781c1cc18218daba2

Observation 6112512b-2920-4f8f-85a2-90a350a19ca2 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lora: Low-rank adaptation of large language models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.566036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.064117Z digest=sha256:230150f896f2a6ad105314e2e83be54ca7fef8d9d3c2eee18f232323ceededf7

Observation 6fd371b2-07d9-4065-acd9-4430f1899597 · outbound

This paper cites Focal loss for dense object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Focal loss for dense object detection,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.553864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.067413Z digest=sha256:1ba28ef17500a3bea1580f5951ca90169db58e260bae7a4ee29eb8732653709b

Observation f57b63ec-c6e1-4abf-b89d-cc156f3b585d · outbound

This paper cites The hungarian method for the assignment problem,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving The hungarian method for the assignment problem,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.071387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.071387Z digest=sha256:7f5b35d594741508e92f953fed73015d518b14a12bf274d9a0b20bd1db518a33

Observation a83705ec-c02e-4de1-ac68-1ebdfa929b7b · outbound

This paper cites Decoupled Weight Decay Regularization.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Decoupled Weight Decay Regularization

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.074870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.074870Z digest=sha256:c7e00b93b20da55d43330acdd66f2b2edf7e765064022a7e342eaef27e3006dd

Observation 3ef449d7-7ebf-4564-a457-296bffa66e82 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cider: Consensus- based image description evaluation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.534683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:53:10.078556Z digest=sha256:1d841b150ed46387ecc740f1fd46c06f20f287f2ca3e8ef6817bd4de376199c7

Pith citing papers

Observation 155de858-f381-424b-992f-64d33929750e · inbound

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation cites this paper.

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:26.525113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:26.525113Z digest=sha256:a8f41622a8fbaa230770318e149f9afd0f0d5edcc8723975479e8236fe8a641d

Observation e782e27f-1a09-42ab-bbfc-eb2796f514da · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.895175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.895175Z digest=sha256:e397fb7fbc07abeb8184a6672456dd37726e5da451996133dff3ed959e66745e

Observation 1e6634ed-9689-4ee4-83a8-c1f5137017f3 · inbound

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning cites this paper.

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:29.207219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:29.207219Z digest=sha256:0a51073f40d4c3705c1af728e55e20fa27cf2321590b91689f86c2eea220429d

Observation a561ba29-651a-49e2-bdc8-6f643e12bc5e · inbound

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation cites this paper.

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:10:25.162426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T17:06:34.398973Z digest=sha256:57c686dd4632adac08ae41a6968b73dd551aed52476d3277cc5932e901677c82

Observation 42e96adf-e3de-4b3a-932b-4faf347be58e · inbound

An interactive enhanced driving dataset for autonomous driving cites this paper.

An interactive enhanced driving dataset for autonomous driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:19:52.534231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:19:52.534231Z digest=sha256:31364dfc3a3aede9bdb11e17ca97c7e1b0e4ab1bb9cbcd8c01af74e9a8427ce0

Observation 57171dc6-a398-4e4b-ab3f-c89d872d27c7 · inbound

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation cites this paper.

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.079222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T05:31:59.676725Z digest=sha256:fc5776990fab6b1572e3fa877b8bdd84ef68a4c4bd2fb63d4bbeca4a62d77a8d