Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:53:10.078556Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 6 inbound Pith citation observations for arXiv:2505.08725.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:53:10.078556Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:26.525113Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T17:10:25.159616Z
100 of 103 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation af0e2db5-97df-48e6-ba34-eef2b8fb4744 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visual instruction tuning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b1dff0-eea1-4800-95eb-7dd44a64bac4 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87bd67b-a921-4263-8acb-e453cf170b50 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15eb341e-5305-44cb-8388-04eba5b584f1 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef1f138-a931-4f50-93bd-25247ad7119c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f29d88-3a74-43d1-bdc2-cf4ea0137144 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a200733a-a8c2-4235-be19-f139feb09e1d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gemini: A Family of Highly Capable Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ffcde34-759e-40dc-b65a-8b9352e44612 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241fb771-5cb1-4263-a7f6-dce7ab92e8d5 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving InternLM2 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2a629a-9792-4350-b954-26c79c47f028 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA: Open and Efficient Foundation Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4472e26-9493-49ce-8ba0-a4559cdb6f64 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150e5978-133c-46b3-901e-938a75c08f7c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da94fea8-9b4a-4abd-a4ed-6ac2c33d19fd · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a885fc-63c3-4161-8387-c7ea46159ffd · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Embodied understanding of driving scenarios,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c513d5-8c01-4cc3-b493-8013c377c597 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2f36b6-edc5-45b5-8277-056f7ea8f29c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivevlm: The convergence of autonomous driving and large vision-language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c4f03a-4367-47b4-bd00-9f2f6084ae81 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1266064-9caa-47b3-af39-d24e24de19d3 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80804ced-713e-4aeb-bfe1-5f89af80d114 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 057b03da-0b8b-4083-97b3-9f1a6283553d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lingoqa: Visual question answering for autonomous driving,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a35f8a-2bbb-4e2e-acbd-042353accd74 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Automated evaluation of large vision-language models on self-driving corner cases,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a66bee5-b1e7-4988-905e-38537054d902 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e553ca0-123d-4a98-a893-c797c35e4a75 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivelm: Driving with graph visual question answering,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc6e766-13cd-4d68-81b5-1b6ea6c962a3 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language prompt for autonomous driving,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ee3faa-f307-44e3-9b6c-d28be498676b · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language-Image Models with 3D Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d731add-91c8-40da-af82-0540c830deeb · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c816d584-4c43-467a-938d-9a22b791c108 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2car: Taking control of your self-driving car,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc639249-fb3a-406e-bdc0-a2413c90e37a · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab54154-d3f1-4af2-8559-f7a4be381719 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Improved baselines with visual instruction tuning,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a236f099-dca0-40ab-ade2-e099f95af1b5 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b77ab0f0-04c0-49da-9d91-68a022e90d8e · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c44b41-cc05-47fd-8529-3f803263c0e4 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Pix2seq: A language modeling framework for object detection,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509210da-73bd-4da1-97e3-0c23bb99be10 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petr: Position embedding transformation for multi-view 3d object detection,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ce9192-cab8-4dea-84d2-aec9baa3765b · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ec23d1-b54e-4a2f-a360-7286e37d526f · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18875db-46df-4a97-8b72-06eda01de60c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounding human-to-vehicle advice for self-driving vehicles,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1f4411-3818-4d74-b3eb-0d99dd8081e6 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c503aa-8639-4c2a-a326-8e878e824e3e · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drama: Joint risk localization and captioning in driving,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4099c1-778f-4a97-bf85-ad9358af885d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076f3704-9b96-4ee8-ad67-a00a728329a6 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c525c83-1dde-4239-b171-898e93d02cdc · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0351992-9a9b-40fc-a141-ae809611e947 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 72584531-c02e-4537-af12-46682891f22e · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llava- next: Improved reasoning, ocr, and world knowledge,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15b41a1e-7211-4ac7-856c-dfebccaf96f4 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Learning transferable visual models from natural language supervision,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62406d58-a7b0-40aa-b1d7-0805bdf5e0b1 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Sigmoid loss for language image pre-training,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 569009a9-a579-4d5d-99f5-95ba72d36488 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Dinov2: Learning robust visual features without supervision,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e91f92a-2861-4361-a337-d7a9656543d7 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vision-language models for vision tasks: A survey,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecec2b5d-db96-4866-afa2-18be6461c90a · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Monkey: Image resolution and text label are important things for large multi-modal models,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9e14063f-7254-4a27-a907-eb4952f5b386 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Uni-moe: Scaling unified multimodal llms with mixture of experts,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7a730da-cb98-432f-a98d-07c5897e15e7 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f16af07-e6cc-4f25-b1e9-18f1ae6e4f7d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret: Refer and ground anything anywhere at any granularity,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96470717-e76d-424b-bbc2-d0e1043c5001 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret-v2: An improved baseline for referring and grounding with large language models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation daad5a10-c313-49e9-a894-0087c9786e26 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Relationlmm: Large multimodal model as open and versatile visual relationship generalist,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0cf9eba9-e938-443c-9dbd-99c564403b34 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Groma: Localized visual tokenization for grounding multimodal large language models,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40a5d163-4b92-494d-9f53-009a63553b0c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Segment anything,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 513d95c0-5368-4b54-845a-297bb16acb97 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Masked-attention mask transformer for universal image segmenta- tion,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8890a925-f7ab-44b7-8f6b-fe1c213b118e · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lisa: Reasoning segmentation via large language model,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7293f22d-464a-45d6-8aa3-e02dc3fedd46 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Psalm: Pixelwise segmentation with large multi-modal model,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95e58821-a84a-4ab3-a11c-13d0928aae29 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 355d12dc-ed99-4b68-89bb-ebbfffa8cd80 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08b527f2-7726-4b34-bb2f-561516633621 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Tod3cap: Towards 3d dense captioning in outdoor scenes,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 56bc0e69-faaa-47f5-92dc-7ef0993400a4 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeeb4bff-6e2e-4bb1-a37f-74f3fd7af46d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a968a4f6-892f-47ca-8685-b032c9d935de · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gpt-driver: Learning to drive with gpt,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation badf7126-f77f-42d7-9abf-e50ca4687240 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Making large language models better planners with reasoning-decision alignment,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 120ab713-3f78-47f2-ab94-13d325e1a21d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Distilling multi- modal large language models for autonomous driving,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c75dc518-574d-49aa-b789-f44dc378b1e9 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Generative plan- ning with 3d-vision language pre-training for end-to-end autonomous driving,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78706104-26be-41cf-a195-1bdadec92af9 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70d8969c-ee11-425b-a5c5-65bb897228a4 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 451ac5ed-dfca-4ea8-828a-f01a0dca9c0d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 34407945-3071-440e-b026-cf31d84494f1 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b970df43-224a-4f28-a776-41ddd56c168c · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lmdrive: Closed-loop end-to-end driving with large language models,
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 337b915b-67f3-424f-92a9-99247d4290c6 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4663e554-d863-45b0-9644-94e37527db6d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Graph-detr4d: Spatio-temporal graph modeling for multi- view 3d object detection,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d988ba15-8d25-4f99-ac2c-019b819fbe97 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Physically realizable adversarial creating attack against vision-based bev space 3d object detection,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af0a7518-2426-4606-a6e0-fcd1212d1553 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Notice of violation of ieee publication principles: Recent advances in 3d object detection in the era of deep neural networks: A survey,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 36daa477-5ef5-4f81-9a69-6658308eb72a · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Obmo: One bounding box multiple objects for monocular 3d object detection,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a91342ba-febe-4693-80fa-49e53d5a7757 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Stereoscopic vision recalling memory for monocular 3d object detection,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 643cd81b-df84-4ca9-aa99-79e395cff86f · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving X-view: Non-egocentric multi-view 3d object detector,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fcfde415-e9c7-4027-abf6-73bcda875309 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Do-sa&r: Distant object augmented set abstraction and regression for point-based 3d object detection,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2ed55f55-0093-499b-b0f8-c1f22cb69acb · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving 3d cascade rcnn: High quality object detection in point clouds,
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14f4dbbe-42fb-4eba-906d-4882482e3aab · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90da7222-d46c-4883-bc41-1523eb343adc · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9199523-fb97-4b64-b608-9838b3aad4fa · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petrv2: A unified framework for 3d perception from multi-camera images,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fb7f6392-75cd-4790-b24d-2cc0897914e9 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cape: Camera view position embedding for multi-view 3d object detection,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 503998d9-5ce9-4875-a161-f574d5ed3826 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f196b30-cefa-46da-8530-3ee3c469f82b · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a44a9a7c-1fa1-4556-948d-3fd3fe9ba1a9 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Open: Object-wise position embedding for multi-view 3d object detection,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 586804eb-24b7-4830-8eea-790c4d954b48 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4d4d5945-1cba-42e6-a12a-42b464cfeb3b · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Far3d: Expanding the horizon for surround-view 3d object detection,
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 31c094b9-5d34-429f-8a0c-d46f01ac88d0 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa0a421-08be-4a19-b4e9-0b8ef187121e · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving nuscenes: A multimodal dataset for autonomous driving,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e332cea-28fc-470f-9fc2-fc05ce598759 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grit: Faster and better image captioning transformer using dual visual features,
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7038bcaa-0acd-43ba-8388-8ab3c01b7303 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ce548a-ed9b-412f-9b67-fa8a857000bb · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Glamm: Pixel grounding large multimodal model,
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6112512b-2920-4f8f-85a2-90a350a19ca2 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lora: Low-rank adaptation of large language models,
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6fd371b2-07d9-4065-acd9-4430f1899597 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Focal loss for dense object detection,
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f57b63ec-c6e1-4abf-b89d-cc156f3b585d · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving The hungarian method for the assignment problem,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83705ec-c02e-4de1-ac68-1ebdfa929b7b · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Decoupled Weight Decay Regularization
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef449d7-7ebf-4564-a457-296bffa66e82 · outbound
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cider: Consensus- based image description evaluation,
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 155de858-f381-424b-992f-64d33929750e · inbound
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e782e27f-1a09-42ab-bbfc-eb2796f514da · inbound
A Survey on Vision-Language-Action Models for Autonomous Driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 160
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6634ed-9689-4ee4-83a8-c1f5137017f3 · inbound
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a561ba29-651a-49e2-bdc8-6f643e12bc5e · inbound
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 42e96adf-e3de-4b3a-932b-4faf347be58e · inbound
An interactive enhanced driving dataset for autonomous driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57171dc6-a398-4e4b-ab3f-c89d872d27c7 · inbound
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.