Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:32:14.153900Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2411.10639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:32:14.153900Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:40.646360Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:26:42.549976Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a286a36e-cfde-4af3-96e6-193e8123dffc · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a97ef46c-5d46-4cbe-8662-c239dd29ab29 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 26319c7d-d4da-4e16-bb57-390c2ef5d34a · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning In- ternLM2 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 392df1de-3ed1-41e6-9cc0-64cb5b00bfe9 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Rehg, and Chao Zheng
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 078278d7-8649-4397-981e-761ba0ffbd8f · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning End-to-end Autonomous Driving: Challenges and Frontiers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2383dca6-332b-4ca5-b93c-5055a0f462aa · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning End-to-End 3D Dense Captioning with V ote2Cap-DETR
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 08695401-f447-4a32-8da9-4c07c9044e04 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning V ote2Cap-DETR++: Decoupling Localization and Describ- ing for End-to-End 3D Dense Captioning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3f141258-fb25-487c-a759-907d35a498bd · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 725fbf58-ff26-4b3f-a44c-e0000d374a51 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eeb2e83e-8c6f-48ed-a4fc-623f43e65fc7 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning ST-P3: End-to-end Vision-based Au- tonomous Driving via Spatial-Temporal Feature Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 639306cd-9d7b-41d7-bcab-0527f4bff043 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Planning-oriented Autonomous Driving
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3d2df9a6-f293-4597-b19c-db9c863ad356 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning EMMA: End-to-End Multimodal Model for Autonomous Driving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e7051d3-6095-4831-8cad-9a075df5a55a · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Le, Yunhsuan Sung, Zhen Li, and Tom Duerig
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6eae8335-299d-47df-8695-6948b558ddc4 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Bench2Drive: Towards Multi-Ability Bench- marking of Closed-Loop End-To-End Autonomous Driving
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b5d55b94-5677-48ac-84af-79cc313a74a4 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning V AD: Vectorized Scene Representation for Efficient Autonomous Driving
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4ef75481-1973-423b-b96c-e103f3384c43 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation add425f1-8a70-43f9-b8e0-5f6a88b087af · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning MaPLe: Multi-modal Prompt Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c973eec2-addd-46ee-b60a-d7c0b968bebd · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Driving Everywhere with Large Language Model Policy Adaptation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8ad6e984-59ac-44a6-b104-7692462b5724 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2b9821-72c2-4293-bcd9-944f4330e743 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFormer: Learning Bird’s-Eye-View Representation from Multi- Camera Images via Spatiotemporal Transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e8b4156f-a2bb-4245-9c97-32dff58c00b6 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning AIDE: An Automatic Data Engine for Object Detection in Au- tonomous Driving
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 112919d5-ea2f-41c4-b779-f12bf4ebf32a · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 05387a8e-eaff-451a-aee2-b99ef078be98 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning ROUGE: A Package for Automatic Evalu- ation of Summaries
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f1320ff-62c4-4e12-8bb3-913be02de2e5 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFusion: Multi- 9 Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a593da1f-e0e1-455b-aa18-400bd765c6bd · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3e59937-63d6-4ee4-aed8-189270d8515c · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Position: Prospective of Autonomous Driving - Multimodal LLMs, World Mod- els, Embodied Intelligence, AI Alignment, and Mamba
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5dc38668-dbe9-43c5-a0e9-f221c782358c · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning DRAMA: Joint Risk Localization and Cap- tioning in Driving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7d5236cf-d935-436f-a806-143825c5f2d1 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning A Language Agent for Autonomous Driving
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 77691cd2-3185-49e6-8cd7-82aec81b005a · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning LingoQA: Video Question Answering for Autonomous Driving
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 10c87793-6202-458d-8091-d7946dc0a88b · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Qi, Runzhou Ge, Kratarth Goel, Zoey Yang, Scott Ettinger, Rami Al-Rfou, Dragomir Anguelov, and Yin Zhou
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 278b70f0-6816-47fb-ba74-170171150e36 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Bleu: a Method for Automatic Evaluation of Machine Translation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation edc1093e-a856-4046-b73b-c4f15788fce1 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning NuScenes-QA: A Multi-Modal Visual Ques- tion Answering Benchmark for Autonomous Driving Sce- nario
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b372ac5-7d9f-4f94-83c3-c87a1c3e63e0 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Learning Transferable Visual Models From Natural Language Supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7515c18-3b16-4d13-b943-91a56a458a50 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning DriveLM: Driving with Graph Visual Question Answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bd17e93-fd3f-43a4-b6bb-1e630851895d · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 11f0f184-521b-4f9e-b466-6313709c4412 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8e0e775d-f927-4a0c-9d8d-f9804ebbcd74 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Lawrence Zitnick, and Devi Parikh
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 84e84d35-fa8d-4d45-8867-05d0ccc8cdd7 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Al- varez
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ea5991d9-253d-4792-80ab-2d8554365293 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning DETR3D: 3D Object De- tection from Multi-view Images via 3D-to-2D Queries
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8705a9cd-bc7f-47c3-a8d0-5ceeef908dab · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 241a726d-53bd-4835-8136-612c80934ebe · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a0d6df6-9dba-4884-aed3-352d29e7d410 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning CoBEVT: Cooperative Bird’s Eye View Semantic Segmentation with Sparse Transformers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4d057410-af34-4e9a-b3ee-fa5e618edfa7 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Wong, Zhenguo Li, and Hengshuang Zhao
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a24f17e-59b8-4e4b-9e27-263125f921d0 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFormer v2: Adapt- ing Modern Image Backbones to Bird’s-Eye-View Recogni- tion via Perspective Supervision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5199b568-f85a-4dc7-a51b-a99d50c2cee9 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 60b98a47-2a22-48fd-8d21-67912726a4f1 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Florence: A New Foundation Model for Computer Vision
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 18ed83e4-f60a-48b2-bcf3-87a291d48b7c · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Cap- tioning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a11d807-7e4b-4a57-9013-ed99e39a5933 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning LLaMA-Adapter: Efficient Fine-tuning of Lan- guage Models with Zero-init Attention
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eedb5c2f-ffc9-4c68-846e-8b7d2887aac0 · outbound
MTA: Multimodal Task Alignment for BEV Perception and Captioning Learning to Prompt for Vision-Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 15cb63b7-9c6a-4f8c-9bbf-8c1e62f2a4a6 · inbound
LTDA-Drive: LLMs-guided Generative Models based Long-tail Data Augmentation for Autonomous Driving MTA: Multimodal Task Alignment for BEV Perception and Captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.