Pith. sign in

Paper Citation Record · LEDGER

MTA: Multimodal Task Alignment for BEV Perception and Captioning

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2411.10639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10639 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:32:14.153900Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:40.646360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:26:42.549976Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a286a36e-cfde-4af3-96e6-193e8123dffc · outbound

This paper cites METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments.

MTA: Multimodal Task Alignment for BEV Perception and Captioning METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.678020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:13.996470Z digest=sha256:d7b4ab73be9475a32b8e0af3bfc28bf2edef7d98f02272d6a02ce48522b488fa

Observation a97ef46c-5d46-4cbe-8662-c239dd29ab29 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.668004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.000851Z digest=sha256:a4c3caf114f1bfca1b03f6af5ffcb6616282b7548a5ebd9214cedcc5892c0927

Observation 26319c7d-d4da-4e16-bb57-390c2ef5d34a · outbound

This paper cites In- ternLM2 Technical Report.

MTA: Multimodal Task Alignment for BEV Perception and Captioning In- ternLM2 Technical Report

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.657779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.005106Z digest=sha256:06957ffbdbefb5f0585f4bb3662cb3d85707fc98ddad183d7b3d4f713c327dc6

Observation 392df1de-3ed1-41e6-9cc0-64cb5b00bfe9 · outbound

This paper cites Rehg, and Chao Zheng.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Rehg, and Chao Zheng

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.646923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.008256Z digest=sha256:17f43e56e40ee1ae3553ac796fc1cf0e6fe1eb762662ae835ec05c1d2b463381

Observation 078278d7-8649-4397-981e-761ba0ffbd8f · outbound

This paper cites End-to-end Autonomous Driving: Challenges and Frontiers.

MTA: Multimodal Task Alignment for BEV Perception and Captioning End-to-end Autonomous Driving: Challenges and Frontiers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.636171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.011154Z digest=sha256:0fdded02cf0fd2c4034b06fb222a31d4e9ee075c6eb4d4fb718b751c06dc6064

Observation 2383dca6-332b-4ca5-b93c-5055a0f462aa · outbound

This paper cites End-to-End 3D Dense Captioning with V ote2Cap-DETR.

MTA: Multimodal Task Alignment for BEV Perception and Captioning End-to-End 3D Dense Captioning with V ote2Cap-DETR

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.626594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.014244Z digest=sha256:ed41d7baeff6b08a766116fdec93260584b62d964e637562368e0b7f507a0434

Observation 08695401-f447-4a32-8da9-4c07c9044e04 · outbound

This paper cites V ote2Cap-DETR++: Decoupling Localization and Describ- ing for End-to-End 3D Dense Captioning.

MTA: Multimodal Task Alignment for BEV Perception and Captioning V ote2Cap-DETR++: Decoupling Localization and Describ- ing for End-to-End 3D Dense Captioning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.615883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.017253Z digest=sha256:1806c0329e5aa6fb1859a8300ba7db4cded37d25af669951ca498a1f0e8a42bb

Observation 3f141258-fb25-487c-a759-907d35a498bd · outbound

This paper cites an unresolved cited work.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:32:14.605993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.019838Z digest=sha256:64efe6a35f0971cb69268e5229b7a1974f3e11828296fd2b51fbfddfcc843e8a

Observation 725fbf58-ff26-4b3f-a44c-e0000d374a51 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MTA: Multimodal Task Alignment for BEV Perception and Captioning InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.595277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.022399Z digest=sha256:6d014c2b7ea1d5d0e70494b8e8b6fde533fc91fbdca7051825cf34a4ed3cfea8

Observation eeb2e83e-8c6f-48ed-a4fc-623f43e65fc7 · outbound

This paper cites ST-P3: End-to-end Vision-based Au- tonomous Driving via Spatial-Temporal Feature Learning.

MTA: Multimodal Task Alignment for BEV Perception and Captioning ST-P3: End-to-end Vision-based Au- tonomous Driving via Spatial-Temporal Feature Learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.585277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.025125Z digest=sha256:3a185092639b7935c86e7ec7afafd42771bb33145dc5d584e32a4125ffcfb9ca

Observation 639306cd-9d7b-41d7-bcab-0527f4bff043 · outbound

This paper cites Planning-oriented Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Planning-oriented Autonomous Driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.573135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.028002Z digest=sha256:31263d0602a602b4c00fb4c1900c3aa17bec047766569167f3b8f2a181304055

Observation 3d2df9a6-f293-4597-b19c-db9c863ad356 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.561424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.030641Z digest=sha256:7cab90ecef6bac993691dd37797eafd4e72883c10209903632cb3abe99fe1140

Observation 3e7051d3-6095-4831-8cad-9a075df5a55a · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.548734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.033178Z digest=sha256:93326f05c4e0ebd9bc2cffb56aa6cb4e2c9abd2d28b8328b0968600a6c1fa415

Observation 6eae8335-299d-47df-8695-6948b558ddc4 · outbound

This paper cites Bench2Drive: Towards Multi-Ability Bench- marking of Closed-Loop End-To-End Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Bench2Drive: Towards Multi-Ability Bench- marking of Closed-Loop End-To-End Autonomous Driving

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.536434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.035757Z digest=sha256:54a9e49c914434495dfdeea3312a48fef729ea5490392110d30f4ff3c61e3c7e

Observation b5d55b94-5677-48ac-84af-79cc313a74a4 · outbound

This paper cites V AD: Vectorized Scene Representation for Efficient Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning V AD: Vectorized Scene Representation for Efficient Autonomous Driving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.524840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.038527Z digest=sha256:cc16817b41dba213da6fd1f7398f3d342748d2a586899722149bbd2556a0bbd5

Observation 4ef75481-1973-423b-b96c-e103f3384c43 · outbound

This paper cites TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes.

MTA: Multimodal Task Alignment for BEV Perception and Captioning TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.513422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.042488Z digest=sha256:24ddabe91f2b39b631188a8ca40f59bee0831f3d1f0c0fcb3f010df2ae485be3

Observation add425f1-8a70-43f9-b8e0-5f6a88b087af · outbound

This paper cites MaPLe: Multi-modal Prompt Learning.

MTA: Multimodal Task Alignment for BEV Perception and Captioning MaPLe: Multi-modal Prompt Learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.501930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.046547Z digest=sha256:78cbb8e42c347193cf046728033141d569bb8fc5227ea01058e3365001c8dc44

Observation c973eec2-addd-46ee-b60a-d7c0b968bebd · outbound

This paper cites Driving Everywhere with Large Language Model Policy Adaptation.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Driving Everywhere with Large Language Model Policy Adaptation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.490592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.050336Z digest=sha256:0596ccad36fb8b6ed7270d55488944379566ca282aa218997d16d77cb98ab52c

Observation 8ad6e984-59ac-44a6-b104-7692462b5724 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:32:14.054228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:32:14.054228Z digest=sha256:cc1afd5bc4e00a24557f5ed95be21321737a7f1a03f0a0982c62826063db3fe3

Observation 9a2b9821-72c2-4293-bcd9-944f4330e743 · outbound

This paper cites BEVFormer: Learning Bird’s-Eye-View Representation from Multi- Camera Images via Spatiotemporal Transformers.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFormer: Learning Bird’s-Eye-View Representation from Multi- Camera Images via Spatiotemporal Transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.473009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.057382Z digest=sha256:8eadb1fbcd2f3db3ebe991c21e2d4b7318ceb5aa699067982d5c5adf9a881081

Observation e8b4156f-a2bb-4245-9c97-32dff58c00b6 · outbound

This paper cites AIDE: An Automatic Data Engine for Object Detection in Au- tonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning AIDE: An Automatic Data Engine for Object Detection in Au- tonomous Driving

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.464892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.060336Z digest=sha256:8d5b136721f2d8c73a3587a1f1412e48111bad2010d761bfb4ae366fa839a0ed

Observation 112919d5-ea2f-41c4-b779-f12bf4ebf32a · outbound

This paper cites BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.455587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.063709Z digest=sha256:58eff003e579a7e789aa205d17e4b3d24dadc825b2455df3592c91a1e1161cc2

Observation 05387a8e-eaff-451a-aee2-b99ef078be98 · outbound

This paper cites ROUGE: A Package for Automatic Evalu- ation of Summaries.

MTA: Multimodal Task Alignment for BEV Perception and Captioning ROUGE: A Package for Automatic Evalu- ation of Summaries

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.446019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.066998Z digest=sha256:3f2c5e9f389c406a9d2660fbb2d919e39c601e394ccc89f5a2c4dbab37c822dc

Observation 8f1320ff-62c4-4e12-8bb3-913be02de2e5 · outbound

This paper cites BEVFusion: Multi- 9 Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFusion: Multi- 9 Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.435380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.070146Z digest=sha256:6bbc9e3517dbf999e692c044ce85a63758d6310720e16d1502d731a826c18398

Observation a593da1f-e0e1-455b-aa18-400bd765c6bd · outbound

This paper cites The Llama 3 Herd of Models.

MTA: Multimodal Task Alignment for BEV Perception and Captioning The Llama 3 Herd of Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.424061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.072938Z digest=sha256:007f34254745dffedb43f5229b5596ce4879b70e345cb2cc0c7e092b88a36d7a

Observation e3e59937-63d6-4ee4-aed8-189270d8515c · outbound

This paper cites Position: Prospective of Autonomous Driving - Multimodal LLMs, World Mod- els, Embodied Intelligence, AI Alignment, and Mamba.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Position: Prospective of Autonomous Driving - Multimodal LLMs, World Mod- els, Embodied Intelligence, AI Alignment, and Mamba

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.414657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.075614Z digest=sha256:8b443a414cad25f0114838ab62b8fc3e5a61dd4691e244accd3a82e12e1200cb

Observation 5dc38668-dbe9-43c5-a0e9-f221c782358c · outbound

This paper cites DRAMA: Joint Risk Localization and Cap- tioning in Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning DRAMA: Joint Risk Localization and Cap- tioning in Driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.402369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.078428Z digest=sha256:ffc90ea84a79172c619f6854e4062b6dd9438dcfe31a6d67d8ef34687800a909

Observation 7d5236cf-d935-436f-a806-143825c5f2d1 · outbound

This paper cites A Language Agent for Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning A Language Agent for Autonomous Driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.390258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.081995Z digest=sha256:e9e1d9e9c54ac46fcbbcbd926d2a9a8c60b823c9de1ad87dd0d7b62d3ba7478e

Observation 77691cd2-3185-49e6-8cd7-82aec81b005a · outbound

This paper cites LingoQA: Video Question Answering for Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning LingoQA: Video Question Answering for Autonomous Driving

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.380841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.084838Z digest=sha256:17452b7ccfb4a139d866cb6709b9bb55f8f901560ca18a1135671566042a588b

Observation 10c87793-6202-458d-8091-d7946dc0a88b · outbound

This paper cites Qi, Runzhou Ge, Kratarth Goel, Zoey Yang, Scott Ettinger, Rami Al-Rfou, Dragomir Anguelov, and Yin Zhou.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Qi, Runzhou Ge, Kratarth Goel, Zoey Yang, Scott Ettinger, Rami Al-Rfou, Dragomir Anguelov, and Yin Zhou

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.369688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.088646Z digest=sha256:65fe94a507c5a519ca4ccef8cd79cd3da0629b028709c8cf92e5224a190da216

Observation 278b70f0-6816-47fb-ba74-170171150e36 · outbound

This paper cites Bleu: a Method for Automatic Evaluation of Machine Translation.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Bleu: a Method for Automatic Evaluation of Machine Translation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.359513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.092525Z digest=sha256:38ec7744091decc46aa2c54c9bc52697b5d06d132e4eb6f4392dae3160be4385

Observation edc1093e-a856-4046-b73b-c4f15788fce1 · outbound

This paper cites NuScenes-QA: A Multi-Modal Visual Ques- tion Answering Benchmark for Autonomous Driving Sce- nario.

MTA: Multimodal Task Alignment for BEV Perception and Captioning NuScenes-QA: A Multi-Modal Visual Ques- tion Answering Benchmark for Autonomous Driving Sce- nario

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.350818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.096584Z digest=sha256:435c002041c487285125a902fb25b6315e691f18eaa736a2b0d1bf910f3bb343

Observation 3b372ac5-7d9f-4f94-83c3-c87a1c3e63e0 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:32:14.099810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:32:14.099810Z digest=sha256:8b8316b65fd420ca115a71d11f48f222aa16d043c61683965bf2bf04666e75e8

Observation b7515c18-3b16-4d13-b943-91a56a458a50 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

MTA: Multimodal Task Alignment for BEV Perception and Captioning DriveLM: Driving with Graph Visual Question Answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.332725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.103303Z digest=sha256:c5afbf610107562bf23f8a1605ec3e095337053785f6cc1333fb34deaf6ffede

Observation 5bd17e93-fd3f-43a4-b6bb-1e630851895d · outbound

This paper cites Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.323116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.106140Z digest=sha256:c1baed57d0547b61c8c365f649db3652bf1386b7d7c2de07c567b1be1f20d311

Observation 11f0f184-521b-4f9e-b466-6313709c4412 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

MTA: Multimodal Task Alignment for BEV Perception and Captioning DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.314317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.109436Z digest=sha256:22acc8f0c11f5fb7db197cd4428a46069525bed9f58c1e2828f84dfade545c0b

Observation 8e0e775d-f927-4a0c-9d8d-f9804ebbcd74 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Lawrence Zitnick, and Devi Parikh

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.306055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.112332Z digest=sha256:c104ff7a66841fbe3f2e2f29ada7666e27f5c2269975b87ef561106bcfd5a02b

Observation 84e84d35-fa8d-4d45-8867-05d0ccc8cdd7 · outbound

This paper cites Al- varez.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Al- varez

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.297987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.115323Z digest=sha256:787f71f3a4c50a4f3d4db20499593cf0c1a8e395ad70a2dad5dfd5d0304beaa7

Observation ea5991d9-253d-4792-80ab-2d8554365293 · outbound

This paper cites DETR3D: 3D Object De- tection from Multi-view Images via 3D-to-2D Queries.

MTA: Multimodal Task Alignment for BEV Perception and Captioning DETR3D: 3D Object De- tection from Multi-view Images via 3D-to-2D Queries

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.288727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.119389Z digest=sha256:e0a93cabae4ece9f98bc1ac956ecb63311def145a0e64acc43096b790a57f1f7

Observation 8705a9cd-bc7f-47c3-a8d0-5ceeef908dab · outbound

This paper cites PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving.

MTA: Multimodal Task Alignment for BEV Perception and Captioning PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.279173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.122040Z digest=sha256:15b304e23a4c47fb1c78dd3e2d3d43c4f571da563b1b2754ac9d8e4cdbd51a10

Observation 241a726d-53bd-4835-8136-612c80934ebe · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.270033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.125088Z digest=sha256:427728e619fd2d1f640674f8a542658c66c80bc594c0b4a0049dec44a9ffa77d

Observation 3a0d6df6-9dba-4884-aed3-352d29e7d410 · outbound

This paper cites CoBEVT: Cooperative Bird’s Eye View Semantic Segmentation with Sparse Transformers.

MTA: Multimodal Task Alignment for BEV Perception and Captioning CoBEVT: Cooperative Bird’s Eye View Semantic Segmentation with Sparse Transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.260408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.128533Z digest=sha256:9a0178f9669a3c0a64963798a2c4811cd465161361e49f9c9e0f5e8561132d29

Observation 4d057410-af34-4e9a-b3ee-fa5e618edfa7 · outbound

This paper cites Wong, Zhenguo Li, and Hengshuang Zhao.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Wong, Zhenguo Li, and Hengshuang Zhao

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.248644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.133172Z digest=sha256:9c1164122bd9b24247156154352b6f9350542cbedcc11c0d8ba18932429b136e

Observation 4a24f17e-59b8-4e4b-9e27-263125f921d0 · outbound

This paper cites BEVFormer v2: Adapt- ing Modern Image Backbones to Bird’s-Eye-View Recogni- tion via Perspective Supervision.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVFormer v2: Adapt- ing Modern Image Backbones to Bird’s-Eye-View Recogni- tion via Perspective Supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.236326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.136525Z digest=sha256:6d74fc9c06be5f62bf700f6d5c7b3a6566c151ec8a874a3ad1b3a0b4e7ce5acf

Observation 5199b568-f85a-4dc7-a51b-a99d50c2cee9 · outbound

This paper cites BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance.

MTA: Multimodal Task Alignment for BEV Perception and Captioning BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.223050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.140134Z digest=sha256:7187cc67e3267b3b5235d13d704dadab4fecd316a848714af47547314ef0345e

Observation 60b98a47-2a22-48fd-8d21-67912726a4f1 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Florence: A New Foundation Model for Computer Vision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.213540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.145025Z digest=sha256:a11e0906c488b86f02b1f898d5aee8817e300332569b9e0156c16deabc9414de

Observation 18ed83e4-f60a-48b2-bcf3-87a291d48b7c · outbound

This paper cites X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Cap- tioning.

MTA: Multimodal Task Alignment for BEV Perception and Captioning X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Cap- tioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.201308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.148248Z digest=sha256:2c67cd7f8ec5a633d0b4ee5abb3bed0bf71446d1953d7f4993972485757f45ec

Observation 3a11d807-7e4b-4a57-9013-ed99e39a5933 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Lan- guage Models with Zero-init Attention.

MTA: Multimodal Task Alignment for BEV Perception and Captioning LLaMA-Adapter: Efficient Fine-tuning of Lan- guage Models with Zero-init Attention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.191412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.151394Z digest=sha256:99c06b4ba9498ec34f78e1cd1da7893d7d23ac6bb0d7e5b0fede979363016b67

Observation eedb5c2f-ffc9-4c68-846e-8b7d2887aac0 · outbound

This paper cites Learning to Prompt for Vision-Language Models.

MTA: Multimodal Task Alignment for BEV Perception and Captioning Learning to Prompt for Vision-Language Models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:32:14.180712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:32:14.153900Z digest=sha256:5a79ba32f22eca2ebe97bee6c5495819b720a42284021c53c1abb463e7e56115

Pith citing papers

Observation 15cb63b7-9c6a-4f8c-9bbf-8c1e62f2a4a6 · inbound

LTDA-Drive: LLMs-guided Generative Models based Long-tail Data Augmentation for Autonomous Driving cites this paper.

LTDA-Drive: LLMs-guided Generative Models based Long-tail Data Augmentation for Autonomous Driving MTA: Multimodal Task Alignment for BEV Perception and Captioning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:26:42.627301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T15:26:40.646360Z digest=sha256:52f4818ded77183b5a257bf4e9070ee8a0f48d559f88d5f761ef1188187d0864