Pith. sign in

Paper Citation Record · LEDGER

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

As of 17 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2507.19370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19370 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:59:05.006971Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T13:44:18.418529Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f036cf7-25b5-4b4d-ba2a-cfe6d0b0efbb · outbound

This paper cites Explaining autonomous driving actions with visual question answering.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Explaining autonomous driving actions with visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.497724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.878331Z digest=sha256:75cccc324e8d1f96d18b4a8e965d26aeb3d367689960e7e7b8a4eac975500604

Observation 7dbcef4f-04c6-49e5-800c-9278c925e901 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.476802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.887726Z digest=sha256:adad35c4c3af415a667d5080265c62c3819c2112b1c7931d7e3a05da590f5a3f

Observation 813b0ede-424b-4f2b-b27e-fc11fd850594 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.460492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.892959Z digest=sha256:aeca23bb3c5c6a0d048dca2366871287548fd162677cc678def67b085c842c6e

Observation aaf05b8e-3dec-48e9-bd2b-f9a1970af63b · outbound

This paper cites The Llama 3 Herd of Models.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:04.897689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:04.897689Z digest=sha256:f0ad3e3757d868f5079398993ed72e9f670eeb5e898db3ab6a71f0d4e61044f0

Observation 1c06142c-ea50-4a50-b1e9-725254abd343 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Lora: Low-rank adaptation of large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.442938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.902562Z digest=sha256:3519bffb32d3d76f16d317dbfe1dac8557bfd403fdbe6b6b0c3ffb7b8defadc6

Observation ef40de80-61eb-4258-b48a-a241b5f16895 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:04.907318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:04.907318Z digest=sha256:5accfa03176c9f5b1fa5984c872c3543294e4c57c33a8e42855b1aa55a4c2ff9

Observation dbbb62e5-1a96-4b7d-a281-717a55472900 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:04.913412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:04.913412Z digest=sha256:56314c4d839a08f37d6635f7021195467a97517eb588d5bed4e4948e7edb4462

Observation 7921eda4-8a38-47a0-91ef-033f2fa295f8 · outbound

This paper cites Adapt: Action-aware driving caption trans- former.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Adapt: Action-aware driving caption trans- former

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.418052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.918298Z digest=sha256:575f863d2f2e4bc4d86e6f9c8229543fbd4ef350cb065055ae54c2d3f8f40741

Observation 0d70e261-5d0f-43ee-bcc6-f9968395b7dc · outbound

This paper cites Tod3cap: Towards 3d dense captioning in outdoor scenes.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Tod3cap: Towards 3d dense captioning in outdoor scenes

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.404009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.923263Z digest=sha256:cf19310ce9814af429a5c75c930448f3a43e856b13d045b34294218b60190ca6

Observation 11025f80-c5d5-423b-95db-5c17b5037c68 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.390739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.927636Z digest=sha256:649a3a5ea5b0411c83ad24728db8ad8d5c6151d90f13cee0216e014ab80ee9cd

Observation ea1c2567-0282-4445-99ac-f70f6a1cfd55 · outbound

This paper cites Improved baselines with visual instruction tuning.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Improved baselines with visual instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.376614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.932234Z digest=sha256:fa02680de6f889f3569de685a6c853662c7536888c9ebc3f98b4a76098744d41

Observation a12bcb04-e76c-4a90-a668-c193cb2340dd · outbound

This paper cites Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.362983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.938761Z digest=sha256:7260ebca33b226df9dcc272ebf5b975e1fe23f00d7a8edd08394c6d36f57821b

Observation 78e54f74-f78f-48d7-95ec-b84f5a434e67 · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Drama: Joint risk localization and captioning in driving

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.349306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.947087Z digest=sha256:42715f6c61fe962bb214c1af1dadaf1d3789b71f3c9fef4403d01d11670bebee

Observation 2fa17cb7-c2bb-4d3a-80fe-b1c834859147 · outbound

This paper cites Image captioning for near-future events from vehicle camera images and motion information.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Image captioning for near-future events from vehicle camera images and motion information

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.332913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.951164Z digest=sha256:049fad8cbaecc2b55d1d82541000b0ba8bb63225f9aad7cf0588449b20cf9830

Observation c20e470c-4ebd-46a7-922b-f0b65c4cae2e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Bleu: a method for automatic evaluation of machine translation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.314809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.955834Z digest=sha256:e19da18dc71eadaf58c90315058b84a05766c272b194353d307aab736a0592c0

Observation 94a6851d-13c2-495b-86e0-b7fca6824231 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Drivelm: Driving with graph visual question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.300455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.960049Z digest=sha256:926bf89faf32b89319950163979b94a7f6571ab5521ffad844b93a8b33ee5132

Observation b8d2aa35-368c-48c4-a04b-d3193d9ad1f5 · outbound

This paper cites Bev-tsr: Text-scene retrieval in bev space for autonomous driving.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Bev-tsr: Text-scene retrieval in bev space for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.286952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.964015Z digest=sha256:ba3ca29f0d4093c70ea7ee9ddcf2177d64678b2d9e40e27e1b090b8014c9eb13

Observation 01220e1d-dfc2-4c93-86ba-d4468b503cbd · outbound

This paper cites Attention is all you need.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Attention is all you need

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.271439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.968432Z digest=sha256:3d6df017da67be9aa398fe187d5f4c29d533ea23a15e31912227c5d60e4cd262

Observation b5fcf54b-2b92-408b-8aac-7b181e77296d · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.256113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.972793Z digest=sha256:591588998e5abbd03353ae5a33b8f212b9fce7fc76ccc7b6d31ce1d065389b61

Observation 5ca3fe39-ac58-498a-bd7c-36291027cd8a · outbound

This paper cites Drivegpt4: In- terpretable end-to-end autonomous driving via large language model.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Drivegpt4: In- terpretable end-to-end autonomous driving via large language model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.241491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.976974Z digest=sha256:1761f117c20b3762011555157918ffc6eaf42ea1465f38a995a7289fc667e136

Observation 37687f01-88b5-435e-8f5a-de101630a4bc · outbound

This paper cites Lidar-llm: Exploring the potential of large language models for 3d lidar understanding.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Lidar-llm: Exploring the potential of large language models for 3d lidar understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.224797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.982560Z digest=sha256:3123476a804398ddd10d868d2819acddd99f16b15aa5c198a43037eb588bcad8

Observation 13fffd37-0574-47b1-80bb-bf527d347642 · outbound

This paper cites Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:04.987813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:04.987813Z digest=sha256:ce0aead5a49bd476275fa1e187b29c6bec7238ef273bdcf768043efb78603d3c

Observation 109a2506-0e51-4546-83e3-6bef805f835e · outbound

This paper cites Tsic-clip: Traffic scene image captioning model based on clip.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Tsic-clip: Traffic scene image captioning model based on clip

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.211119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:04.992585Z digest=sha256:4d9b1bc8cb8b53d68c8708e7063ac4bdd0e95efef02a837c33d6939b25cbc9eb

Observation 762bff50-b218-4dc5-b1af-322e11279379 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:04.996323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:04.996323Z digest=sha256:e3a6f638e7dbb1462b3af7d9e5572c23aa05901a8d6110d0eabf80b366502580

Observation 35770311-6dfd-4b02-bb18-df8a01c4a631 · outbound

This paper cites Bertscore: Evaluating text generation with bert.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving Bertscore: Evaluating text generation with bert

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:59:05.195720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:59:05.000928Z digest=sha256:40995068052a8ce58c18593f630b29548cd3bff75bf6ecbeb695cabc3dba08d7

Observation be8e2f3b-f0eb-4978-8fb3-5090d61dc67d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:59:05.006971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:59:05.006971Z digest=sha256:b2370841ebd856041385527854edd14e3e7736e74d505d85b05c6c3a886ac1ec

Pith citing papers

Observation 4f9c9b50-f0d3-482f-8feb-dc6ad8db4541 · inbound

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations cites this paper.

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T13:44:18.418529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:44:18.418529Z digest=sha256:8525bda7fc423dfc6b1b612c6d3f6b0682e363131effad7b8a9711786d13a739