Pith. sign in

Paper Citation Record · LEDGER

NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2503.12772.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12772 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:52.227293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:37.328933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26b95bed-e2df-41a0-af29-13a9f2c25b97 · inbound

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning cites this paper.

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:46:44.184782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:46:43.955825Z digest=sha256:b36d9593aeba6b71dc2a84400efff9e4c5da8c923916085122a5f7df6f7fb7e5

Observation 8da40cf6-3837-4f95-af12-b276475c9050 · inbound

Understanding Driving Risks using Large Language Models: Toward Elderly Driver Assessment cites this paper.

Understanding Driving Risks using Large Language Models: Toward Elderly Driver Assessment NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:55.861011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:55.861011Z digest=sha256:4263fa8ebe2b450af10e49b8682df0e211c80dfc2534608889e40e001fd26df7

Observation 45f13c46-d0bc-415c-9fb2-f32c2a34fb0d · inbound

Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding cites this paper.

Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:27.966664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:03:27.966664Z digest=sha256:d0363c1a2ac6dd9c320da3854060ef76ea89fb152683fbe1e74ec595e1a6396c

Observation 2f733e9a-c35b-4f28-862c-487477fa3dfc · inbound

NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving cites this paper.

NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:46:24.173753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T12:44:24.574082Z digest=sha256:fbf76e70b5a139aec9e9fe036e6b7a5af86279b75aff9e2aa4dc40ccb89990ca

Observation b8263c0b-38fc-4a72-aea9-e32ef8f6514c · inbound

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios cites this paper.

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:52:03.162024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:52:03.162024Z digest=sha256:260cac71f1723a5d7230714aed44152e2bbc196c3206d372ab6c7599c9c33ec0

Observation 911d829a-d7de-4a33-a7f6-5cdc442c59c9 · inbound

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model cites this paper.

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:10.915896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:34:29.624029Z digest=sha256:41c25b571450f8da511dc4b1b74c77aab8bea0ce1b9f26331b2fa5d06d3c48a6

Observation 451577da-0761-4d38-9627-505ed2c7f9b4 · inbound

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving cites this paper.

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.122346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:13:37.421188Z digest=sha256:bbdbecb8df66030f54c40a227144cc934614f4557dc5131defe5b1685c0a6ae3

Observation 0d90753c-b061-4758-b82c-be26d173c739 · inbound

FleetAgent: Teleoperation Assistant for Autonomous Fleets via Vectorized V2N Messages cites this paper.

FleetAgent: Teleoperation Assistant for Autonomous Fleets via Vectorized V2N Messages NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:37.331858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:14:23.976821Z digest=sha256:c6217e239af5abe67ae9105fd6f95219720b338c54c4ea8056c7ca821118c043

Observation 315823d6-7348-4ab6-a7dd-49eadfcfdeb5 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 168

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:557e492c14d7b9bdadc3c27015f164cbe93baf947f7c89636970e61ff17cc808

Observation dda87fdc-9093-4ff6-b53a-33cc0d1666f9 · inbound

Vision-Language Assistant for Emotional Reactions to Risky Driving cites this paper.

Vision-Language Assistant for Emotional Reactions to Risky Driving NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:10:55.554549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:10:55.554549Z digest=sha256:148917859f47ffa47d64015d6f20d92214eec2d21f71d6446c0d224487cecade

Observation 5ac5ac03-794d-49fc-a791-4158bcbd2c00 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:52.227293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:52.227293Z digest=sha256:47d7726988a95b9dcd584106258fde8a4ba49169163f933a15c7fe11c8411dbf