Pith. sign in

Paper Citation Record · LEDGER

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

As of 17 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 6 inbound Pith citation observations for arXiv:2506.06218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06218 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:13.187292Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:08:21.388777Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T05:46:01.450385Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy45
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 578c97c1-444d-4a6e-b1e1-13d8c2717b6f · outbound

This paper cites LLaMA 3.2: Open Foundation and Instruction Models, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving LLaMA 3.2: Open Foundation and Instruction Models, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.323684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.796168Z digest=sha256:041bb8e6b3882789abb76adafdeda84d5c0f0c399ccdc62513bed06e475bb5c9

Observation 63cdeb96-d1e5-4291-a68f-ce8c9d47efe4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.311045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.800668Z digest=sha256:df5aa9c3d1f9e43921768071924981c15c761b30f6fa02f9f3c3cd4a956c1a90

Observation 905ea11b-5021-431c-8a9d-558743278fe3 · outbound

This paper cites CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.298239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.805120Z digest=sha256:8a40715e3ebc8eb7e5b37fc27d0e5ea65ede8ce5556ac2e41df211d21021a1cd

Observation 13c45644-8a12-4bbd-a3a6-e11adc0b4cb0 · outbound

This paper cites Qwen2.5-VL Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.813523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.813523Z digest=sha256:54a14fe2c7f684786b09d657b284a0de1ef5ce93d2e851820a852836bdea28eb

Observation 7392d7c1-043c-49b9-bff4-33f1a3996cf2 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.285747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.817562Z digest=sha256:713bdb2aa3803ee928a915ee0a93986a7a2032e4478b51e2a657e90adf41667c

Observation eaa9f319-698d-4da0-b5ab-8f591cb92aa1 · outbound

This paper cites MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.272732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.821890Z digest=sha256:386f09809023ceb85493b9f97d6e5d0e5b1a8428ae238f2b48a85b3caf65f290

Observation 29e0b0a2-5b83-403e-8b58-c4e0f603a994 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.825529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.825529Z digest=sha256:a46119446aa99d488656909bea39ae7a5f909324aa7ce54d27c880cbcf981aba

Observation be51db96-d4d2-40ce-b9b5-0fbd258ccd95 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.829100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.829100Z digest=sha256:c5927d14e89f6a4db63041905530f8273e6becef26171c0a16458f5c4c0d7848

Observation 75893966-066d-4887-8676-651159a8b5fd · outbound

This paper cites The Cityscapes Dataset for Semantic Urban Scene Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving The Cityscapes Dataset for Semantic Urban Scene Understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.242535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.832601Z digest=sha256:c02623f7a1790ee4c2edad3d4eeea835e2f3e722f2dde4c88a3e044e79f63444

Observation efd415f6-72e2-431c-b563-cf6736bd4c49 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.230572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.836540Z digest=sha256:c1f03e40e05827d981408be2652b859ac21e048bcdf4236e38f8ab7e1a40fa02

Observation 2dc79cff-9d03-4d53-809d-9ff63e8df15b · outbound

This paper cites DeepSeek-V3 Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.840182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.840182Z digest=sha256:99fc2fdfa6758788a0c04e20ead96a864ca1fd2a6af551181c0c8210eca77610

Observation f4de31c4-3919-481f-a0d0-74aaf4f3a52f · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.218383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.844016Z digest=sha256:cdfa3028a92242a9767737e73091b4cb40c272f3a5b55fe2148349971bdb266a

Observation 6def661e-8df5-4e99-bb0b-ff7a47980a23 · outbound

This paper cites Talk2Car: Taking control of your self-driving car.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Talk2Car: Taking control of your self-driving car

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.206159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.847511Z digest=sha256:0ee42bdc68e98282a3491ede5cde68d6dea39cab308a77494843ae3d88e9544b

Observation 88631727-4d17-4ebb-aa1a-f5478a3cf72f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.850966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.850966Z digest=sha256:792464c5c5dab7d0715599875dc52842c1efb5093de2917be9f39d8b7ad55325

Observation e4a84a4e-1f60-4d2c-b8b8-495c0e52d839 · outbound

This paper cites Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.194298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.854590Z digest=sha256:cc19110b82699ca6561d7939da8b2a9edcafa2d78f550e7cfeb2e2a570cce7de

Observation 601b8b08-f511-4c9f-98aa-b2353f4078fd · outbound

This paper cites CARLA: An Open Urban Driving Simulator.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving CARLA: An Open Urban Driving Simulator

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.181740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.858238Z digest=sha256:3bebda0477c900ed815aafc3d9d6bcaafdfe1e73953cfdb6d3360feee5a61319

Observation d40f0edc-972c-4867-aa84-a7a60001d6e0 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.861632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.861632Z digest=sha256:d0aadb0b0bbacf27705a242cf71236062c19f0543501e12213c970570148343b

Observation ab061dd3-fe7b-4ad6-adcb-b025364a38fe · outbound

This paper cites Are we ready for autonomous driving? The KITTI vision benchmark suite.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are we ready for autonomous driving? The KITTI vision benchmark suite

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.169479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.974684Z digest=sha256:8ff62bde038953a925dd0aa876e23f5323c747437a937ae955305367d32ddf54

Observation eebbe239-95f0-44ab-85c8-545ce79f0b2b · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.978819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.978819Z digest=sha256:4efee495b495632c2dfd853510bf5a049ff681cc2a324318b4e360e8562952cd

Observation 9e87d558-b20f-4cef-ac97-7223b62a16b7 · outbound

This paper cites Planning-oriented Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Planning-oriented Autonomous Driving

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.156715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.982852Z digest=sha256:71a96d7037ac71348b7139ff92c908f23ab384d0e5949642a62ad4acc61614df

Observation 83e38495-3cfb-4db7-937a-892a4f4ae3f1 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.986482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.986482Z digest=sha256:7cf67636cb82dd7da6752152ea70b31c12da12bced26947d536017d4dd5163b9

Observation d24a4d13-87c9-47ef-b6dd-2082928fa097 · outbound

This paper cites Making Large Language Models Better Planners with Reasoning-Decision Alignment.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Making Large Language Models Better Planners with Reasoning-Decision Alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.144030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.990385Z digest=sha256:4d0582a5f057c1c7d01490a548bcfa67cd41282746979ce80be287b39f300394

Observation cefcec4e-63ab-43cc-9f01-be84f09718ae · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.993977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.993977Z digest=sha256:fbc716d3d1b437cd6bbbf40997e1152846cffec50ec4473826d280acc67cbdb5

Observation 4dd1bbdd-79ef-4d20-a514-4a992137034b · outbound

This paper cites NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.130940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:12.998184Z digest=sha256:3137053e023d8fdbcc6f45bcaa3824f3d0595c67654f062f4cfc60ea607d3f28

Observation 5fd9b368-d0d3-4e47-881c-78b07fcce794 · outbound

This paper cites DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.001866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.001866Z digest=sha256:15b390e23f104523bd70fbe9f8e13fb11b6769cc3a7b710137fff758b0c56045

Observation ddd1bf57-9e7b-4b46-8e0f-8c99f47b4e06 · outbound

This paper cites Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.117108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.005973Z digest=sha256:4523451499f02fd8f803d02b5efbd69e5ce25a8b26e694a09138d15a18473309

Observation ed812cee-cb1b-45e8-aa2c-1fff4ba8c535 · outbound

This paper cites V AD: Vectorized Scene Representation for Efficient Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving V AD: Vectorized Scene Representation for Efficient Autonomous Driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.104167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.009414Z digest=sha256:1d2645f0573766423d2fa78eba098b4852a5d65738a7e549128cbd4cb0cbcb5d

Observation 41261a53-7866-4512-91a1-9a6668329127 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.012947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.012947Z digest=sha256:cdac1e3a644a698ca0fa658702f03cc8ef60daee4e2e8495d3c3db771931cdc3

Observation 94ebaa67-f270-4961-aa0e-bd90ebb95002 · outbound

This paper cites TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.091617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.017175Z digest=sha256:bb718ab8be7554f3b15dd53eecdaf99cb70b7a9779c56e1f9bec9e7f73239046

Observation 19f1f03c-586a-497f-a89a-7ac07cf885c6 · outbound

This paper cites Textual Explana- tions for Self-Driving Vehicles.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Textual Explana- tions for Self-Driving Vehicles

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.078929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.020700Z digest=sha256:7c7dcf955ca0f553b04992d45c92c8b7fea06df55645bc533e6dfa75080a8939

Observation 3551ae27-52c8-49a8-a625-373da8b077b1 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:14.066062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.024159Z digest=sha256:8d72f4784ca3b43e57d88381ace0b29e477536b2623eae262d288add230d4ec1

Observation dbaa139d-cb07-4b0c-b657-91268b39a258 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.027750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.027750Z digest=sha256:955e53321ae35ec31e6d6dec7512c81cbda23e6fea818ee6dc94d15ada796b1a

Observation 8b01edf7-5b07-49ad-a107-6d521729b595 · outbound

This paper cites Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:02:13.516006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.031254Z digest=sha256:b08840c0d187ad3ad18976d5f83d9cd51674011c76477de525c4bc38327576ae

Observation 132f6d3e-81e1-402f-a17b-9bad4f2ea6cb · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.044325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.035268Z digest=sha256:6c31122d3f8216bad35652a29e880eb81ebaa3941ef8dea5ccf4059d68ff103d

Observation 2700579a-bb60-4ee0-9928-9f50f1406279 · outbound

This paper cites BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.031477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.038932Z digest=sha256:15bfb3035a25144f6cb8860bbfeb1e6a8080463329309cd9d3a097ceb6c4cc57

Observation 240f776d-cfbb-458d-bf3c-6dbcf0b47e46 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:14.018771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.042382Z digest=sha256:611ff857cfb4b64ad7467198cfed107a225c4c8689222022101aeb1c824f0583

Observation 002482ce-25a7-45be-8c9b-1c2510cab280 · outbound

This paper cites Visual Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.045969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.045969Z digest=sha256:85da62503edf6c41ab8f222634a52e35b237d5aba6d7e7a547f065b128290d1b

Observation 318f772c-2a14-4b7e-b0b1-3ab52449499d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Improved Baselines with Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.049457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.049457Z digest=sha256:846b4ced96544c71f830755176d305bb2348879d298c91411b95b699ac387ea7

Observation bbd26351-f20e-4f2c-a10b-ca53f9602d7b · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.052905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.052905Z digest=sha256:563a371be98ec4f90f1c7a5049e6614bd0dfd3c6eea52713c075a3635883dffd

Observation 71a49614-32f8-45c9-8c13-79e52a879db5 · outbound

This paper cites Dolphins: Multimodal Language Model for Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.979738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.056348Z digest=sha256:ff264eb7252e530e9b58889b5e65e227e45c8ab8739c4cd07f4b79624a41b501

Observation 013047aa-4b10-48d5-8b9b-9d903c2d2029 · outbound

This paper cites DRAMA: Joint Risk Localization and Captioning in Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DRAMA: Joint Risk Localization and Captioning in Driving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.967085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.060019Z digest=sha256:57cfd504cc03bf9eb3b06cad64499203ad4c97c829cb9d2a0c6e3c90606b749a

Observation 41a19cf7-1158-4448-a836-ef7911dd0dc7 · outbound

This paper cites One Million Scenes for Autonomous Driving: ONCE Dataset.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving One Million Scenes for Autonomous Driving: ONCE Dataset

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.952444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.063837Z digest=sha256:ef0f0096922d2f125876705b4232155ff2816abe268e4a68ece9d15387385686

Observation 780325de-d703-4d4a-a403-3389f88e7dae · outbound

This paper cites LingoQA: Video Question Answering for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving LingoQA: Video Question Answering for Autonomous Driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.939101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.067430Z digest=sha256:603b825764a3d24d22346c4bd691645a2e7a1c982f03b8223e2d688bb3760fcf

Observation cb004d6a-7285-419a-972f-081708cce445 · outbound

This paper cites Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.926264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.071430Z digest=sha256:7a83e8c412645dcc5d09fc3d48418852058022bbd9fa9a259211b23eaa7ca38c

Observation cf4f7d79-26b2-4640-b2c3-9b424c0eaba3 · outbound

This paper cites GPT-4 Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.074966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.074966Z digest=sha256:df57d176ed3fcb5df8b27d70242a2b23243632fcdf4f0863a3fd8a8a4019ac62

Observation 134754d4-a4e7-447c-809d-0dce2daf5388 · outbound

This paper cites VLP: Vision Language Planning for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving VLP: Vision Language Planning for Autonomous Driving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.912933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.078694Z digest=sha256:674d6b9f62a0a3ad1cd2d81e874d23b27877d62320e31b999458815a2a5489c9

Observation bcc8709b-3780-4a6a-af70-212dca4e3f26 · outbound

This paper cites NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.899507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.082205Z digest=sha256:85abe0d6fea723d1e532ed58e2ea15efe6bebfbd53c3cbf967d5aab41224e63e

Observation 97156a9e-4eca-40cb-9148-7172c33361f5 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Learning Transferable Visual Models From Natural Language Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.086057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.086057Z digest=sha256:cc70c83e0bbbbcf79318b86f3bebe92d7f9dec299f45cd21e95de36a681d5e1c

Observation 9763a031-1998-44a4-86f8-1f149144cc7b · outbound

This paper cites Rerun: A Visualization SDK for Multimodal Data, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Rerun: A Visualization SDK for Multimodal Data, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.877564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.089541Z digest=sha256:37e2dd31220ac82f1c371cc7e71af3fdcb357f276f8df10fcd6fecd8e8bd8b09

Observation 134d81e8-f02b-46f6-9c80-612713c386b6 · outbound

This paper cites Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.865560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.093398Z digest=sha256:4250a26111376da7b3492b2f49d496cbc86eac871204f7d5fc9f013dfd9dcfbe

Observation 1593a902-c933-491f-b71a-d6ebb71ea76f · outbound

This paper cites Waslander, Yu Liu, and Hong- sheng Li.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Waslander, Yu Liu, and Hong- sheng Li

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.852220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.097032Z digest=sha256:d07433f147ec464e9d1c6c7631c8191796fceb623547bb53c5aec2d1b11771aa

Observation 87b7c163-2644-4d34-b0d8-dffa49d438b4 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveLM: Driving with Graph Visual Question Answering

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.838832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.100678Z digest=sha256:7aa1b8034618cbe6f5d758e885a374fa1c278dc98115df000ae607a13ab73e2e

Observation 41fddff6-2566-4c61-a976-f856dc4612b2 · outbound

This paper cites InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.104094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.104094Z digest=sha256:91e1889d37c1fa80befe76929bda1c89ccdd1ad5dd1b287b83dc149a193b57e2

Observation 7b77396c-1291-433d-82d4-6ebf7af29670 · outbound

This paper cites Scalability in Perception for Autonomous Driving: Waymo Open Dataset.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Scalability in Perception for Autonomous Driving: Waymo Open Dataset

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.826552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.107627Z digest=sha256:864a9687308df2f31261cf2b36bac4a8c8c6887e4bf443d814f8b86d2c69a361

Observation bb8925f4-fb50-4f58-8bc6-6e3acc12d6ab · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.111100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.111100Z digest=sha256:64d95a0852bb3b94356fb589650e8cde9f6414f9b21b8e880b61521ab729cc04

Observation f1ce3a9e-3718-4ac9-a315-490601e658e7 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.814190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.114668Z digest=sha256:aaf8b606992c2cb30d5c8a35dedfbde363ce1f0988f0c62b3c05918333643301

Observation 860292ef-8592-4576-9437-ee71171cbe26 · outbound

This paper cites Object Referring in Videos With Language and Human Gaze.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Object Referring in Videos With Language and Human Gaze

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.801960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.118138Z digest=sha256:dc89d830688f8ae54d9898cb42b3e305b8ebbe14e994fe9420daab5d2aa3a418

Observation 33cdbd46-4a8c-4df3-959f-0e485c860f1c · outbound

This paper cites Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.789264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.121934Z digest=sha256:aed122ec9f12b21724a92fe9ef8de624d78f2890ff973b73f8306eb8ddfdc952

Observation f89684e8-d231-42ff-8cd8-d8119f7c033c · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.125473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.125473Z digest=sha256:f799b54f3dd83155d65edb67794b0c760812f0c8f13f8fcc2ac5fce216dd6f34

Observation cb4cbe43-40cb-4e80-8d76-ddc71f041efa · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Chain-of-thought prompting elicits reasoning in large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.129189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.129189Z digest=sha256:c3f17f119eea208883c7e8c8ef3bd8dfb2eedcf9a8c39ac94b9aa117ba18b2cc

Observation d88d3f1b-67ba-4b72-8aba-72c12dd19460 · outbound

This paper cites Argoverse 2: Next Generation Datasets for Self-driving Perception and Forecasting.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Argoverse 2: Next Generation Datasets for Self-driving Perception and Forecasting

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.768279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.132528Z digest=sha256:b03e6bf905847f099bc701bbead57e50a2ca82de898c02a037ccb5a338dae78c

Observation dfaf150f-7bbd-48cc-a930-5b5a60c4a71a · outbound

This paper cites BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.136091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.136091Z digest=sha256:3af60236a94a5acef341518fd0c101dd022119c27119a777ec5b47fb673fa977

Observation ad5f8d31-1693-43e6-83ed-ccb971147ec1 · outbound

This paper cites Language Prompt for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Language Prompt for Autonomous Driving

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.756163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.139784Z digest=sha256:e65b499222cfd2f55ab251c2863040fe39546d1de01124105a473845f6bc4a7f

Observation d4d6aed3-8deb-41f1-8318-014e26189f73 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.143198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.143198Z digest=sha256:64a128c8f1099fda1cd2ce5c27e3bb595b4ee3bfb4ad52373eda6db81e277462

Observation 07fa72f2-17f2-4265-9bb6-c46b7207bb8c · outbound

This paper cites Explainable Object-Induced Action Decision for Autonomous Vehicles.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Explainable Object-Induced Action Decision for Autonomous Vehicles

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.744300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.146937Z digest=sha256:0fbdb18979d7b3ab06ee0c09605bdfd585740f36776727c22b1741db45e5d59d

Observation e032a932-7c28-4897-8aa4-fbd41d5d5bad · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.732394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.150620Z digest=sha256:01545f3511abe186375783accb6d3336bceada0d4d6f1b6dbde21bc243ae7639

Observation 223ca51b-e1e8-43c3-b3d9-18af9342e671 · outbound

This paper cites BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.719320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.154080Z digest=sha256:86f2f57dfdd917fbf7c7e7b7d7286f2a5f8c85d7c6fcc1c69e691588f85bf983

Observation 292ac349-af25-448f-a8da-beacad958f26 · outbound

This paper cites Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.157665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.157665Z digest=sha256:bcc4a2fe8f46cd42c934e96a318b8fff7a2ac4ef0a6bc73e0559b449b0d3f38b

Observation c7220aac-9563-4639-9e3e-d34b828a3844 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Sigmoid Loss for Language Image Pre-Training

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.706389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.161399Z digest=sha256:0e709acbbc4542d3a093ef663ad8a2cf15a1417677276ce770a799154f896938

Observation dc59d335-10d2-41a7-887b-feecbf0eedf1 · outbound

This paper cites Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.693280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.165120Z digest=sha256:cb373bc4099721d83c047d16207e82bbcbe7644e8acec658c8dd32eaf6e7080b

Observation dd65be83-e74e-4cc0-a07d-55f577fadc0f · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Large Language Models Are Not Robust Multiple Choice Selectors

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.680927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.168578Z digest=sha256:a87535fd8d3f98d98c03803e1f0bb28f2d21b372ea888444c916ef55866150c0

Observation 0c025d56-caaf-488c-8a4b-a9c382724fa6 · outbound

This paper cites Doe-1: Closed-Loop Autonomous Driving with Large World Model.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Doe-1: Closed-Loop Autonomous Driving with Large World Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.171979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.171979Z digest=sha256:33aaf9fbb3621f62638148cceee42c77754a57a0862f322192dd674421f77795

Observation 213a1e93-b9ef-4b95-b707-d07d4e80dfbd · outbound

This paper cites HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.175610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.175610Z digest=sha256:5c90c9ccf209e38031600b4741ecbb660fa7895edfd9cb9d21dc18df0d376767

Observation eafc358d-10b8-4a46-b629-3beb1bc903f1 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.179479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.179479Z digest=sha256:b0ebb997d82660f692a85fccbf2b173fcc36549546a1515ab7d6996ddb93f0ae

Observation 05f11f15-3548-4c0d-9279-c56be095989d · outbound

This paper cites Embodied Understanding of Driving Scenarios.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Embodied Understanding of Driving Scenarios

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.668516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.182943Z digest=sha256:5f628660fb064bb20ec1783a038701bad27651433a450055d7d2a60380ade878

Observation c267b539-9536-4bf2-8eaa-69e75c2ce4c1 · outbound

This paper cites Ego turning left,.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Ego turning left,

Reference 77

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T06:02:13.655652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T06:02:13.187292Z digest=sha256:76da1f5cf3f16a4b12592fd7edde8da599494556318e35d65414c8f0830ff5f8

Pith citing papers

Observation 57803ea2-d4b9-4db3-bc67-e67f7b813e83 · inbound

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving cites this paper.

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:43:45.701944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T17:42:49.902997Z digest=sha256:e0e5de486dd11aead2bf6888158155df7437491e93154251f442032d7cd295ba

Observation 2a0ecd49-8cb3-4f76-8ee6-0f4f7b711920 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.170067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T05:08:46.145616Z digest=sha256:2d8e7ce83c626c7147571a03f2a587453c48a31c1ad5800c485d9d1f6608f0a7

Observation 2d3d9db3-be22-4727-97a6-fc5ab8d8f675 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:37:24.199322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T21:35:33.280591Z digest=sha256:8178df636ef061956e283580dde8781b79927241fcc75b2c650fa2a14ec8a381

Observation 761ad56b-b155-434d-a190-58ebc25690a6 · inbound

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis cites this paper.

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:46:01.453100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-09T05:40:54.265224Z digest=sha256:b99c8d01f2846c010ee4f4f329e95706a05679b357904e85cb032b8bce6f2fe3

Observation 13b4fbaf-9454-4e66-91ec-dc9f651d025f · inbound

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness cites this paper.

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T19:51:03.570643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T19:51:03.570643Z digest=sha256:ebda876c89ed6a6454463573b669001d953bf2af167bb5388301729fd8abaf4d

Observation 9b79c5e9-0f91-48c2-a21f-3c066ff5d836 · inbound

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision cites this paper.

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:21.388777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:08:21.388777Z digest=sha256:0d8541e9b75eb27d7ba60af86a867892d715b3d75708bfb116cb78ebf36b8120