Pith. sign in

Paper Citation Record · LEDGER

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

As of 16 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 6 inbound Pith citation observations for arXiv:2506.06218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06218 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:13.187292Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:08:21.388777Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T05:46:01.450385Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy45
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 578c97c1-444d-4a6e-b1e1-13d8c2717b6f · outbound

This paper cites LLaMA 3.2: Open Foundation and Instruction Models, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving LLaMA 3.2: Open Foundation and Instruction Models, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.323684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.796168Z digest=sha256:1b3985f5606dd521204499b7a6eb2cf6918dee65c5454f14135ec1de10d9e1ac

Observation 63cdeb96-d1e5-4291-a68f-ce8c9d47efe4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.311045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.800668Z digest=sha256:24203e3587006129ed7453a3c14198f5ce81758cacc16701d01cc2549128f3b2

Observation 905ea11b-5021-431c-8a9d-558743278fe3 · outbound

This paper cites CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.298239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.805120Z digest=sha256:88677a24efa9ae81ba98e7445dfbc6ed7c5b97c08f33001c845a55475e7fbf11

Observation 13c45644-8a12-4bbd-a3a6-e11adc0b4cb0 · outbound

This paper cites Qwen2.5-VL Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.813523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.813523Z digest=sha256:54a14fe2c7f684786b09d657b284a0de1ef5ce93d2e851820a852836bdea28eb

Observation 7392d7c1-043c-49b9-bff4-33f1a3996cf2 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.285747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.817562Z digest=sha256:927fbf80c37ca2d14b52965c4c09e6c5e324cfb75aad1bc11fd8e4cc076fc396

Observation eaa9f319-698d-4da0-b5ab-8f591cb92aa1 · outbound

This paper cites MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.272732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.821890Z digest=sha256:3355596e4edf13893b29761388a573430fe0144335687611356ab8f444f152cb

Observation 29e0b0a2-5b83-403e-8b58-c4e0f603a994 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.825529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.825529Z digest=sha256:a46119446aa99d488656909bea39ae7a5f909324aa7ce54d27c880cbcf981aba

Observation be51db96-d4d2-40ce-b9b5-0fbd258ccd95 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.829100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.829100Z digest=sha256:c5927d14e89f6a4db63041905530f8273e6becef26171c0a16458f5c4c0d7848

Observation 75893966-066d-4887-8676-651159a8b5fd · outbound

This paper cites The Cityscapes Dataset for Semantic Urban Scene Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving The Cityscapes Dataset for Semantic Urban Scene Understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.242535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.832601Z digest=sha256:957cb0aee01bbca568475d8192f5475f4bca732f575506ce759e489407543277

Observation efd415f6-72e2-431c-b563-cf6736bd4c49 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.230572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.836540Z digest=sha256:80037fad617f31bfaf2acb2937aa0eda4d350abadffd273d881a466a1ceb20a5

Observation 2dc79cff-9d03-4d53-809d-9ff63e8df15b · outbound

This paper cites DeepSeek-V3 Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.840182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.840182Z digest=sha256:99fc2fdfa6758788a0c04e20ead96a864ca1fd2a6af551181c0c8210eca77610

Observation f4de31c4-3919-481f-a0d0-74aaf4f3a52f · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.218383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.844016Z digest=sha256:ba38fda98a8fb87140b8c231949db6f4e163b678687acd68030a5e0526ba0509

Observation 6def661e-8df5-4e99-bb0b-ff7a47980a23 · outbound

This paper cites Talk2Car: Taking control of your self-driving car.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Talk2Car: Taking control of your self-driving car

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.206159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.847511Z digest=sha256:9bd175aa329de298b12de6769312dde99261414bbdfce03bca4a853c4012bf8c

Observation 88631727-4d17-4ebb-aa1a-f5478a3cf72f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.850966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.850966Z digest=sha256:792464c5c5dab7d0715599875dc52842c1efb5093de2917be9f39d8b7ad55325

Observation e4a84a4e-1f60-4d2c-b8b8-495c0e52d839 · outbound

This paper cites Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.194298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.854590Z digest=sha256:75b20a887541a8cb676caaddf298217ac8beb94c35cbf2953a967664ada57cd2

Observation 601b8b08-f511-4c9f-98aa-b2353f4078fd · outbound

This paper cites CARLA: An Open Urban Driving Simulator.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving CARLA: An Open Urban Driving Simulator

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.181740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.858238Z digest=sha256:51a4a91b2015d58ec05f997dd57d08ca18ffdf0b7fa8d93c0d12722711355067

Observation d40f0edc-972c-4867-aa84-a7a60001d6e0 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.861632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.861632Z digest=sha256:d0aadb0b0bbacf27705a242cf71236062c19f0543501e12213c970570148343b

Observation ab061dd3-fe7b-4ad6-adcb-b025364a38fe · outbound

This paper cites Are we ready for autonomous driving? The KITTI vision benchmark suite.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are we ready for autonomous driving? The KITTI vision benchmark suite

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.169479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.974684Z digest=sha256:bfcd6c51c3105d029519064da7a5115318d0f340dc50e404eab17ade41ca8e4e

Observation eebbe239-95f0-44ab-85c8-545ce79f0b2b · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.978819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.978819Z digest=sha256:4efee495b495632c2dfd853510bf5a049ff681cc2a324318b4e360e8562952cd

Observation 9e87d558-b20f-4cef-ac97-7223b62a16b7 · outbound

This paper cites Planning-oriented Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Planning-oriented Autonomous Driving

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.156715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.982852Z digest=sha256:d6441ffe394d838638ac98efb23961f5189af23aac70a734e9a57d06cf2f5bc7

Observation 83e38495-3cfb-4db7-937a-892a4f4ae3f1 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.986482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.986482Z digest=sha256:7cf67636cb82dd7da6752152ea70b31c12da12bced26947d536017d4dd5163b9

Observation d24a4d13-87c9-47ef-b6dd-2082928fa097 · outbound

This paper cites Making Large Language Models Better Planners with Reasoning-Decision Alignment.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Making Large Language Models Better Planners with Reasoning-Decision Alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.144030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.990385Z digest=sha256:3e85834a6dc8ea091af9066badefbd7b759a8c7844f39fd16a19b9c16fff7c0b

Observation cefcec4e-63ab-43cc-9f01-be84f09718ae · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:12.993977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:12.993977Z digest=sha256:fbc716d3d1b437cd6bbbf40997e1152846cffec50ec4473826d280acc67cbdb5

Observation 4dd1bbdd-79ef-4d20-a514-4a992137034b · outbound

This paper cites NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.130940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:12.998184Z digest=sha256:dc35e5d442d94a04f185b7653ee223ac0a7740c8b6abea632c6853c5eeff6b9e

Observation 5fd9b368-d0d3-4e47-881c-78b07fcce794 · outbound

This paper cites DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.001866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.001866Z digest=sha256:15b390e23f104523bd70fbe9f8e13fb11b6769cc3a7b710137fff758b0c56045

Observation ddd1bf57-9e7b-4b46-8e0f-8c99f47b4e06 · outbound

This paper cites Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.117108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.005973Z digest=sha256:7febf16e09598e8556d0c84476814eca469b50732790ed56c0e9d7eb72a81d57

Observation ed812cee-cb1b-45e8-aa2c-1fff4ba8c535 · outbound

This paper cites V AD: Vectorized Scene Representation for Efficient Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving V AD: Vectorized Scene Representation for Efficient Autonomous Driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.104167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.009414Z digest=sha256:fad9c075a1768820a2305e04f63992f9ce53428e47c5da88bdb9e0bb61bb074a

Observation 41261a53-7866-4512-91a1-9a6668329127 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.012947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.012947Z digest=sha256:cdac1e3a644a698ca0fa658702f03cc8ef60daee4e2e8495d3c3db771931cdc3

Observation 94ebaa67-f270-4961-aa0e-bd90ebb95002 · outbound

This paper cites TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.091617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.017175Z digest=sha256:fa80ab760be191a82413aef608b5e5eda07e448485fe3aaf1184985eb875dd9c

Observation 19f1f03c-586a-497f-a89a-7ac07cf885c6 · outbound

This paper cites Textual Explana- tions for Self-Driving Vehicles.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Textual Explana- tions for Self-Driving Vehicles

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.078929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.020700Z digest=sha256:559dbbc320f1e7512ffe4c749f7ca5dbf5b3e70d4d6d250d20100faa88804218

Observation 3551ae27-52c8-49a8-a625-373da8b077b1 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:14.066062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.024159Z digest=sha256:292352e5fd4b8b1f560fbf5bc1da4a7af2ce575bab66e5ac87e66bf28cf13f3e

Observation dbaa139d-cb07-4b0c-b657-91268b39a258 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.027750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.027750Z digest=sha256:955e53321ae35ec31e6d6dec7512c81cbda23e6fea818ee6dc94d15ada796b1a

Observation 8b01edf7-5b07-49ad-a107-6d521729b595 · outbound

This paper cites Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:02:13.516006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.031254Z digest=sha256:ce9ee21b4438c05231b532d58df9f55675dec4cb01d6044d930a4f2be959b613

Observation 132f6d3e-81e1-402f-a17b-9bad4f2ea6cb · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.044325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.035268Z digest=sha256:d68f9cd8292ac7daa6ebb0c6c97a0590cc80aa0b196d634b0c26737e2e1525a6

Observation 2700579a-bb60-4ee0-9928-9f50f1406279 · outbound

This paper cites BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:14.031477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.038932Z digest=sha256:cc64018cbbf1121644f13b306ae948cd9c69d4f8726e696370e170b25f75aeb5

Observation 240f776d-cfbb-458d-bf3c-6dbcf0b47e46 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:14.018771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.042382Z digest=sha256:d6d1505a32b60a64230082deb4ed5e520487ec3f8361370bea2fbe0dfd6e3ba7

Observation 002482ce-25a7-45be-8c9b-1c2510cab280 · outbound

This paper cites Visual Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.045969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.045969Z digest=sha256:85da62503edf6c41ab8f222634a52e35b237d5aba6d7e7a547f065b128290d1b

Observation 318f772c-2a14-4b7e-b0b1-3ab52449499d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Improved Baselines with Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.049457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.049457Z digest=sha256:846b4ced96544c71f830755176d305bb2348879d298c91411b95b699ac387ea7

Observation bbd26351-f20e-4f2c-a10b-ca53f9602d7b · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.052905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.052905Z digest=sha256:563a371be98ec4f90f1c7a5049e6614bd0dfd3c6eea52713c075a3635883dffd

Observation 71a49614-32f8-45c9-8c13-79e52a879db5 · outbound

This paper cites Dolphins: Multimodal Language Model for Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.979738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.056348Z digest=sha256:e6ef66b8947155a7886a2ba664fce5c17a849372f90acae9b2e19e0aa178dd59

Observation 013047aa-4b10-48d5-8b9b-9d903c2d2029 · outbound

This paper cites DRAMA: Joint Risk Localization and Captioning in Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DRAMA: Joint Risk Localization and Captioning in Driving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.967085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.060019Z digest=sha256:14707d654e2c764323d9fa0a30b042a82afd2438de5d1e0fe10086b55c29c3ea

Observation 41a19cf7-1158-4448-a836-ef7911dd0dc7 · outbound

This paper cites One Million Scenes for Autonomous Driving: ONCE Dataset.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving One Million Scenes for Autonomous Driving: ONCE Dataset

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.952444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.063837Z digest=sha256:e4c64e6eb8c1e717a5b3c89bf9def0517ed743e88bbf21c822d1bc8346a65d8e

Observation 780325de-d703-4d4a-a403-3389f88e7dae · outbound

This paper cites LingoQA: Video Question Answering for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving LingoQA: Video Question Answering for Autonomous Driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.939101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.067430Z digest=sha256:25a7200da477bdd89e42fe887ce34ed738d748cc4a4bfca1f16582d7148aaa21

Observation cb004d6a-7285-419a-972f-081708cce445 · outbound

This paper cites Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.926264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.071430Z digest=sha256:3e90524e6e1119d26295de58e3eb0fd5f79e4dd3da12ccbf4ee93c1893e1ef4b

Observation cf4f7d79-26b2-4640-b2c3-9b424c0eaba3 · outbound

This paper cites GPT-4 Technical Report.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.074966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.074966Z digest=sha256:70a65956b4e4417e0a39d4c470c4047e84f6d60afa0407c6a016eb0520331ab6

Observation 134754d4-a4e7-447c-809d-0dce2daf5388 · outbound

This paper cites VLP: Vision Language Planning for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving VLP: Vision Language Planning for Autonomous Driving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.912933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.078694Z digest=sha256:efb5d6ad44ab35e7642a65eae1a69fbfe527da332895e3b18f8a996ccc5b40fe

Observation bcc8709b-3780-4a6a-af70-212dca4e3f26 · outbound

This paper cites NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.899507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.082205Z digest=sha256:9502265abfbad5a5414930f88951f2f06d9a2c1e4dbcc6b9ddb4baf074c25a54

Observation 97156a9e-4eca-40cb-9148-7172c33361f5 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Learning Transferable Visual Models From Natural Language Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.086057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.086057Z digest=sha256:cc70c83e0bbbbcf79318b86f3bebe92d7f9dec299f45cd21e95de36a681d5e1c

Observation 9763a031-1998-44a4-86f8-1f149144cc7b · outbound

This paper cites Rerun: A Visualization SDK for Multimodal Data, 2024.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Rerun: A Visualization SDK for Multimodal Data, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.877564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.089541Z digest=sha256:65d608c6ed9ff09d47d5b42c9663a899f078a2e2076d48e7a0266a3039919f80

Observation 134d81e8-f02b-46f6-9c80-612713c386b6 · outbound

This paper cites Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.865560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.093398Z digest=sha256:0bbad53c863efc1624fe2511f8fc3414d4baf1a597681ccb843945051834e059

Observation 1593a902-c933-491f-b71a-d6ebb71ea76f · outbound

This paper cites Waslander, Yu Liu, and Hong- sheng Li.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Waslander, Yu Liu, and Hong- sheng Li

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.852220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.097032Z digest=sha256:366f7eebe411d761851bdab1ad9208bfabb9e6ec418317dee7e1603a7d8d7b1b

Observation 87b7c163-2644-4d34-b0d8-dffa49d438b4 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveLM: Driving with Graph Visual Question Answering

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.838832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.100678Z digest=sha256:0cbceffddefe073fb376e62bb1bda20e7934b7c4854cfc7c4fc9d50806a75301

Observation 41fddff6-2566-4c61-a976-f856dc4612b2 · outbound

This paper cites InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.104094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.104094Z digest=sha256:91e1889d37c1fa80befe76929bda1c89ccdd1ad5dd1b287b83dc149a193b57e2

Observation 7b77396c-1291-433d-82d4-6ebf7af29670 · outbound

This paper cites Scalability in Perception for Autonomous Driving: Waymo Open Dataset.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Scalability in Perception for Autonomous Driving: Waymo Open Dataset

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.826552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.107627Z digest=sha256:e6f3f42349fdbe84f4dc474456e142b86a102a5757ba1275af2a3d4171a3c4a6

Observation bb8925f4-fb50-4f58-8bc6-6e3acc12d6ab · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.111100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.111100Z digest=sha256:64d95a0852bb3b94356fb589650e8cde9f6414f9b21b8e880b61521ab729cc04

Observation f1ce3a9e-3718-4ac9-a315-490601e658e7 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.814190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.114668Z digest=sha256:612f6be351b4fe59575b05712ae092f1c33601ad15f93753e45a165bf6d70208

Observation 860292ef-8592-4576-9437-ee71171cbe26 · outbound

This paper cites Object Referring in Videos With Language and Human Gaze.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Object Referring in Videos With Language and Human Gaze

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.801960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.118138Z digest=sha256:72e1c49bf5a7f059f2779331ff9f96f4634225607689e22cf5fde132aba28f8b

Observation 33cdbd46-4a8c-4df3-959f-0e485c860f1c · outbound

This paper cites Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.789264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.121934Z digest=sha256:995f2354bcbbc6e19ed9510ba788aaa9467089640959adc9cd4ff041ab10fffa

Observation f89684e8-d231-42ff-8cd8-d8119f7c033c · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.125473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.125473Z digest=sha256:f799b54f3dd83155d65edb67794b0c760812f0c8f13f8fcc2ac5fce216dd6f34

Observation cb4cbe43-40cb-4e80-8d76-ddc71f041efa · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Chain-of-thought prompting elicits reasoning in large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.129189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.129189Z digest=sha256:c3f17f119eea208883c7e8c8ef3bd8dfb2eedcf9a8c39ac94b9aa117ba18b2cc

Observation d88d3f1b-67ba-4b72-8aba-72c12dd19460 · outbound

This paper cites Argoverse 2: Next Generation Datasets for Self-driving Perception and Forecasting.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Argoverse 2: Next Generation Datasets for Self-driving Perception and Forecasting

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.768279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.132528Z digest=sha256:01735bba051d8fd60c8525538a73f8b7e233f881c1996de0b467aec9730e1c66

Observation dfaf150f-7bbd-48cc-a930-5b5a60c4a71a · outbound

This paper cites BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.136091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.136091Z digest=sha256:3af60236a94a5acef341518fd0c101dd022119c27119a777ec5b47fb673fa977

Observation ad5f8d31-1693-43e6-83ed-ccb971147ec1 · outbound

This paper cites Language Prompt for Autonomous Driving.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Language Prompt for Autonomous Driving

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.756163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.139784Z digest=sha256:063cb60614a82060064288fbe2cc163ea72ce9536a3c578ef26342f9e9483303

Observation d4d6aed3-8deb-41f1-8318-014e26189f73 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.143198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.143198Z digest=sha256:64a128c8f1099fda1cd2ce5c27e3bb595b4ee3bfb4ad52373eda6db81e277462

Observation 07fa72f2-17f2-4265-9bb6-c46b7207bb8c · outbound

This paper cites Explainable Object-Induced Action Decision for Autonomous Vehicles.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Explainable Object-Induced Action Decision for Autonomous Vehicles

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.744300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.146937Z digest=sha256:f4cb9340cad30f4c78f147c6866f96c12064ddafc12b9c07c07b4d0a1002753e

Observation e032a932-7c28-4897-8aa4-fbd41d5d5bad · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.732394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.150620Z digest=sha256:d7633babd1f398733ca9a4c3ada74e00b65ac2790c782c0bfa368135146943d2

Observation 223ca51b-e1e8-43c3-b3d9-18af9342e671 · outbound

This paper cites BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.719320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.154080Z digest=sha256:6b7cad135450f37aba908b054a24bcff7526aa57f4a427ec0748ce5a39451115

Observation 292ac349-af25-448f-a8da-beacad958f26 · outbound

This paper cites Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.157665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.157665Z digest=sha256:bcc4a2fe8f46cd42c934e96a318b8fff7a2ac4ef0a6bc73e0559b449b0d3f38b

Observation c7220aac-9563-4639-9e3e-d34b828a3844 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Sigmoid Loss for Language Image Pre-Training

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.706389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.161399Z digest=sha256:9980bbb5351ccb9f3c65250354a0f091ea8b5e0ddaaee24eb55ddd2dbef55dbb

Observation dc59d335-10d2-41a7-887b-feecbf0eedf1 · outbound

This paper cites Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.693280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.165120Z digest=sha256:70db7734f208788a52cbf84883f9b6af37fd929ec24dafd4397ed8fab6fbbaab

Observation dd65be83-e74e-4cc0-a07d-55f577fadc0f · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Large Language Models Are Not Robust Multiple Choice Selectors

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.680927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.168578Z digest=sha256:0e66ab839d9576f2fb2aa0ddbdfe7c12ca8d4b8650ba6741d9f997ea818ad60d

Observation 0c025d56-caaf-488c-8a4b-a9c382724fa6 · outbound

This paper cites Doe-1: Closed-Loop Autonomous Driving with Large World Model.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Doe-1: Closed-Loop Autonomous Driving with Large World Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.171979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.171979Z digest=sha256:33aaf9fbb3621f62638148cceee42c77754a57a0862f322192dd674421f77795

Observation 213a1e93-b9ef-4b95-b707-d07d4e80dfbd · outbound

This paper cites HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.175610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.175610Z digest=sha256:5c90c9ccf209e38031600b4741ecbb660fa7895edfd9cb9d21dc18df0d376767

Observation eafc358d-10b8-4a46-b629-3beb1bc903f1 · outbound

This paper cites an unresolved cited work.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:13.179479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:13.179479Z digest=sha256:b0ebb997d82660f692a85fccbf2b173fcc36549546a1515ab7d6996ddb93f0ae

Observation 05f11f15-3548-4c0d-9279-c56be095989d · outbound

This paper cites Embodied Understanding of Driving Scenarios.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Embodied Understanding of Driving Scenarios

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:13.668516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.182943Z digest=sha256:fab110828f6fd99c5b86af2f17e4f1ad0c3c889c01c9e1383804f0414ed35d26

Observation c267b539-9536-4bf2-8eaa-69e75c2ce4c1 · outbound

This paper cites Ego turning left,.

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving Ego turning left,

Reference 77

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T06:02:13.655652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T06:02:13.187292Z digest=sha256:3eabe3f5b3c5d81aa52efa3a848d15980797cc36efa9b298cf269864a431be74

Pith citing papers

Observation 57803ea2-d4b9-4db3-bc67-e67f7b813e83 · inbound

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving cites this paper.

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:43:45.701944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T17:42:49.902997Z digest=sha256:7673e181cd4bb1e38fd0db2ff5a08d6696750e23f75fa7b5d268481e209da70b

Observation 2a0ecd49-8cb3-4f76-8ee6-0f4f7b711920 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.170067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T05:08:46.145616Z digest=sha256:ac432b3d6b6c4f303a04812b12419817fbbde192e079448401201b3ff235d936

Observation 2d3d9db3-be22-4727-97a6-fc5ab8d8f675 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:37:24.199322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T21:35:33.280591Z digest=sha256:e65481db986d99cf09b0d5012ac7aac7e8fed50fd03045819143d788999c622a

Observation 761ad56b-b155-434d-a190-58ebc25690a6 · inbound

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis cites this paper.

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:46:01.453100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-09T05:40:54.265224Z digest=sha256:f27f6b07faa9d0d26ca695a16f5e4b82596520bf573cfa5b53af104ca53e7f1b

Observation 13b4fbaf-9454-4e66-91ec-dc9f651d025f · inbound

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness cites this paper.

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T19:51:03.570643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T19:51:03.570643Z digest=sha256:ebda876c89ed6a6454463573b669001d953bf2af167bb5388301729fd8abaf4d

Observation 9b79c5e9-0f91-48c2-a21f-3c066ff5d836 · inbound

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision cites this paper.

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:21.388777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:08:21.388777Z digest=sha256:0d8541e9b75eb27d7ba60af86a867892d715b3d75708bfb116cb78ebf36b8120