Pith. sign in

Paper Citation Record · LEDGER

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

As of 20 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2505.24139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24139 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:48.142792Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31c9008d-33aa-44b6-af9a-559fbd3d502c · outbound

This paper cites GPT-4 Technical Report.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.676305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.676305Z digest=sha256:1f9fb45a4085fa43ec31375f9055a3e9ffc80343e290af000277d55625b64487

Observation 2ed61a3a-8073-4315-877b-5a3a05bd4045 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773765Z digest=sha256:44dd756fdd72e3ad00572a46b7e4b04e1bc8ae95ad55463fb1696960687ae008

Observation beddcfe4-0057-4cd1-8d22-de14e5b055cc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885701Z digest=sha256:8d8b1ea942472c35b3df010eb4e63b57753855da7927156c0e9f13a93bed1de4

Observation 66aac8d1-ebcc-4f26-8130-3fbccb7b9677 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010726Z digest=sha256:33616fc3522b0f90b3db80ba8c9c2ced5270ff4ba6a6f67212e4e1efc3366443

Observation bed03f34-f29f-4830-a37a-767fcaf7edfc · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation nuscenes: A multi- modal dataset for autonomous driving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.115143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.115143Z digest=sha256:868549dd925020af5833f2fb28b71125fc576152a8eff5d7f0a7a730738edc92

Observation cb07947d-7a04-40f8-b8fd-1dce74a61729 · outbound

This paper cites Mp3: A unified model to map, perceive, predict and plan.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Mp3: A unified model to map, perceive, predict and plan

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.952192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:39.247829Z digest=sha256:e62f4b79510025293ff090dc53962816bc66e408111ad91b3057a1c511a037ed

Observation 74c6998d-0bb9-484f-9753-cf4ef7dc9242 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414992Z digest=sha256:f0252b2836fc3b11cc8eeb4a07beab4fc919ee83503e3f6d6a3a801690366ca2

Observation 61e4c13f-0271-4849-8a80-d10972dd5686 · outbound

This paper cites Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.689795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:39.514849Z digest=sha256:1a6a00c68001e20572369babdefb895f2937de88c7cbeb1b0591db284a15a6bc

Observation cc1d5806-5d52-4b51-87e6-86b48484e907 · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Pali: A jointly- scaled multilingual language-image model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.421405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:39.632649Z digest=sha256:252495e711819fd4d8122993d3844a32001cc8492d86615ccf1d7bf021954d53

Observation 7ef6ce2e-1f01-4d44-aac5-6835a73a4daa · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786935Z digest=sha256:e3f2f25cdb314a9689a4d2ebf821fb0435dd721bbaba0b70fbfd384aec13d18e

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:43320a60247fb1bcb843382f51284d16f7dc180acc9f32675d6a05519cad6e13

Observation 1c0425df-174d-4689-8fb1-2c0ecc62c389 · outbound

This paper cites Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.065086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.065086Z digest=sha256:ddc52d573ec6659f275ba9c6ee59382008bc509288fe412af8b009c910d948df

Observation fe1abe47-c89e-4e09-b562-e7577a5048b2 · outbound

This paper cites End-to-end driving via conditional imitation learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation End-to-end driving via conditional imitation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.196150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.196150Z digest=sha256:ed7d1ade3a59e6ae32d5b8fe257445f5bb6dd8467c51b15e19c7a0694b477a48

Observation 9f760794-0c1f-42ac-8532-55d501c2ad73 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368657Z digest=sha256:bd3dbbfec73039a29b65232b3b9672caca8a7b79b46de976abcd03f32ae83655

Observation 96922c51-2048-4dd1-9fd2-44256feccb66 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.166691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:40.524852Z digest=sha256:78a61db394358c840de4eb0a2cb7be14a6212f041abd9ddb77bfafd10a784b84

Observation cb7e1ec2-84e9-4a6b-8a38-e17b054a1e1c · outbound

This paper cites Carla: An open urban driv- ing simulator.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Carla: An open urban driv- ing simulator

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663443Z digest=sha256:a23ce4312eda6f21d0b835fa4c6345a6cf5062393d79329b4f4fd45765637b85

Observation 0013014d-7d77-4c1e-baee-dae7af983233 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.779008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.779008Z digest=sha256:6c05bda7eda24eae788f3992fe1382a75780627f4b3f403d36bd1b8289eaf9e3

Observation f8049e6c-3ac0-4c77-801f-7df1528b2882 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.921078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:40.900656Z digest=sha256:3bacbac04ca1568f54d1994d841608c8ebc4500557b4c285c3fdb02ec4c12546

Observation f954c2c1-d40a-4d9f-953e-175b80a289ba · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.646465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:41.060670Z digest=sha256:90a83f74df5824e669c06ad8ddd4239547674256f891fac08872c2f882d37368

Observation 55954c71-019d-45fb-9b36-c3e7d0a0adc3 · outbound

This paper cites Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.341971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:41.217284Z digest=sha256:81f23ad6f21263a1fb00ec5b57078810ddcfc099ab61ead12cf4a10d31e6dbd9

Observation ae08d836-f8f3-4380-a4f6-34f825239a14 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation The Curious Case of Neural Text Degeneration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.350155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.350155Z digest=sha256:dacf394e1ead6fa110131a81d22d87492986080f6a58d10d2b7e824a1fc1c573

Observation 15c7e1e6-df82-4de6-8f25-392ef1cecf0e · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.475718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.475718Z digest=sha256:73f632b21534a92957ed3280959e0698c06130cf283a1cb1dabbc229c64af7fd

Observation b59e1dfc-2918-4289-b8cc-7f1207360a04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.597287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.597287Z digest=sha256:80035e262f9838ed50aad21d1e99941c7b401f98916a993f8ec6ede2e71c3b66

Observation f18ee96d-3712-4bea-909d-5cd3c557066a · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.723680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.723680Z digest=sha256:3bed98381f301f73668aefbb8a33d001dfc8acd2342c72a8ef8ba189eadf34ab

Observation 35d21fce-3f9f-4d2d-a6d9-a1cfab86d2eb · outbound

This paper cites Planning-oriented autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Planning-oriented autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.030218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:41.903017Z digest=sha256:9a4c32258bb1cd84ad81d707ae081901a758aa2d411d944558dcb9d98f61b9e8

Observation aa993d85-ab48-4e48-849a-7611211b010f · outbound

This paper cites Emma: End-to-end multimodal model for autonomous driving, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Emma: End-to-end multimodal model for autonomous driving, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.747632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.052634Z digest=sha256:9182de9f819d656ec151979cd23007b91223ad2490502b090545d7cd89a84661

Observation 6754c6eb-422a-492d-8fc9-68165356f283 · outbound

This paper cites Sym- phony: Learning realistic and diverse agents for autonomous driving simulation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sym- phony: Learning realistic and diverse agents for autonomous driving simulation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.497172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.191165Z digest=sha256:51470501aac39e321de13ce0db36f63c91137ef23898dc8966970dcd1b636638

Observation 664aad22-6a91-4e60-bfb3-8104a111ae04 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.250573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.281582Z digest=sha256:00c4e660fd162c61e1b7a38e3bd245e8f2390fd72065ea963befb31e503c9884

Observation 2be857d5-4059-46d5-9d42-28e4b1e3daf9 · outbound

This paper cites Textual explanations for self-driving ve- hicles.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Textual explanations for self-driving ve- hicles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.976491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.410239Z digest=sha256:402fa41fed12957a1bfd0c1c2eb943b6ad40ff5caa2ad38b4975b748835971a8

Observation 3c205576-dc4f-4a24-9360-742215e8f925 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.471180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.471180Z digest=sha256:8f15ad9df56c49f3a19ebfe88a5acbaaa0fb22372d62829677fdda9c5df4df5c

Observation 5bf21733-b6ee-4834-8b13-17e90502d7aa · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.676964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.638491Z digest=sha256:20872a118827a5d029eca5978c3f255e979e546332758f2ca9855d55b6b2d07f

Observation e9c72479-8522-4d78-bbcc-03ffe838c2e5 · outbound

This paper cites an unresolved cited work.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:38:54.395892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.760321Z digest=sha256:c220475ecc4cbeb00decc08965c2ebfaacc85cbfba78f85aebb82117ad147923

Observation 4bda4cc1-3aed-4cc9-9c3b-47b8eb5bca8c · outbound

This paper cites Maptr: Structured modeling and learning for online vectorized hd map construction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Maptr: Structured modeling and learning for online vectorized hd map construction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.075130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:42.821141Z digest=sha256:938e4f60e478006d414b91c58a02d976839af4f744edd7eace24612671255b0f

Observation 228d6bf1-2abc-471a-88f1-f6bdc6a814bd · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.932843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.932843Z digest=sha256:e9c11ef1d7055c343261b9c622670eb0fdc265c3fe67c65c949e380f4c734b28

Observation d01682a3-6782-4c96-85f1-e44173bc6431 · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.771640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:43.036990Z digest=sha256:cf0377898486612fdba91b57688a7ddf4d2ba2c84c50b9dabbc095feb4e1869e

Observation c2ddc1c8-05c5-4301-99e0-ce15d4d69c00 · outbound

This paper cites When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.162693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.162693Z digest=sha256:fd0a7204e2cfafeab94ad1f0af02eda6c9e9d00c44d7fd06278267b72a817991

Observation a5409517-a9f0-48ae-a5f2-665bc24bc018 · outbound

This paper cites Dolphins: Multimodal Language Model for Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Dolphins: Multimodal Language Model for Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.273602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.273602Z digest=sha256:226075451ec25081af14f6df730c1b3564d92c6167e5e65763b518c345fba227

Observation 2bfa5ad5-60c5-40f4-9171-974bef9222c6 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-Driver: Learning to Drive with GPT

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.401500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.401500Z digest=sha256:323b8a0276cec263d6f1f2f0218efb34764f50047402b518854d6053f1445320

Observation bd6c8412-f4d4-4fd7-afe4-ec11081536a1 · outbound

This paper cites A Language Agent for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation A Language Agent for Autonomous Driving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.517282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.517282Z digest=sha256:f308027c6208547cdeb237cf7bd640285f52532992990115548e385dd892e531

Observation 9e39d9be-e59f-403f-976d-c0213474736c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.703061Z digest=sha256:e603890b8745af0db315ccd25bed1b0976358416287c16ad6f48be934d34b105

Observation f61e82be-4f81-4ba0-8913-d1fad4f12ba0 · outbound

This paper cites Wayformer: Motion forecasting via simple & efficient attention networks.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Wayformer: Motion forecasting via simple & efficient attention networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.486190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:43.832722Z digest=sha256:7fa15e509c0b25dc2ec4f3aa1ffcd4997999205f1f72796c2a07807d8c443e18

Observation e67a1496-3d2d-4c88-ae80-1e58be36b576 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gpt-4v(ision) system card, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.955298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.955298Z digest=sha256:3852166c3e93fda51670d91b3bc4c92b12caae262085da1b9cf7b5804eabb060

Observation f28f4a2c-ab29-4a21-8347-cb967f1b32f6 · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.211246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:44.059307Z digest=sha256:80d0997f04a875987347d568e2bf0258019548af486c55d3075ae328a07782e3

Observation 164843cb-a90f-4556-ac38-254ef3a26b16 · outbound

This paper cites Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.982600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:44.176812Z digest=sha256:853bb7bb332b8b6011a4b50b3aa5884eeeb9fbc8e030fd78117514ddcfc2243c

Observation 2e711174-492a-42da-b0b0-db22b9d36e76 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.272624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.272624Z digest=sha256:887a9ab2c6b2c98387baec3ff0f1daa2203ec8413992167499885e0570abf6a7

Observation b3e9b324-3513-4cc7-93ea-a422d1914fe1 · outbound

This paper cites Motionlm: Multi-agent motion forecast- ing as language modeling.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Motionlm: Multi-agent motion forecast- ing as language modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.748592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:44.450630Z digest=sha256:8b8dfda865d35b6b8ff059471179ecdc8c4b88f2b172cb930d944141e5a836e3

Observation d89266e2-3fe4-4b3b-8013-48662518bcc4 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.598445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.598445Z digest=sha256:4e036b2dadc8ac13fb15e52690ef8ce977f12709c573b183776fa60715a88057

Observation 2c5eda30-d8a7-4be8-8e5c-70e790622df1 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Lmdrive: Closed-loop end-to-end driving with large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.504616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:44.772508Z digest=sha256:fac9f88f2647fee0cb138ec674d511ab179cb7bac6125d8c61b9a4bd4aa36b57

Observation c18b1de5-5d85-418d-9a19-45888623b327 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivelm: Driving with graph visual question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.276600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:44.920320Z digest=sha256:035cadde021ac58dd9d5493425c291cdcad697cdc503df4509c2934b3fb40210

Observation 9f5b6328-35fb-47f7-b67c-3ccee3ab4dc0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.070721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.070721Z digest=sha256:a18302788a6d92bc9d4430161ee038cf0c4a94a2e75eac170a068ed01eec11bc

Observation 879eb1ed-e2e9-4b61-84ef-e2af356c1312 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.214604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.214604Z digest=sha256:00bc9c80ce4193214d69b83edc9f09a798e8edd3e358593e0f6f60243eeeca31

Observation 7888535f-c224-4d0a-9828-58511b917915 · outbound

This paper cites Tokenize the world into object-level knowledge to address long-tail events in autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Tokenize the world into object-level knowledge to address long-tail events in autonomous driving

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.003873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:45.377764Z digest=sha256:822e34aee1e01295f3f20d3461b1bcabbe4b20766bdde325624fb7e7dfe4c96d

Observation 9b28db14-2292-474c-ae4f-7fa7680c4b4b · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.549532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.549532Z digest=sha256:de3c0d97368ba51f7a4165d634fe6fdc5c866334e155d8e088da9dcd6655e96e

Observation 0a03b6eb-83bf-4bae-9457-f30b89c83a66 · outbound

This paper cites Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.699027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.699027Z digest=sha256:3ab4550178658c9250eee4d1ce876403fd5512c97a13fe557b09cfc01add8148

Observation f776f205-7415-4d28-abd2-ed2c8eda2da1 · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.850206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.850206Z digest=sha256:8d57f986d90ee99d17ac6b043d4fcf2776cc3bb4f54b6419a55bc53ea991031b

Observation 26a79353-1439-4b73-8101-767c0ba465eb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.988947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.988947Z digest=sha256:91782305a0a2a6311b14647ec2b51a11cfffd9afe8f66d629608a8b5f8e38c6c

Observation 8d694b7e-af38-40e2-9fb7-028f7c06024d · outbound

This paper cites Para-drive: Parallelized architecture for real- time autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Para-drive: Parallelized architecture for real- time autonomous driving

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.716475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:46.128576Z digest=sha256:567ea9c4bdc145076efa8d60f8f2c3c4a1a4ea0c7c172f79336948d86e15d1b6

Observation 2ca55e1d-e88e-442e-a77a-233eda41c84c · outbound

This paper cites Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.402951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:46.345104Z digest=sha256:af4cc738909a212c8088da74f89c0023af4a14ef93a17f4e3fa8a3403527460f

Observation ba4c6815-35bc-404c-9a2a-03768f953b19 · outbound

This paper cites Grok-1.5 vision preview, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Grok-1.5 vision preview, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.146156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:46.496905Z digest=sha256:63ee201063ae3993662a544bb11bbfcf0fd3591e07cb3e66b649c34bfe89c272

Observation 82a398da-2cd3-41a4-8527-1dae377c99d3 · outbound

This paper cites Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.858663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:46.704458Z digest=sha256:69ad460bbcab2fcd697734fe40249687b0ad3bdf48931b86e7dc2200254e0a16

Observation 590f44fb-28c3-4cdc-86c2-d8339a9b5484 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.578166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:46.855253Z digest=sha256:752fc91f87d7ca9e91762f1650691037f910216eebf45ce27de76b95ac8da72c

Observation d2c482d2-a73b-4bfe-9cf6-f7ee97d7fef2 · outbound

This paper cites Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.007757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.007757Z digest=sha256:2902dc1f714cf610bb4b7e36ea7ff36c018acefd888ec165c95e057726c3a445

Observation c4cfb3a3-b7f1-4f2b-a5d0-80d6e43ded02 · outbound

This paper cites Sigmoid loss for language image pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sigmoid loss for language image pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.381230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:47.131659Z digest=sha256:e09cd218cdef39f8438c3c28e3d89d2fb898e0e82be7d95bb635f0596238c9ed

Observation a57a1a17-8e8e-485d-8ff8-4aafe132ef36 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation P Xing, Hao Zhang, Joseph E

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.071638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:47.261925Z digest=sha256:40cef902b9fab3a520f1d4940e411094629eabbdc7d7106aac04318fc65df26b

Observation 201df00b-e06b-4070-8860-a2f6dee2fa7f · outbound

This paper cites OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.398500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.398500Z digest=sha256:e5d96de55d97ed0f2b441bcf1c3813b8569259620beae02a720c4c738f77c51f

Observation eab5a7d9-eafd-418e-bcb6-80cc2d23eb9e · outbound

This paper cites GenAD: Generative End-to-End Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GenAD: Generative End-to-End Autonomous Driving

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.557681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.557681Z digest=sha256:c7b61029e08f325d4dcccd73c35beb1d248e2e43c209cb737aa9d79f7072d33f

Observation 73dfb289-b232-4559-bb1a-f4a08fca04c0 · outbound

This paper cites go straight forward.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation go straight forward

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:38:49.760921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:47.711590Z digest=sha256:67f0d6792d88382ef3bf28e4502d914357a46730d3157189b6e2906ea35cbf73

Observation 496eae92-decf-4a09-8c35-75b441213256 · outbound

This paper cites 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.444570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:47.858650Z digest=sha256:4111da8d1974bd7a48ac12910639511872ac569756317a0e0c39ee70a560de96

Observation 8c2c5b16-bcd8-4a11-a3dd-437c8fde576a · outbound

This paper cites 10, we visualize more planning results on WOMD-Planning-ADE.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 10, we visualize more planning results on WOMD-Planning-ADE

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.104731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:48.010515Z digest=sha256:029083ba259f5cd0a30d6847bd7f678cef1cffaf46a764bc771b5e9b236f1270

Observation 2256c5ff-8d20-4b75-b615-7ac5f6e3322e · outbound

This paper cites Camera configuration.We apply different configurations of camera sensors in Tab.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Camera configuration.We apply different configurations of camera sensors in Tab

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:48.812486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:38:48.142792Z digest=sha256:ff042d0265d7373a171f02f3a8e346e67f22fb3ad9f8f23c4005f8131e81ea17

Pith citing papers

No inbound Pith citation observations are available.