Pith. sign in

Paper Citation Record · LEDGER

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

As of 13 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2505.24139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24139 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:48.142792Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31c9008d-33aa-44b6-af9a-559fbd3d502c · outbound

This paper cites GPT-4 Technical Report.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.676305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.676305Z digest=sha256:287a1713f5170023b89153dff41c8587c6bbd39fced81be0ca7037eec9d859b7

Observation 2ed61a3a-8073-4315-877b-5a3a05bd4045 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773765Z digest=sha256:5c7b523640e79a6ef2414464295f6037dc7810a71035ef1745dfedcb540a32c8

Observation beddcfe4-0057-4cd1-8d22-de14e5b055cc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885701Z digest=sha256:ffa4b3cd9e9db660a50394c5755e82f49f34da69170549984e70fc69a33ee7a8

Observation 66aac8d1-ebcc-4f26-8130-3fbccb7b9677 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010726Z digest=sha256:77eac35c1a0cea77e06896e850fa491cd16f4a2a40394b18818242324c20991b

Observation bed03f34-f29f-4830-a37a-767fcaf7edfc · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation nuscenes: A multi- modal dataset for autonomous driving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.115143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.115143Z digest=sha256:7fa447b6307800a6c43dc6de3f9a1b61cd07b4bd12b83e87fc50bfedc131b08e

Observation cb07947d-7a04-40f8-b8fd-1dce74a61729 · outbound

This paper cites Mp3: A unified model to map, perceive, predict and plan.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Mp3: A unified model to map, perceive, predict and plan

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.952192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:39.247829Z digest=sha256:bb9eb20cbc6b12c3136e04648acf6a720c11d198e57863b0a6bd939f9b48308c

Observation 74c6998d-0bb9-484f-9753-cf4ef7dc9242 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414992Z digest=sha256:2f964afd549b74dee66b97ca127c0a0789b1bb17588b28b7113b8ac1a956b4dc

Observation 61e4c13f-0271-4849-8a80-d10972dd5686 · outbound

This paper cites Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.689795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:39.514849Z digest=sha256:3e66c9f0d26832671022c3804ec41173d4200a773276cf28507527db81849a29

Observation cc1d5806-5d52-4b51-87e6-86b48484e907 · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Pali: A jointly- scaled multilingual language-image model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.421405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:39.632649Z digest=sha256:0b4a926df27f4239d1d526f4b17124970fb63404a52c9f17a7cc45372407abe7

Observation 7ef6ce2e-1f01-4d44-aac5-6835a73a4daa · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786935Z digest=sha256:c1e643de77569eb8d02cd394fd32c22e1581e84f426f9793851f5450cc120249

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:48a228beeca8ed914e7aceacf4b01a6f2bec7980f3f49c8cd052a97f18e26840

Observation 1c0425df-174d-4689-8fb1-2c0ecc62c389 · outbound

This paper cites Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.065086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.065086Z digest=sha256:f07facebb9b6d27488db2a9c34ffd8671593277bca119cfa79ed8beab0e9ffe1

Observation fe1abe47-c89e-4e09-b562-e7577a5048b2 · outbound

This paper cites End-to-end driving via conditional imitation learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation End-to-end driving via conditional imitation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.196150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.196150Z digest=sha256:a44faa1fd89a795bcbd1c8f106cc57c1bd0b19da08729699842c2fce631f084b

Observation 9f760794-0c1f-42ac-8532-55d501c2ad73 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368657Z digest=sha256:c5ceb86858ec1552d1aa69423dfbd4f0749364205861f09a335c45ac1c86cf15

Observation 96922c51-2048-4dd1-9fd2-44256feccb66 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.166691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:40.524852Z digest=sha256:d7c5fa8fcb33f802949af057548f4e0078aab7ec5274cb40e221dc167fdee292

Observation cb7e1ec2-84e9-4a6b-8a38-e17b054a1e1c · outbound

This paper cites Carla: An open urban driv- ing simulator.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Carla: An open urban driv- ing simulator

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663443Z digest=sha256:8a4b3bc1af965a2406e0e082e183dc052a69783b64abe199a3b0c4c8078064ab

Observation 0013014d-7d77-4c1e-baee-dae7af983233 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.779008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.779008Z digest=sha256:8c39da08002f86e1dfd530e461b590519981a612754e2bbb823f5b09821a9895

Observation f8049e6c-3ac0-4c77-801f-7df1528b2882 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.921078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:40.900656Z digest=sha256:954b8a17d003b8a05c062744497e3978f23f53725de1cba39770229db1aab73a

Observation f954c2c1-d40a-4d9f-953e-175b80a289ba · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.646465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:41.060670Z digest=sha256:960eaa131b44e11abdc61366660b6c7cf9a87a2d89f72aeb8e8b2854d8c8eeed

Observation 55954c71-019d-45fb-9b36-c3e7d0a0adc3 · outbound

This paper cites Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.341971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:41.217284Z digest=sha256:978be4a9aa7025a62d3c87636323f41d985fe5954ae2214160a34361dda8b764

Observation ae08d836-f8f3-4380-a4f6-34f825239a14 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation The Curious Case of Neural Text Degeneration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.350155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.350155Z digest=sha256:4541c59d47b83a9ef1502e3eace2fc90f6bd14caea01195a064b0669b6cf458b

Observation 15c7e1e6-df82-4de6-8f25-392ef1cecf0e · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.475718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.475718Z digest=sha256:6701a8828f2115e96d20b77879fe0093ca7efd7b1fec14bc066474814d478c77

Observation b59e1dfc-2918-4289-b8cc-7f1207360a04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.597287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.597287Z digest=sha256:cfb12551ef8db33857bab14ac3ee920a53a6e13c7f58ab9b9068653cc1cd33f5

Observation f18ee96d-3712-4bea-909d-5cd3c557066a · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.723680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.723680Z digest=sha256:23dfc4f3884e6af6a2b77f551319f1a68f6f44c4cfe965480613a776ad9e8704

Observation 35d21fce-3f9f-4d2d-a6d9-a1cfab86d2eb · outbound

This paper cites Planning-oriented autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Planning-oriented autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.030218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:41.903017Z digest=sha256:b4bfae7fbf8c7850d4591553cd449a5716bdd6f216e45e2bbdd775a2f88c4371

Observation aa993d85-ab48-4e48-849a-7611211b010f · outbound

This paper cites Emma: End-to-end multimodal model for autonomous driving, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Emma: End-to-end multimodal model for autonomous driving, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.747632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.052634Z digest=sha256:f8d43eeb3ccba41e99e36891538d3efa2a05ecfbf1941b99e99f515c2f3a11fc

Observation 6754c6eb-422a-492d-8fc9-68165356f283 · outbound

This paper cites Sym- phony: Learning realistic and diverse agents for autonomous driving simulation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sym- phony: Learning realistic and diverse agents for autonomous driving simulation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.497172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.191165Z digest=sha256:0c6ebb3347cac06471bc3623d6d60cd6e02ba7e53510ae0bd8e5c367389be554

Observation 664aad22-6a91-4e60-bfb3-8104a111ae04 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.250573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.281582Z digest=sha256:17eea9a059b4e14b3d525c94c0a069c80144cb515bcd26f524362a2b87487caf

Observation 2be857d5-4059-46d5-9d42-28e4b1e3daf9 · outbound

This paper cites Textual explanations for self-driving ve- hicles.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Textual explanations for self-driving ve- hicles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.976491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.410239Z digest=sha256:5056c332aba073f5082b32c8e278b1768f89dd600ec79d6f7f194345a49115ab

Observation 3c205576-dc4f-4a24-9360-742215e8f925 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.471180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.471180Z digest=sha256:c35c089a6db63f9e82726789cb10c289b6d397126a41cf451150bed15ddd7b04

Observation 5bf21733-b6ee-4834-8b13-17e90502d7aa · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.676964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.638491Z digest=sha256:67303151cdc6df45d78f50c7a77b35895169109d38e96d33eabbd05d8dd0d0bf

Observation e9c72479-8522-4d78-bbcc-03ffe838c2e5 · outbound

This paper cites an unresolved cited work.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:38:54.395892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.760321Z digest=sha256:41b33f8e55827fc653369b9a273a17b3a383c7337b3cd75b740c76da2798f439

Observation 4bda4cc1-3aed-4cc9-9c3b-47b8eb5bca8c · outbound

This paper cites Maptr: Structured modeling and learning for online vectorized hd map construction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Maptr: Structured modeling and learning for online vectorized hd map construction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.075130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:42.821141Z digest=sha256:e46cdb06facff66482b96430143631e1c157784e7d5f13de952e5ea1d969381f

Observation 228d6bf1-2abc-471a-88f1-f6bdc6a814bd · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.932843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.932843Z digest=sha256:f49035f4bed6fd9e11f763a520faf75f73648c7a688f1b4e6ee41fc87272abd1

Observation d01682a3-6782-4c96-85f1-e44173bc6431 · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.771640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:43.036990Z digest=sha256:8886a3608ed0775a21914bc36390dab3f1d3c49f7ce73dfb9a3eeaea0ee66612

Observation c2ddc1c8-05c5-4301-99e0-ce15d4d69c00 · outbound

This paper cites When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.162693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.162693Z digest=sha256:432bdd8e5fa821ce92e808419aea66c3e51d20ebe7ebbb296b652a0dafef67dc

Observation a5409517-a9f0-48ae-a5f2-665bc24bc018 · outbound

This paper cites Dolphins: Multimodal Language Model for Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Dolphins: Multimodal Language Model for Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.273602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.273602Z digest=sha256:98374ba1afad521a8e9593f7806f18a0b4f94ee3a87433acf70a95638d5ad512

Observation 2bfa5ad5-60c5-40f4-9171-974bef9222c6 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-Driver: Learning to Drive with GPT

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.401500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.401500Z digest=sha256:801f52fc6593a69bc17150a3e03a7f9e293c40ad48f3eb2b6fe6aa879c0f581f

Observation bd6c8412-f4d4-4fd7-afe4-ec11081536a1 · outbound

This paper cites A Language Agent for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation A Language Agent for Autonomous Driving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.517282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.517282Z digest=sha256:c3ffe0d2e64e286fabd47c661c6cc7b46b944f4587ea8bc8e9650134bfc4958c

Observation 9e39d9be-e59f-403f-976d-c0213474736c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.703061Z digest=sha256:0e77480d3aefe20f6618d3b8a86f10a97ff4473b5f83bce4e7765d6c13855bda

Observation f61e82be-4f81-4ba0-8913-d1fad4f12ba0 · outbound

This paper cites Wayformer: Motion forecasting via simple & efficient attention networks.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Wayformer: Motion forecasting via simple & efficient attention networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.486190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:43.832722Z digest=sha256:835bdeb1dc8362c53de55ba57de36916ad58dcae3907930f4b94d1812c8e86e1

Observation e67a1496-3d2d-4c88-ae80-1e58be36b576 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gpt-4v(ision) system card, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.955298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.955298Z digest=sha256:00993f2aee209e1a7d35033b77ca10161ba566020d9a00dad2cd49ddfded4a09

Observation f28f4a2c-ab29-4a21-8347-cb967f1b32f6 · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.211246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:44.059307Z digest=sha256:d6dc4b873757faf71537bd18eb859d11dfd65dd941746eb67534259948fc2abc

Observation 164843cb-a90f-4556-ac38-254ef3a26b16 · outbound

This paper cites Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.982600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:44.176812Z digest=sha256:9dab4e73df47fe4501e3fd1cb179ace73628fc2f48f77ac85edb578a72fbc980

Observation 2e711174-492a-42da-b0b0-db22b9d36e76 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.272624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.272624Z digest=sha256:e1a0d1bc1e847b7420e5b6d71b432980e02fbfe5ea764ed980de0fb9a97c2b31

Observation b3e9b324-3513-4cc7-93ea-a422d1914fe1 · outbound

This paper cites Motionlm: Multi-agent motion forecast- ing as language modeling.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Motionlm: Multi-agent motion forecast- ing as language modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.748592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:44.450630Z digest=sha256:800eda74f3ebbf9592373aa6c0698db568d64273663f91f1cc83dfc43b6b449b

Observation d89266e2-3fe4-4b3b-8013-48662518bcc4 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.598445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.598445Z digest=sha256:b8fdee6ef89523ef00080e15967a0059a164cca5103af3a71f39d3e492c75b29

Observation 2c5eda30-d8a7-4be8-8e5c-70e790622df1 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Lmdrive: Closed-loop end-to-end driving with large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.504616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:44.772508Z digest=sha256:b28ac281b9e9cd82277d062d3661d0f65f30b3dac7828a428fefdde4dddf23f5

Observation c18b1de5-5d85-418d-9a19-45888623b327 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivelm: Driving with graph visual question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.276600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:44.920320Z digest=sha256:b3dbd856ff41149876cac40ba018c1a5eed60784000655213c11531f6ff65ce3

Observation 9f5b6328-35fb-47f7-b67c-3ccee3ab4dc0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.070721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.070721Z digest=sha256:10329b0ced8561e5a9cffb565d351fb34a44c1bf36f213ff3c92a399e0506a9f

Observation 879eb1ed-e2e9-4b61-84ef-e2af356c1312 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.214604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.214604Z digest=sha256:c535c416df942ceb828a03ee921fe8cce71389130444d460083c8638c9727359

Observation 7888535f-c224-4d0a-9828-58511b917915 · outbound

This paper cites Tokenize the world into object-level knowledge to address long-tail events in autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Tokenize the world into object-level knowledge to address long-tail events in autonomous driving

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.003873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:45.377764Z digest=sha256:7d62d70c011ae011de2503d41db58d202babd93abd96f4d3718cde7f6cdf5884

Observation 9b28db14-2292-474c-ae4f-7fa7680c4b4b · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.549532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.549532Z digest=sha256:49dad804be794cae30eac24a24fc34ab4a90db97412062ab81d8689f73729848

Observation 0a03b6eb-83bf-4bae-9457-f30b89c83a66 · outbound

This paper cites Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.699027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.699027Z digest=sha256:5049234ca1b5d672a04fdd8177aadd249124a5799d59e2938f3820b49e63e7e4

Observation f776f205-7415-4d28-abd2-ed2c8eda2da1 · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.850206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.850206Z digest=sha256:1b2e93a3cdc629ce08189b68a5558b0c65b8000236b051892c87cc7a258aa703

Observation 26a79353-1439-4b73-8101-767c0ba465eb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.988947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.988947Z digest=sha256:e1bfb80ed993ae4786f1b5d378807c33bac1e04313ce8bcbdb0e2d84e14e2110

Observation 8d694b7e-af38-40e2-9fb7-028f7c06024d · outbound

This paper cites Para-drive: Parallelized architecture for real- time autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Para-drive: Parallelized architecture for real- time autonomous driving

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.716475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:46.128576Z digest=sha256:0425bce27c8020a501f27c583810fcee5183bc7fc589bfea68c0e2c894e49942

Observation 2ca55e1d-e88e-442e-a77a-233eda41c84c · outbound

This paper cites Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.402951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:46.345104Z digest=sha256:da9488ca439ea1baaf0fb12018b4ebcc7cb316dd9a78838639ccbc0592cfce07

Observation ba4c6815-35bc-404c-9a2a-03768f953b19 · outbound

This paper cites Grok-1.5 vision preview, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Grok-1.5 vision preview, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.146156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:46.496905Z digest=sha256:c6048e4aa4a0f59a02bc0e801ea67c512acc927bdea2494c03559df4b1b44b57

Observation 82a398da-2cd3-41a4-8527-1dae377c99d3 · outbound

This paper cites Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.858663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:46.704458Z digest=sha256:2954557ece1e6767df2c4eaa00c29e731d814a360a17f04a18356752f172a788

Observation 590f44fb-28c3-4cdc-86c2-d8339a9b5484 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.578166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:46.855253Z digest=sha256:08e39c70827463b54819d1d3875aca4a0ef7f0e385a4f40a90257c2abbf28035

Observation d2c482d2-a73b-4bfe-9cf6-f7ee97d7fef2 · outbound

This paper cites Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.007757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.007757Z digest=sha256:da34bd7e120a71d856235b24ae2eab6291e816690db29e8aa085b25502d931ea

Observation c4cfb3a3-b7f1-4f2b-a5d0-80d6e43ded02 · outbound

This paper cites Sigmoid loss for language image pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sigmoid loss for language image pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.381230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:47.131659Z digest=sha256:08afb55f8966feb75449db79d3daebc96331f5d44d7f8252d507d0eac8362c94

Observation a57a1a17-8e8e-485d-8ff8-4aafe132ef36 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation P Xing, Hao Zhang, Joseph E

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.071638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:47.261925Z digest=sha256:00547d4b8869a5a536e0d6883015c40b64ac1ef0ad392d790ed827f41b7e7c40

Observation 201df00b-e06b-4070-8860-a2f6dee2fa7f · outbound

This paper cites OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.398500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.398500Z digest=sha256:4d0a65e97f1742f53d197391f84dcbd336cb53af7626741fb327ff1f372c4dc1

Observation eab5a7d9-eafd-418e-bcb6-80cc2d23eb9e · outbound

This paper cites GenAD: Generative End-to-End Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GenAD: Generative End-to-End Autonomous Driving

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.557681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.557681Z digest=sha256:694551a8e71351adc5de4d23b35408574dcbb66bef71d582b56df4e945c5e405

Observation 73dfb289-b232-4559-bb1a-f4a08fca04c0 · outbound

This paper cites go straight forward.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation go straight forward

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:38:49.760921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:47.711590Z digest=sha256:09aad3ace2a755141b0445c032877036cd272e083b9706989387f1013bd199d1

Observation 496eae92-decf-4a09-8c35-75b441213256 · outbound

This paper cites 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.444570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:47.858650Z digest=sha256:d8664d7955aca1686bb7a5ece974ad8b2df1bdc4478d9b561310bba29a2d9136

Observation 8c2c5b16-bcd8-4a11-a3dd-437c8fde576a · outbound

This paper cites 10, we visualize more planning results on WOMD-Planning-ADE.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 10, we visualize more planning results on WOMD-Planning-ADE

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.104731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:48.010515Z digest=sha256:8a312d50db9c0b245462112d012f1ee3a706c297407a9515f096ecc36814c01e

Observation 2256c5ff-8d20-4b75-b615-7ac5f6e3322e · outbound

This paper cites Camera configuration.We apply different configurations of camera sensors in Tab.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Camera configuration.We apply different configurations of camera sensors in Tab

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:48.812486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:38:48.142792Z digest=sha256:39bdfd9d1f03ab1f48defce17a5caf2c038cf3f82ea3bc22222d96b9d386e97f

Pith citing papers

No inbound Pith citation observations are available.