Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:52.291782Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2505.12670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:52.291782Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:31:04.357239Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T06:48:01.003547Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a18a2696-e94f-4aca-8fda-e6786c62b45b · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning End-to-end autonomous driving: Challenges and frontiers,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0d960f-0af8-4ea1-8860-e78a3ff44a8c · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning A Survey for Foundation Models in Autonomous Driving
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189e9e1a-09d9-4c4b-9131-6eafe3977ea0 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vision+ language applications: A survey,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 605a68a2-84aa-4a9e-bb01-de143d5b5f47 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning DriveGPT4: Interpretable end-to-end autonomous driving via large language model,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c4f705-89c5-4b84-896b-d6f35abd379c · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vision language models in autonomous driving: A survey and outlook,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bddb246b-8ea8-4d90-9ca9-655402259a73 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning When do we not need larger vision models?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0ed303bb-cf76-44f8-979a-c62e2578acd7 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Separable Self-attention for Mobile Vision Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b7a076-3c3e-44ff-9516-602c024403d8 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning DriveLM: Driving with Graph Visual Question Answering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c59b47-b36a-4fe5-817b-5a3d77ac8231 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Learning transferable visual models from natural language supervision,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4ff659-68f5-4189-a91e-dcc344ee8268 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fad5269-0a4c-4390-860d-2f0a6e416878 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28847b15-ef5a-4bf8-aa9f-9e9477df5776 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Scene graph refinement network for visual question answering,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 38ad1e65-ca51-43e6-965b-2405c02f1d40 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2cf9eb-8b24-4aa1-855e-8e77cc4217e8 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d1ff74-fb13-4c9e-872f-e97568eea4b8 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Multimodality self- distillation for fast inference of vision and language pretrained models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 47516cda-4c33-4530-9fc4-9773dfd08b5a · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Flamingo: a visual language model for few-shot learning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b17a0b-bb8a-4f11-a23e-959351b913ca · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b15d4e4-32b0-4ee2-aae7-e18279db2980 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning OPT: Open Pre-trained Transformer Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 682b9e78-b76d-4b55-93b0-56bfd9d2fb00 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76450d2-bfb9-409b-b2f5-5b6719bd3410 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Flava: A foundational language and vision alignment model,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 862d9f8d-d441-4457-a556-3fd7efc80e16 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50477efa-e9d2-4ae9-8b7a-3c2076f6a8fb · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Openscene: 3d scene understanding with open vocabularies,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b603bd1f-0935-473c-a204-59643f940dee · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Clip2scene: Towards label-efficient 3d scene understanding by clip,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3c5ed886-96cd-447d-8cea-1db146b14ad6 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vldadaptor: Domain adaptive object detection with vision-language model distillation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4247254-dbce-4dbe-bfad-f97d89673f77 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Anomalyclip: Object-agnostic prompt learning for zero-shot anomaly detection,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7688d2-cdee-47f8-b26c-520eb55e8968 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Semantic anomaly detection with large language models,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f7a1ad8e-3648-44e4-85a4-01a59e229244 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning GPT-Driver: Learning to Drive with GPT
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 273436b4-d44e-430c-84ab-4a4c6f0d9d69 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Lmdrive: Closed-loop end-to-end driving with large language models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47be580d-563c-4e01-bff0-0d3c85c22457 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6847050-a82f-426a-804a-39ab590f745b · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db55f0ca-cb8a-47e4-924c-73b442dd698d · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c8a9be-9f40-47fa-93b7-656ab242e9fd · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning BEV-CLIP: Multi-modal bev retrieval methodology for complex scene in autonomous driving,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2a1e0d84-50fc-4a71-9d8a-fb34f96cc3ca · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Bleu: a method for automatic evaluation of machine translation,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe5dbac-a806-444e-bfbb-1baf55cee1fd · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning METEOR: an automatic metric for mt evalu- ation with improved correlation with human judgments,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8262256f-858d-48f5-8f0c-39c311210d09 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Rouge: A package for automatic evaluation of summaries,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8600a1a-9d85-4567-986e-57154939e31b · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Cider: Consensus- based image description evaluation,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35677d33-0331-4b6f-9650-1ac351cdae70 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca07e764-599a-468e-899d-1e7d92c95108 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Driving with LLMs: Fusing object- level vector modality for explainable autonomous driving,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ddeb0211-a43c-476e-af80-cb22c490c9c9 · outbound
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Learning latent per- mutations with gumbel-sinkhorn networks,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88023e48-d173-4a4c-a9d3-d4966c1b92f8 · inbound
A Survey on Vision-Language-Action Models for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcf9d86-99c1-456d-b038-bd242dcb871c · inbound
DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 627aa47c-1cfa-4c1e-9669-aa899392543e · inbound
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33116493-0463-46f0-95d1-2259158c7857 · inbound
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.