Pith. sign in

Paper Citation Record · LEDGER

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2505.12670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12670 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:52.291782Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:31:04.357239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T06:48:01.003547Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a18a2696-e94f-4aca-8fda-e6786c62b45b · outbound

This paper cites End-to-end autonomous driving: Challenges and frontiers,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning End-to-end autonomous driving: Challenges and frontiers,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.111566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.111566Z digest=sha256:128dddcb4cad246a515b0f1d7bffd05fdb9f921521ef6246681181ee337e093e

Observation ba0d960f-0af8-4ea1-8860-e78a3ff44a8c · outbound

This paper cites A Survey for Foundation Models in Autonomous Driving.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning A Survey for Foundation Models in Autonomous Driving

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.117266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.117266Z digest=sha256:252204c30e65b09ce166d0d3cdd8705fced6f19068f54a120c16b9ea4bf65eea

Observation 189e9e1a-09d9-4c4b-9131-6eafe3977ea0 · outbound

This paper cites Vision+ language applications: A survey,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vision+ language applications: A survey,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:53.017879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.122656Z digest=sha256:21bca6fac5b616610b1d9e946311bfd322260d367d59dcf9efed8b999bd77f63

Observation 605a68a2-84aa-4a9e-bb01-de143d5b5f47 · outbound

This paper cites DriveGPT4: Interpretable end-to-end autonomous driving via large language model,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning DriveGPT4: Interpretable end-to-end autonomous driving via large language model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.127318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.127318Z digest=sha256:bd97c109b2c292a5c4bce8ce5536c700c11a1b8fcf3a2da3fbf61a5741b95b77

Observation 73c4f705-89c5-4b84-896b-d6f35abd379c · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vision language models in autonomous driving: A survey and outlook,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.991645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.132489Z digest=sha256:dcd2e4e60d4c7a4b8edea456f0022bd256884c40d857c6adcca14b6baf541647

Observation bddb246b-8ea8-4d90-9ca9-655402259a73 · outbound

This paper cites When do we not need larger vision models?.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning When do we not need larger vision models?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.976621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.137613Z digest=sha256:66c23ffb56e52f5472e7fb172fd4d351c63d89b747e9a3d097b87f1b371a3db2

Observation 0ed303bb-cf76-44f8-979a-c62e2578acd7 · outbound

This paper cites Separable Self-attention for Mobile Vision Transformers.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Separable Self-attention for Mobile Vision Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.143208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.143208Z digest=sha256:ff0c9b27d460c929fa36e47d8d0a6a46b33f087aec2a362966ce78bc1b239fe4

Observation 76b7a076-3c3e-44ff-9516-602c024403d8 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning DriveLM: Driving with Graph Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.148098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.148098Z digest=sha256:7faa1e84cb17dd9b586d07e4cd243660b07b0bef7d71847bde9c2c416bc51a75

Observation f3c59b47-b36a-4fe5-817b-5a3d77ac8231 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.153109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.153109Z digest=sha256:b748db94c1c0e016f6f8304e2865758e999d9a725c2237f6e4dc1767448138b1

Observation da4ff659-68f5-4189-a91e-dcc344ee8268 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.157711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.157711Z digest=sha256:5ae82b1b6b5950a2d97e321a715753a0d58f1b91f495e8ba15c9a6b47c502911

Observation 8fad5269-0a4c-4390-860d-2f0a6e416878 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.162848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.162848Z digest=sha256:15d527006cdd6ca15eee58532f44934ea0173476247497f2e1be7b53071bd84a

Observation 28847b15-ef5a-4bf8-aa9f-9e9477df5776 · outbound

This paper cites Scene graph refinement network for visual question answering,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Scene graph refinement network for visual question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.951780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.167841Z digest=sha256:5a1ac064366e633e8c89ead32724a0c0fe08fa31eff83d85332cf1ec6dc2a5c8

Observation 38ad1e65-ca51-43e6-965b-2405c02f1d40 · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.172544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.172544Z digest=sha256:a7bb86232cfab1a2f759880da28615e2e06976815f127152b784bf55c4b68ded

Observation aa2cf9eb-8b24-4aa1-855e-8e77cc4217e8 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.177381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.177381Z digest=sha256:49d6d94c1d0ad264acd3d41f25014a3e18f5322360cc0d072d830f0360b45867

Observation 49d1ff74-fb13-4c9e-872f-e97568eea4b8 · outbound

This paper cites Multimodality self- distillation for fast inference of vision and language pretrained models,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Multimodality self- distillation for fast inference of vision and language pretrained models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.927111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.181992Z digest=sha256:2a0fb21c05f8482c38e501362b56bad8109b11646ab8ab59b1d5dd12404cf689

Observation 47516cda-4c33-4530-9fc4-9773dfd08b5a · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Flamingo: a visual language model for few-shot learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.186562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.186562Z digest=sha256:bfaf0a1c3d72ad7b44f6aa55903826517e55ca43dc28d0bfadc71fc0e5205b51

Observation 26b17a0b-bb8a-4f11-a23e-959351b913ca · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.191121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.191121Z digest=sha256:16f697bb8e58b2caa6cd1c94792a46d5e162c6f0e983857fb2ea4621bc14ac2b

Observation 1b15d4e4-32b0-4ee2-aae7-e18279db2980 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning OPT: Open Pre-trained Transformer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.195567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.195567Z digest=sha256:feef5a0f097f29719b624ea437a29b2db6695c26a08c1ca28f151efc5bd50795

Observation 682b9e78-b76d-4b55-93b0-56bfd9d2fb00 · outbound

This paper cites Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.200530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.200530Z digest=sha256:c2613e327c0c52dc239ba938d95e610b07a16207263509a44213c4db9ecc0d47

Observation c76450d2-bfb9-409b-b2f5-5b6719bd3410 · outbound

This paper cites Flava: A foundational language and vision alignment model,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Flava: A foundational language and vision alignment model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.882863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.205093Z digest=sha256:5818492e9196dc9335581760111134e29b367d76a0df7d60ea2508264f593413

Observation 862d9f8d-d441-4457-a556-3fd7efc80e16 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.210012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.210012Z digest=sha256:d5938eb8a571afa11b87c5d750631d6a8f5b7d1532e26f474e1aca09838c37d6

Observation 50477efa-e9d2-4ae9-8b7a-3c2076f6a8fb · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Openscene: 3d scene understanding with open vocabularies,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.214494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.214494Z digest=sha256:044f7b537c14d00178a03b3fe60de790654b20f5436316e3d96e110f1e79118b

Observation b603bd1f-0935-473c-a204-59643f940dee · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.858095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.219508Z digest=sha256:054ce612d8fd7cdeb376eae57b3bd98430af6b3534de25c7a67f47515d5f1961

Observation 3c5ed886-96cd-447d-8cea-1db146b14ad6 · outbound

This paper cites Vldadaptor: Domain adaptive object detection with vision-language model distillation,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Vldadaptor: Domain adaptive object detection with vision-language model distillation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.842891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.224171Z digest=sha256:8c7be8e3389d01438ee5d9f947dd6b7b90d10db33123bacf487bb7086337748a

Observation e4247254-dbce-4dbe-bfad-f97d89673f77 · outbound

This paper cites Anomalyclip: Object-agnostic prompt learning for zero-shot anomaly detection,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Anomalyclip: Object-agnostic prompt learning for zero-shot anomaly detection,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.228335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.228335Z digest=sha256:2524014fe071be3c161b7a02b8c290e2a366b4b14095e2aac04ed8f7e4ce6e08

Observation ff7688d2-cdee-47f8-b26c-520eb55e8968 · outbound

This paper cites Semantic anomaly detection with large language models,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Semantic anomaly detection with large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.826461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.232771Z digest=sha256:0316fd499ae5c66dd2abd96a1e9884c7c5f45fe8ea49c451ba86f37ca852f3b7

Observation f7a1ad8e-3648-44e4-85a4-01a59e229244 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning GPT-Driver: Learning to Drive with GPT

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.237289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.237289Z digest=sha256:64a91f49d3f8afb9c7e0408c996987f687f465623ffc6063a145e93f1e8fab65

Observation 273436b4-d44e-430c-84ab-4a4c6f0d9d69 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Lmdrive: Closed-loop end-to-end driving with large language models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.241971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.241971Z digest=sha256:95eeaece061295d6615503591ad2c2a19d7ad83704b4004eb2ceae94208ec387

Observation 47be580d-563c-4e01-bff0-0d3c85c22457 · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.247011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.247011Z digest=sha256:5834aeb42e8fe0c520e9024538c7042771df26b5d2f6991e1dbc4c5fa7e34471

Observation b6847050-a82f-426a-804a-39ab590f745b · outbound

This paper cites Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.251636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.251636Z digest=sha256:f3452f27ba1bf4c0900f6348c7cc95a458f3e40eb446bf2852a272e6c4f31fe8

Observation db55f0ca-cb8a-47e4-924c-73b442dd698d · outbound

This paper cites MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.256242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.256242Z digest=sha256:82f30a703ceee8307b892bf1b1afa07e195533d970a16bec87e5fd12d9c536fb

Observation 96c8a9be-9f40-47fa-93b7-656ab242e9fd · outbound

This paper cites BEV-CLIP: Multi-modal bev retrieval methodology for complex scene in autonomous driving,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning BEV-CLIP: Multi-modal bev retrieval methodology for complex scene in autonomous driving,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.800453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.261175Z digest=sha256:0c33edc09aa6018a1dd9824dfaae95cf23c118d3995b634e9a84617088158e60

Observation 2a1e0d84-50fc-4a71-9d8a-fb34f96cc3ca · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Bleu: a method for automatic evaluation of machine translation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.265475Z digest=sha256:919e0431f9a325607b2c6aaad3c29fbe2170b5dc2b62eddbb192b3b249bddf64

Observation 2fe5dbac-a806-444e-bfbb-1baf55cee1fd · outbound

This paper cites METEOR: an automatic metric for mt evalu- ation with improved correlation with human judgments,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning METEOR: an automatic metric for mt evalu- ation with improved correlation with human judgments,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.774414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.269773Z digest=sha256:d40c491f798a7a8f0d9fc7a8b702790eed1fbdbfa486e198b23522df92a17eaa

Observation 8262256f-858d-48f5-8f0c-39c311210d09 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Rouge: A package for automatic evaluation of summaries,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.274008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.274008Z digest=sha256:b15f181685fec49b40a44ccd12da637cbc1b35b84d51b15f8f13523bdf2f1327

Observation a8600a1a-9d85-4567-986e-57154939e31b · outbound

This paper cites Cider: Consensus- based image description evaluation,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Cider: Consensus- based image description evaluation,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.278309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.278309Z digest=sha256:c724d39c66d453ddf9348ffdb8c8680af8ee5e85556515055599f183e453c0b3

Observation 35677d33-0331-4b6f-9650-1ac351cdae70 · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.283010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.283010Z digest=sha256:949b4c1852de125e141196a0576d8174a51f417fa819f7e77f50875f46de5ea9

Observation ca07e764-599a-468e-899d-1e7d92c95108 · outbound

This paper cites Driving with LLMs: Fusing object- level vector modality for explainable autonomous driving,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Driving with LLMs: Fusing object- level vector modality for explainable autonomous driving,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.740459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.287524Z digest=sha256:0f7463324d59fa9df4e1a3cd03f9dcc838b4711997909155c07ce64b75c6cbfe

Observation ddeb0211-a43c-476e-af80-cb22c490c9c9 · outbound

This paper cites Learning latent per- mutations with gumbel-sinkhorn networks,.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning Learning latent per- mutations with gumbel-sinkhorn networks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:52.724730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:32:52.291782Z digest=sha256:738e89ae1a4a0d080cc4972793337b59e55fe4f4b59c65576670524e9cc3db2d

Pith citing papers

Observation 88023e48-d173-4a4c-a9d3-d4966c1b92f8 · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.357239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.357239Z digest=sha256:35355124f4c861cf6eea833d2e6388594db0a54f6ae8ee24efe4958d8a62a0dd

Observation bfcf9d86-99c1-456d-b038-bd242dcb871c · inbound

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving cites this paper.

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:48:01.006961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T06:48:00.943591Z digest=sha256:d23da13b95ac857dbac7abedd5ca3e4ac843a8670c6874a39e828ce119df8f50

Observation 627aa47c-1cfa-4c1e-9669-aa899392543e · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:28.631229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:48:36.717026Z digest=sha256:cc779350695923bb32e4a92d2addc68ec8595cba194880afdf09fd020c6a3b90

Observation 33116493-0463-46f0-95d1-2259158c7857 · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.679725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:36:52.396245Z digest=sha256:df47faf41ae67cc07ca26eaf3c631a0c593a249202a922aaf4282ff3b9b1651b