Pith. sign in

Paper Citation Record · LEDGER

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2508.13470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13470 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:06:25.327098Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.379714Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1ccbd97-d541-4f94-8c02-1e42ddfd6af8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:22.867967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:22.867967Z digest=sha256:854de3123237feb5119d97b86c73548b13cfe1eb9d6e9478b17a3e6d45b3059a

Observation 273c710f-7f77-4bba-87e7-804e49fe16e2 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen2.5-vl technical report, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.641597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:22.954759Z digest=sha256:2158496be83458b20406c6045b08f7439f6da9a9aee9b65c4fac3aa9b956df68

Observation 2ee33d05-da6e-4be0-89fc-4f0f0d3a715c · outbound

This paper cites Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.503846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.004322Z digest=sha256:1c195f8aa4c73ae82aed077b65c5f441e3497f3676e0367bb97be805b8b8df1b

Observation afc6916c-e97f-43e6-b66b-f4a376c76f49 · outbound

This paper cites Cityllava: Efficient fine- tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine- tuning for vlms in city scenario

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.338854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.047424Z digest=sha256:8ef361d8125c62bd35f7e295f136d31c3a697e7cfa189fb0c9abd33b0b3f0582

Observation 97a332a3-6fda-4503-ab1c-039cbffc7f9f · outbound

This paper cites Cityllava: Efficient fine-tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine-tuning for vlms in city scenario

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.215166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.121343Z digest=sha256:db4644a55c7cb78dbd7e2e909e309eaa0631a8d10740859e2068f04986e77141

Observation 683612bf-d7a1-485c-b5c1-8417c7c23426 · outbound

This paper cites Optimal gradient checkpoint search for arbitrary computation graphs.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Optimal gradient checkpoint search for arbitrary computation graphs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.085557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.188507Z digest=sha256:9d47bb7a3794214cff367e326ae3e5a42cce93e586feb823adc92432f56df07f

Observation 73c536d8-c260-472d-896d-4d0875328523 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.871406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.240284Z digest=sha256:d9e344fa0d4bbe020f3efc5b499fe16a6d1ca54eb002d704a64390bd46b6fbd2

Observation cc28c621-9006-4e38-b4a2-a4e01cc313bb · outbound

This paper cites Better zero-shot reasoning with role-play prompting, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Better zero-shot reasoning with role-play prompting, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.728846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.340402Z digest=sha256:db966de84dbad9573a21121299f5fe7284b854dd2f43dcaba5db1db4c6ef29d1

Observation 18c180c3-c576-499b-b6df-f33e92a77418 · outbound

This paper cites Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.580566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.441211Z digest=sha256:0fbeb53e0526628d2831c047ce6ec2165869198d085bdb3dec680e39b53a7318

Observation 323ef9e9-47cc-41dc-8691-b32fc052179e · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.616871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.616871Z digest=sha256:cda5e77def86984dbcc75ce0e62b921bafd9f853c9d8c0f33f78fc110c7ee89e

Observation 4b115999-15e5-4575-8c98-d5ceb4a12853 · outbound

This paper cites SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.701939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.701939Z digest=sha256:1ca400fd280c89b197d639a2a1fd13d60e4302a6c9352df6ca06a780b9115a62

Observation 56de8027-29a1-481e-8680-aa6ead3cb860 · outbound

This paper cites Improved baselines with visual instruction tuning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.256587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.781156Z digest=sha256:48ae49950d34df2e6c31fc8acdf6fab022c8b4526ad232a8af22292ace0f496a

Observation 7f6b4a87-5206-469f-9fef-79ed1d70dd9d · outbound

This paper cites Improving generalization in visual reasoning via self-ensemble, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improving generalization in visual reasoning via self-ensemble, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.857493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.857493Z digest=sha256:ff7aa39d4d8ebad3c351881fa633c84343e5f72b848a6c5810fe82bc8d5d270d

Observation 6fd23e2d-cd9c-40d0-9a4d-56cb10567450 · outbound

This paper cites Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.124916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.920763Z digest=sha256:45c850c584c236b5c0154f1fd42980d1cab41687d3bebef5fdaaf8a7c103a777

Observation a6764805-6aa6-4800-b0fd-0c01c7d9b09c · outbound

This paper cites Le, and Quang-Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang-Vinh Dinh

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.921115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.979582Z digest=sha256:de1b4029b6a1fbfcb931e78615c43a7c1b3ed0d3bb5326e62e6a1a62532e26a1

Observation 87a592ff-d671-481f-8425-e07c624a32d4 · outbound

This paper cites Gpt-4v(ision).

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Gpt-4v(ision)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.797475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.016832Z digest=sha256:fa51e2de523af8d04b0e09c4255fd25e5f1c41e89dadf2b8bcd6455e6c511e59

Observation b69b4d08-62d1-41bc-95b4-b66e6049efc0 · outbound

This paper cites Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.678940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.075072Z digest=sha256:f3b404797484d0f3bd29daa3dc76579333984cb51688c5bd28d6753f013554dd

Observation 5b6bc513-0163-4956-85fc-894ebf82f210 · outbound

This paper cites 1, 2, 3, 5.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models 1, 2, 3, 5

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.401999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:23.527136Z digest=sha256:ec88f8ea8fa6cfb221c84834089f02069e2b861c0f26214a72abb80be40ce7c1

Observation 6f9b1c7d-cf42-4e37-a67e-0d548291fcac · outbound

This paper cites Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.509429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.137473Z digest=sha256:57b2510182a80f522a2ef2cc317351070951fcb0d7014f5aaff04db1308e35af

Observation 6184360f-bbac-4af5-ab0c-a8d93fdd0657 · outbound

This paper cites Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.365440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.221004Z digest=sha256:880c65db8c55c16fd79b46738e1278f4d3ba9efdfdfde0a5408c1ab3d7478a93

Observation 84a8b82a-efdc-4e8e-b6b4-d69547a4cf23 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.181553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.286445Z digest=sha256:e6571cf4333b8166168a93f6f95421262d91aa0db25020386454e53cec789874

Observation 7fed3d03-dbc0-40ed-b51f-eec4ec332711 · outbound

This paper cites Le, and Quang- Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang- Vinh Dinh

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.981603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.346433Z digest=sha256:693e94306099392dc97a66ac814b7d4ddbfbdee51de5cdaa6a4c498d2b899952

Observation 8977a4f4-d9bc-4267-9041-e6ab6ab98974 · outbound

This paper cites Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.832408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.381382Z digest=sha256:3e81e5adfc86c5a169070e6b5a2d5f5e592e18895960d89e38337533a1c65691

Observation 0cf37018-ff54-4d37-8b33-bd3c090bbe75 · outbound

This paper cites Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.668464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.470008Z digest=sha256:ec49a2d625581f9c25f8da336f65a14b5521f2a0410ec58189c9a088523f3541

Observation 23da6946-b704-42ff-8433-1fadd6fe12e4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.496382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.572649Z digest=sha256:fd1be9c4848dc5a24ab7ee4182a495bae68190e3da3b882e906a4c5aa2c46558

Observation ebbe07ed-b501-4724-9d36-7ff24729e2b8 · outbound

This paper cites SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.379777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.632142Z digest=sha256:a1cd838e196fe0643fab7f1d14d4770ad182ddef70f111a9487c3bbe0554b8b5

Observation 117f22c4-c441-4ad8-9258-a6695f8841c0 · outbound

This paper cites Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.263756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.730131Z digest=sha256:bc2f24cbdea9dc4413f3fc0dddbd0f9e782ff78a49b95159a9f213eaf7deef6c

Observation 688e393a-b896-4e19-b5df-d7386dda7b8a · outbound

This paper cites CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.771980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.791229Z digest=sha256:c6c4dbf4a11269980ac29ce2d1f74e01d08003a6510b9204242e32edbabd084e

Observation ea8fc92d-7dbb-49f8-9275-9a6d9b86247a · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.105600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:24.856071Z digest=sha256:cf9372c543afdc090c1691b00a6fb32cb9cbcd6bfe49189b2494383d53865343

Observation 84d2afda-c73a-4913-a5ae-9c33b2006d0b · outbound

This paper cites A Study of Situational Reasoning for Traffic Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models A Study of Situational Reasoning for Traffic Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:24.973033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:24.973033Z digest=sha256:b5ba0d5fef16edc1b30047edf47bea93f2942e48aadda4cfee5f4cfa8dc7d91d

Observation b0cbf26b-8b6a-47ac-92ab-18eea1b1bb09 · outbound

This paper cites When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.069634Z digest=sha256:a3816085024125c16083c961b1bebae6796c0f3e6f4d995978fef3816b7cc779

Observation 41fe9ed6-f67d-4514-a3ac-efed3ca5e00f · outbound

This paper cites CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.539818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:25.134699Z digest=sha256:e01481390f9d44ab74a131e12ec873425e4ec8fbfee4f689cfd9a0df9331f428

Observation e66332e3-e7c8-49e4-b33e-672f422eade7 · outbound

This paper cites TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.219559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.219559Z digest=sha256:2a5e8f73a933023553af3b47852d31fdd8ac3bfcba14eda765f6054a91cf49b9

Observation 3610eae5-e259-4029-b6d7-b2658d72b531 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:25.927402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:06:25.327098Z digest=sha256:503fb38e2fe390a6923c5fa15daaeb576ab83c3d526f34fc397152c1af00effd

Pith citing papers

Observation ca2de7d7-d1d7-4c89-9448-115a09479bea · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:e7260b6fdec32495f5bef535e2c55303863cd8ca576da67de4f7296afd824c8e