Pith. sign in

Paper Citation Record · LEDGER

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2508.13470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13470 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:06:25.327098Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.379714Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1ccbd97-d541-4f94-8c02-1e42ddfd6af8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:22.867967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:22.867967Z digest=sha256:f9b56666df0d47775a0f1e21462eaa749daabd697b067c3a7c8a00a70b47b0cf

Observation 273c710f-7f77-4bba-87e7-804e49fe16e2 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen2.5-vl technical report, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.641597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:22.954759Z digest=sha256:f4c309b53ba114f7b382957f45b97f25f1f2e4117e21c2277c2028c4b6d05504

Observation 2ee33d05-da6e-4be0-89fc-4f0f0d3a715c · outbound

This paper cites Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.503846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.004322Z digest=sha256:d10d6a0510292ec08a27cfde461d2a1f83d64a0586b90c96e3b70e1db94e610d

Observation afc6916c-e97f-43e6-b66b-f4a376c76f49 · outbound

This paper cites Cityllava: Efficient fine- tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine- tuning for vlms in city scenario

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.338854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.047424Z digest=sha256:904de88d67c8cd30d33839c187e023b7a79d335a8f7387139ce1fda4be0fd487

Observation 97a332a3-6fda-4503-ab1c-039cbffc7f9f · outbound

This paper cites Cityllava: Efficient fine-tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine-tuning for vlms in city scenario

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.215166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.121343Z digest=sha256:947a1ee8cf8ba42c1a43aff9dbb165eb2c7fb223be74293fe3e00b3a9ae97271

Observation 683612bf-d7a1-485c-b5c1-8417c7c23426 · outbound

This paper cites Optimal gradient checkpoint search for arbitrary computation graphs.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Optimal gradient checkpoint search for arbitrary computation graphs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.085557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.188507Z digest=sha256:a43eb0a5728e549d7653a5975449e680f10aa83c49f45b073bd1040ae4559eba

Observation 73c536d8-c260-472d-896d-4d0875328523 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.871406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.240284Z digest=sha256:742939d436c6e0766605701191359c83ca0c8c3919ba05fa341c88250824d9fd

Observation cc28c621-9006-4e38-b4a2-a4e01cc313bb · outbound

This paper cites Better zero-shot reasoning with role-play prompting, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Better zero-shot reasoning with role-play prompting, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.728846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.340402Z digest=sha256:bf2c1f690182b58b7c07e1c92c1d1e7174828780e3204482d8e0a58e9384e5ea

Observation 18c180c3-c576-499b-b6df-f33e92a77418 · outbound

This paper cites Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.580566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.441211Z digest=sha256:c3324117efc1a3f6b33531c456b60a3723caa15d87f940886f8ba683025a36e7

Observation 323ef9e9-47cc-41dc-8691-b32fc052179e · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.616871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.616871Z digest=sha256:9fcded64c619453a23f65c71be2aa9dcd9e9efee9bc420bec215aa7e93d2fa17

Observation 4b115999-15e5-4575-8c98-d5ceb4a12853 · outbound

This paper cites SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.701939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.701939Z digest=sha256:2dd15d4966fc2be21a79a5d0f228503e192d06d38ccab75a7778fe43628e44e6

Observation 56de8027-29a1-481e-8680-aa6ead3cb860 · outbound

This paper cites Improved baselines with visual instruction tuning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.256587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.781156Z digest=sha256:010402e7a47eda2a102563ed2a08bb0e990186bfeb2aecce36d4f6adf7cbaf2f

Observation 7f6b4a87-5206-469f-9fef-79ed1d70dd9d · outbound

This paper cites Improving generalization in visual reasoning via self-ensemble, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improving generalization in visual reasoning via self-ensemble, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.857493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.857493Z digest=sha256:ec4e4586ac6f9964cbd866f1e5a9ee11cb46027a3f5c010fae2ad73d0edb25a7

Observation 6fd23e2d-cd9c-40d0-9a4d-56cb10567450 · outbound

This paper cites Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.124916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.920763Z digest=sha256:3d032e68a62613abbfac124744448d012d2e583d7ff4d065aef82b173719f7f1

Observation a6764805-6aa6-4800-b0fd-0c01c7d9b09c · outbound

This paper cites Le, and Quang-Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang-Vinh Dinh

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.921115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.979582Z digest=sha256:069f8470705180cc1d532ba0a35ab561053c6d9e723f0baf95afb49a9c41de03

Observation 87a592ff-d671-481f-8425-e07c624a32d4 · outbound

This paper cites Gpt-4v(ision).

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Gpt-4v(ision)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.797475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.016832Z digest=sha256:c97ec1b4b9ed5fc1a10bb562908d8c7a41344526f0fe924e07e475c30085d792

Observation b69b4d08-62d1-41bc-95b4-b66e6049efc0 · outbound

This paper cites Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.678940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.075072Z digest=sha256:596fe4176eaf9cd9d864ac6098fa1cd006946e6b93428511872c7a0cfda83fe9

Observation 5b6bc513-0163-4956-85fc-894ebf82f210 · outbound

This paper cites 1, 2, 3, 5.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models 1, 2, 3, 5

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.401999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:23.527136Z digest=sha256:1769a9e9b2f3a90855df074f01f0fb629da0b0391819bb135487e4e12443f0f4

Observation 6f9b1c7d-cf42-4e37-a67e-0d548291fcac · outbound

This paper cites Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.509429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.137473Z digest=sha256:3ce79d7d7c7ba499d6d1480c00556a5152e48718b26b9d38c2b720519bb2a03b

Observation 6184360f-bbac-4af5-ab0c-a8d93fdd0657 · outbound

This paper cites Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.365440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.221004Z digest=sha256:a33dad43e83ad766149efcf58b9ea00b0067f23bc36ffc22317e1706cc261539

Observation 84a8b82a-efdc-4e8e-b6b4-d69547a4cf23 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.181553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.286445Z digest=sha256:08ccb0bc01744cc817d1612ee15a96b18a2d9c51f810fd3d90935180e785eb99

Observation 7fed3d03-dbc0-40ed-b51f-eec4ec332711 · outbound

This paper cites Le, and Quang- Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang- Vinh Dinh

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.981603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.346433Z digest=sha256:e0dbcdd7e2ac848c08acd531892d0f0046ac19760c467fc412df4dc8912dd854

Observation 8977a4f4-d9bc-4267-9041-e6ab6ab98974 · outbound

This paper cites Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.832408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.381382Z digest=sha256:98d619fa9da3a70b95e4af4b533a6e3cb4aeef48e5e1ef8ab5affc30c083873e

Observation 0cf37018-ff54-4d37-8b33-bd3c090bbe75 · outbound

This paper cites Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.668464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.470008Z digest=sha256:452b71ea1355cbd29ca8794599777970d3ec50ba01be9ed16fbeb1ace8e9fb24

Observation 23da6946-b704-42ff-8433-1fadd6fe12e4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.496382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.572649Z digest=sha256:dc21f596cd1f1b73af6e6510729bfa8f8764175f002912250cc125794ec93351

Observation ebbe07ed-b501-4724-9d36-7ff24729e2b8 · outbound

This paper cites SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.379777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.632142Z digest=sha256:d109b5ea035cf3b8ef9d6a44b4c3f378acae577bf90c389bdf8fbd98504005f3

Observation 117f22c4-c441-4ad8-9258-a6695f8841c0 · outbound

This paper cites Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.263756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.730131Z digest=sha256:e0ae73eb45e60a697d90fea18a36bd181c70d05fa2596f149e8cec98512356de

Observation 688e393a-b896-4e19-b5df-d7386dda7b8a · outbound

This paper cites CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.771980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.791229Z digest=sha256:d7526b4c24e0d66fa3b5ccea1b7b20944cd19c1e48cd7aebf8ef7d08a1c0fe24

Observation ea8fc92d-7dbb-49f8-9275-9a6d9b86247a · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.105600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:24.856071Z digest=sha256:76af8f90fcac0e9c3a68ed55dad1c7488d8e932a4bfc294ee03890404080a07f

Observation 84d2afda-c73a-4913-a5ae-9c33b2006d0b · outbound

This paper cites A Study of Situational Reasoning for Traffic Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models A Study of Situational Reasoning for Traffic Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:24.973033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:24.973033Z digest=sha256:59af969e1ca54dd5f6a362a2be5e5e7067a433771ff9059968fdc75924757864

Observation b0cbf26b-8b6a-47ac-92ab-18eea1b1bb09 · outbound

This paper cites When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.069634Z digest=sha256:8dd665e693192aeea20d786b4ee6afc2853b79263401a0f883dee162beeef777

Observation 41fe9ed6-f67d-4514-a3ac-efed3ca5e00f · outbound

This paper cites CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.539818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:25.134699Z digest=sha256:ac746b37184a5b4a78c15bf26badf56c168c6e42ac6527c4044a61b433a23cca

Observation e66332e3-e7c8-49e4-b33e-672f422eade7 · outbound

This paper cites TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.219559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.219559Z digest=sha256:af19cd4e950d88b2f002bebc63fc3637993154973f6d67045dede4bf4e7b8c07

Observation 3610eae5-e259-4029-b6d7-b2658d72b531 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:25.927402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T19:06:25.327098Z digest=sha256:8eeda7d19b66471d3051036ebcc1b33e8f7a464916800f0ead381c64491ac905

Pith citing papers

Observation ca2de7d7-d1d7-4c89-9448-115a09479bea · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:0290f59c52cdb432a223119c890c8144568a06c56078cac5a5dfd77405e1daf3