Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:08.780881Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2412.01132.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:08.780881Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:16.204012Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T00:15:52.849428Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 129e97fc-118e-48b7-a6ed-537ce680ce53 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering: Datasets, algorithms and challenges, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfad524f-48b6-47a8-843c-d10e471f9518 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videoqa in the era of llms: An empirical study, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 44aedcde-13e9-4ce0-96ad-770869e13167 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Francis, and Alessandro Oltramari
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f603df1-b510-4b89-9090-6979884feb52 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Carla: An open urban driving simulator, 2017
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4761f4e7-63f7-48f4-8915-06f870a3c798 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Lingoqa: Visual question answering for autonomous driving, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 20c7a9e5-ca6e-4f0d-967c-a60cf68fbd61 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c9995d22-7b77-4980-8a48-530f2f7d47e3 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks GPT-4o: System Card
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 872b3005-0b4b-4fe1-a58a-d34a9b1d47f3 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5248eca9-14de-4926-9522-8ba932fb945a · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Recognizing an action using its name: A knowledge-based approach
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1f6253e7-a2ae-458e-8b2b-88901f668cc4 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video representation learning with deep neural networks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f79878ee-26e8-49f2-a65b-bff1bcd36640 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering via gradually refined attention over appearance and motion
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 288832a7-3821-4f8b-a103-91085ed142b5 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Vilbert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b4658756-8c62-4aa0-967b-6cc2596ce35a · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Compositional attention networks with two-stream fusion for video question answering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eff6c892-b8f5-4b73-bede-f3b5f70725ca · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Visualbert: A simple and performant baseline for vision and language, 2019
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e93753ee-4fb1-42ea-af3b-7a21806dd383 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Activitynet: A large-scale video benchmark for human activity understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c73926d5-4343-4d0b-af80-0c22f0799a30 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Tgif: A new dataset and benchmark on animated gif description, 2016
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d00888a-1c89-44a0-af8f-a916f2ee03a5 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Next-qa:next phase of question-answering to explaining temporal actions, 2021
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79cac09c-c190-41f0-b3dd-f544ee9bf5d8 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gemini: A family of highly capable multimodal models, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 08f1446b-302c-487e-971f-60fdc007da69 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video instruction tuning with synthetic data, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c726ee96-410e-46ec-9cbb-1a3f447843dd · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Ego4d: Around the world in 3,000 hours of egocentric video, 2022
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9d99f4f-780b-4100-9884-961ae8c0217f · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7573b07a-f401-4f63-8698-9066dda52739 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4d09cf-3d89-4d80-96f0-5dc03a22fba5 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mist: Medical image segmenta- tion transformer with convolutional attention mixing (cam) decoder
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 757f607f-27dc-47d1-b85b-0de084d1c596 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A multi-world approach to question answering about real-world scenes based on uncertain input
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5da1ab1b-8c51-41a9-b746-be81af226e7f · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b0ec725-2eeb-40b2-bb12-70ed4eb13ed5 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d18766e-e981-475a-bb57-8f40c2e22232 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Argos vision: Advanced computer vision solutions, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 702591f2-84be-46f0-8228-b9617433be06 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Can i trust your answer? visually grounded video question answering
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 61059adc-ffe0-4a6c-9a91-a63ae3a9cb7b · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc40cf6-3779-4f84-adfb-f31fc1f4023c · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9366bec2-442a-4b46-9588-cd422ecda1d4 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2562f93c-e1f9-4275-88af-c45b25505be9 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534606d4-8008-4126-b81e-c515c2655036 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc90ef3-4729-45c5-a76b-58a6d51ae024 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Improved baselines with visual instruction tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bcfcc618-7819-43ac-b05d-3f39601e250d · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09fad8df-814b-4bc4-a41d-6e3584624b82 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609fe016-13b9-49e0-91f5-1451828ffdfa · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM2 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87acb9a-73b4-425c-936a-2e63e780a536 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185f6007-f6c3-4765-aa97-855bd88e3305 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Minesh Mathew
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2ac603b-7010-4854-ad84-de5d025caa2d · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2987ac11-a28d-4c06-8098-a0ae3a7b7454 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Infographicvqa
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 59ea8c61-fff6-4480-89f1-6123798168da · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Towards vqa models that can read
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a21812c-f164-45cd-97ac-fd6c8ce900a6 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0bb42d2-7d75-4433-9fe4-2636430cfaee · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f90c4b2-758f-4fed-ba1c-122d09534858 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Learning transferable visual models from natural language supervision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 824b2fb0-b4b9-43eb-ae50-573ec8937a95 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Beats: Audio pre-training with acoustic tokenizers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 40d87ff2-f9ca-4180-a1f5-dfbe488898d5 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8fb43b-7709-4f52-b0a4-a0b2a3109792 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd311754-fd02-46e3-b8ed-c8f7b2ddd62a · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0de5971-23ab-4a76-a6ac-dc29b659efeb · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bdeeb4e0-ac6c-41f9-a072-332ba319b5f3 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b27448-3032-4290-8177-014191844ef8 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f6484ece-c312-4ca4-8b39-8a7d9917b913 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac9974fc-3fc7-408f-802f-d469af7619dd · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683195b3-4202-4a26-8627-2b7242cd4580 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gonzalez, Ion Stoica, and Eric P
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59587831-6f21-4d66-b0dd-7251b5939791 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce52ce43-89c2-4ffe-a81b-12acc45d3999 · outbound
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How many cars can you see?
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f316b1d1-7c9a-494e-96fa-35b68ef33342 · inbound
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.