Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.853704Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.08455.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.853704Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b89677d-3b55-4f0b-b0fa-46f60380ab2e · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Robovqa: Multimodal long-horizon reasoning for robotics
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c878ab8e-0a4b-4286-9bb6-0c705479fbd3 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236615c5-5e79-4ada-ba59-d2b445ddf7cc · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 144edce9-8fb6-4610-bc43-0d037e06a919 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Openeqa: Embodied question answering in the era of foundation models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d80e3b03-69f1-440e-8910-a0d0c85dc28a · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f59abb2-0476-4a6d-a816-b4485f291313 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45022173-519c-455a-bc87-57177bbf8318 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6cb698aa-a7bd-4323-8b10-e43c707f888d · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c22944-33a7-4cec-936e-4a58d7efeb44 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b998e8-84cd-4ebd-82ad-3ab02563bc02 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50418e7-bf9e-4d51-86fd-1264daa20d9b · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning transferable visual models from natural language supervision
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cdff0e5f-ee66-4d28-a96b-3f5716613f57 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sigmoid loss for language image pre-training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 82f532e7-71c2-45e0-b07e-e60c41ae0a9c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96598b34-9e8d-4787-afd8-98e1cfe4e53e · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7270f04b-5650-4b02-bbef-ae2a4a921ff8 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605c03e1-405b-4afc-b7ee-1ac3815bde4b · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a418e650-accc-4d06-8ae2-41cfbd230895 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae912172-be3e-4f6d-ac41-6e33bfe252ab · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50caf98d-008e-4c68-b7f7-a27bc88dd2a9 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf4d981-c735-4225-bb99-f5f4764d190b · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bb79dd26-bed0-4924-b083-bc95c8cfb7b3 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfdc15e9-adde-4fe9-a13d-624b11eb2769 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5439363-9c59-44d3-a269-b8b7db28832c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578520e5-b3bb-4afe-9cc2-8bb64ce22f25 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c1aa7a-e644-415c-9837-9a358b3c61b3 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ebae743-b61f-48fa-b400-fbdbd49df6e4 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc00dbc-3b25-4895-8253-99514193ed0c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen-vl: A versatile vision-language model for understanding, localization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ed7d0b5a-a7ca-4ac5-b33c-f162c980f986 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5-VL Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2f1091-ee92-4f41-b884-c575b4d27450 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Intentqa: Context-aware video intent reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 852a0e58-1b22-4e8e-903b-0ef34fe7db90 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Rex- time: A benchmark suite for reasoning-across-time in videos
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95eb2e73-826e-45ec-ba3f-6e5aa90e6e69 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9ef2c91-79b1-4a9a-ab86-7279b66c08b5 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Long Context Transfer from Language to Vision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906e0ef9-b0e3-40b6-a7d0-d3b554e4d1d6 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Moviechat: From dense token to sparse memory for long video understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c725de8b-4d3c-42b8-bbb1-e77eda694019 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec3dd096-487a-4ce7-8741-b02ab8275c7c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a1d4b7-fae9-465e-a4fd-a4c44af26b47 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1f5f85-d232-4aab-8ea3-c116692f5b5e · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video understanding with large language models: A survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 920f45ec-7d6e-4994-a205-3e17288a260c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Foundation Models for Video Understanding: A Survey
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae3780d-0c83-40ce-a81c-fa06d730fe3f · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca70cb1-cdc6-4f87-93d4-b83935edb694 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video question answering via gradually refined attention over appearance and motion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a023add-8aef-4e6e-8b2f-07d2cf103131 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Next-qa: Next phase of question- answering to explaining temporal actions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation acea9151-c1c2-4783-9c9f-1d659178ccc0 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e75fe3f7-da28-4ef7-8baf-7a88741e82dc · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39db794-266f-4f71-b8c6-afea1340d505 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efe9d6d-0a96-4d3d-b9df-09d0c1139a53 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TempCompass: Do Video LLMs Really Understand Videos?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd700fdb-2a23-43c3-9a61-04b000dd0e50 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861cd7f6-85bf-478e-bea8-bf3dc51f9482 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3b54af-c5be-4614-b2b4-575ef17ec51b · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Longvideobench: A benchmark for long- context interleaved video-language understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4af261fa-79dd-48c9-a700-bc67770e9b37 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0eb0d0b-aa1a-4a6d-83de-645e6b7d1a3a · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5e342c-f4bb-4d8b-a56c-51bb49b5abcd · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b934793b-b950-48ce-93e9-5c77b1a3a9c0 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa4837b5-47b0-42d8-8d65-e4e6cd901d7a · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78f15c0-b7e3-4db3-8dc0-380def64992e · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8e6e87-15f4-430c-9521-2b29d9083cbf · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Chain-of-thought prompting elicits reasoning in large language models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8580e34b-d30c-4b28-b1e6-c0848ed85041 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Deep reinforcement learning from human preferences
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 89df4697-23fa-4319-9127-b731d1cc5aac · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Proximal Policy Optimization Algorithms
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97814645-f452-4faa-8251-db811fa094db · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24680908-ecc4-4b80-bf19-e826a679e737 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1edf25-7ef5-4916-a861-67fdbd79e8fd · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning to summarize with human feedback
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8c36977-6c37-45b7-9b12-17e04a61dd27 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models WebGPT: Browser-assisted question-answering with human feedback
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e00432-1208-4893-87e9-0d48cb358aa9 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7501a49a-7ab1-4c2a-b32f-8c0ba6a17baf · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Cross-task weakly supervised learning from instructional videos
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcfd53b-1f87-4dc8-b886-63934c8076a9 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4 Technical Report
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7f9d09-8164-4e31-b0ff-cf7cb4912456 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Procedure planning in instructional videos
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 45816025-77a8-4996-8343-f5f185a6bce3 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b696177-b007-40e3-a160-eb62669a7c28 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8f6bae-b1e5-4f52-8979-691d0a85fb06 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models The Llama 3 Herd of Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577db0fc-0d97-4733-8a58-afaccc14980c · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Identifying and mitigating vulnerabilities in llm-integrated applications
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf3fe687-b0a0-444a-87ea-e87f80ecc16e · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen Technical Report
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 047be0b7-5194-461e-9856-6be2066e33e2 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5 Technical Report
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0015e8-0aeb-4867-97b3-f8b93ab7d2bc · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ef27645-4dab-43fc-85a8-edf929a3dbc1 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Improved baselines with visual instruction tuning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation faf3c1a1-325b-4625-ae12-0e58912ab2c3 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DINOv2: Learning Robust Visual Features without Supervision
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53afbc61-2f7a-425b-a164-06134bde3f79 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unmasked teacher: Towards training-efficient video foundation models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 443f75ea-042d-40ab-bbbf-82c5d5f71f07 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-alignment of large video language models with refined regularized preference optimization
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95008de-bb64-4ce2-8fda-886c9f83364d · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models NVILA: Efficient Frontier Visual Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df600f0-39b1-42a5-8f74-b0b2ab60c761 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d2c112-4f3f-477f-9dd0-beaee1a5d3d4 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4o System Card
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3caf5fc-588c-41b9-9866-da8cedf6f0d6 · outbound
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.