Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:51:15.431541Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2412.10471.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:51:15.431541Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:05.146191Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T06:34:19.511206Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3d6d84b9-b5e6-4fe4-8014-55e0c34eba9f · outbound
VCA: Video Curious Agent for Long Video Understanding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51523699-844e-4da4-85e4-bc474baadc22 · outbound
VCA: Video Curious Agent for Long Video Understanding Pixtral 12B
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261889dc-d7b7-4bc8-bffe-c72d02f478d1 · outbound
VCA: Video Curious Agent for Long Video Understanding Vivit: A video vi- sion transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6865a22d-37c3-4bdd-838f-4988f968c1b8 · outbound
VCA: Video Curious Agent for Long Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fc42dc-693d-4ecd-b0dd-fd17260f792f · outbound
VCA: Video Curious Agent for Long Video Understanding Memory consolidation enables long-context video understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8515d0a1-12a8-4620-a969-eb49c4c311b9 · outbound
VCA: Video Curious Agent for Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3abac2-5cce-4289-af18-bfc1dc724b49 · outbound
VCA: Video Curious Agent for Long Video Understanding Is space-time attention all you need for video understanding?,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14f7899-0ab9-4bfd-a0ea-b9a7ff6bda41 · outbound
VCA: Video Curious Agent for Long Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset, 2018
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef648ba7-9177-4563-a7ce-4c4bc552e631 · outbound
VCA: Video Curious Agent for Long Video Understanding PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd7a1a4a-20e4-407a-8323-d6ebf5831f0d · outbound
VCA: Video Curious Agent for Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3732499b-b4b4-46c0-94be-ea5a3901ea72 · outbound
VCA: Video Curious Agent for Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd94c65-4c7c-4883-a74b-5ef4939ffc3d · outbound
VCA: Video Curious Agent for Long Video Understanding Long Story Short: a Summarize-then-Search Method for Long Video Question Answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5b0cb3-e987-447c-9c90-fb5c27115ec2 · outbound
VCA: Video Curious Agent for Long Video Understanding Neural mechanisms of selective visual attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b5829a-899c-4632-be9b-04fe5d604886 · outbound
VCA: Video Curious Agent for Long Video Understanding An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5577ee1c-92ed-47d2-ac91-7265309d36e2 · outbound
VCA: Video Curious Agent for Long Video Understanding Agent ai: Surveying the hori- zons of multimodal interaction, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47980ef-9da4-4686-8445-036e384d9c6f · outbound
VCA: Video Curious Agent for Long Video Understanding Multiscale vision transformers, 2021
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4d0e99-9438-4a05-ad98-7580bec6bd84 · outbound
VCA: Video Curious Agent for Long Video Understanding Videoagent: A memory-augmented mul- timodal agent for video understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216eef99-e433-4e15-ba2d-d43f10802772 · outbound
VCA: Video Curious Agent for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f7901a5-8f93-4ca5-b274-028bec8a148d · outbound
VCA: Video Curious Agent for Long Video Understanding Convolutional two-stream network fusion for video action recognition, 2016
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9006ca0-3711-47a8-a612-5c9a2078e9b1 · outbound
VCA: Video Curious Agent for Long Video Understanding Slowfast networks for video recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef14b41-08f2-42f8-ad88-f8edb94d5787 · outbound
VCA: Video Curious Agent for Long Video Understanding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3389d05-c3d7-4a46-96bb-45eda648ba62 · outbound
VCA: Video Curious Agent for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ba67c1-aebc-4b9b-bec4-a041a6acb94f · outbound
VCA: Video Curious Agent for Long Video Understanding Ego4d: Around the world in 3,000 hours of egocentric video
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c2a2dc18-7077-474d-9a6b-e4de37afa7f2 · outbound
VCA: Video Curious Agent for Long Video Understanding Cogvlm2: Visual language models for image and video understanding, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e8fa3bc7-3dbc-485e-8264-4a12f7377600 · outbound
VCA: Video Curious Agent for Long Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f6678cd-d257-4b71-af64-7cb5170f03a5 · outbound
VCA: Video Curious Agent for Long Video Understanding Cogagent: A visual language model for gui agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d351d7f-b698-4c05-853f-767d1c889ed0 · outbound
VCA: Video Curious Agent for Long Video Understanding GPT-4o System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 477c62d8-1208-4a0a-9f42-4e31cb5725ec · outbound
VCA: Video Curious Agent for Long Video Understanding Videowebarena: Evaluating long context multi- modal agents with video understanding web tasks, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f44c37ae-b272-4ebd-8bbd-5200b06ef533 · outbound
VCA: Video Curious Agent for Long Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video under- standing, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3457cfb-8e41-482b-8dd7-fa7ba192d3bf · outbound
VCA: Video Curious Agent for Long Video Understanding Visualwe- barena: Evaluating multimodal agents on realistic visual web tasks, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5005d86a-8d58-41a3-8e5d-9c944600af86 · outbound
VCA: Video Curious Agent for Long Video Understanding Text-conditioned resampler for long form video understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 880f3b83-c31c-426d-8fcb-fba47ec844e0 · outbound
VCA: Video Curious Agent for Long Video Understanding Mmctagent: Multi-modal critical thinking agent framework for complex visual reasoning, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f21a43c9-f4a2-4950-abe8-2db653c866ff · outbound
VCA: Video Curious Agent for Long Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35849f4b-8d40-4b19-9e9c-ed9b54d7e59d · outbound
VCA: Video Curious Agent for Long Video Understanding Llms meet long video: Advancing long video comprehension with an interactive visual adapter in llms, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2c05279b-68fe-4e6a-a474-16ab5fe9c366 · outbound
VCA: Video Curious Agent for Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3665cbd0-7ea9-40dc-ad99-aa7113cadbbe · outbound
VCA: Video Curious Agent for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bbff455-264e-455a-a1b6-0bf79aa81699 · outbound
VCA: Video Curious Agent for Long Video Understanding Vila: On pre-training for visual language models, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46a8d58a-6029-4d54-b466-009ec2bf1a2b · outbound
VCA: Video Curious Agent for Long Video Understanding Vila: Efficient video-language alignment for video question answering
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a1e798b-10dc-479c-86fd-cc36d57bd13f · outbound
VCA: Video Curious Agent for Long Video Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612e2c43-93f8-4c0b-811d-b46df35b197b · outbound
VCA: Video Curious Agent for Long Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c35b66f-85ab-40d0-b766-f318258234e7 · outbound
VCA: Video Curious Agent for Long Video Understanding DrVideo: Document Retrieval Based Long Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fb01ea-7a3a-45ac-9755-4104a306c77d · outbound
VCA: Video Curious Agent for Long Video Understanding Drvideo: Document retrieval based long video understanding, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46e6c3fb-224b-45ad-9dca-67673151045b · outbound
VCA: Video Curious Agent for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cb7f4b-97b4-4be9-b505-2b0c4a993db4 · outbound
VCA: Video Curious Agent for Long Video Understanding Beyond short snippets: Deep networks for video classification, 2015
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47a4c27d-d36b-4cc9-8205-b49832ef4134 · outbound
VCA: Video Curious Agent for Long Video Understanding A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de5b2391-1995-4716-975e-ba8516b10142 · outbound
VCA: Video Curious Agent for Long Video Understanding Too many frames, not all useful: Efficient strategies for long- form video qa
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5955f83d-b136-40fb-8dac-d4d57c812536 · outbound
VCA: Video Curious Agent for Long Video Understanding Cinepile: A long video question answering dataset and benchmark
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 00f39dc4-3c86-4e6e-8c63-c06ad75c3438 · outbound
VCA: Video Curious Agent for Long Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d0596a-2ccc-4e89-ae0e-6d3cfe265492 · outbound
VCA: Video Curious Agent for Long Video Understanding Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bb21bf32-36d8-4c41-87be-b037535dab71 · outbound
VCA: Video Curious Agent for Long Video Understanding Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4f4958a5-1ca7-450c-be0e-eb9382c7b21a · outbound
VCA: Video Curious Agent for Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4a6df2-4941-45a3-b2ea-c2d322bc254d · outbound
VCA: Video Curious Agent for Long Video Understanding Videobert: A joint model for video and language representation learning, 2019
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6b354a4-655a-47d7-95f2-dbe1febb5d04 · outbound
VCA: Video Curious Agent for Long Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a894cc53-7abf-4b75-adbb-d886e7773247 · outbound
VCA: Video Curious Agent for Long Video Understanding Long-form video-language pre- training with multimodal temporal contrastive learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 814c598b-edbf-4d3c-9d89-4f4317151126 · outbound
VCA: Video Curious Agent for Long Video Understanding Video understand- ing with large language models: A survey, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de543687-321a-44ee-b080-cbf3b9d5466b · outbound
VCA: Video Curious Agent for Long Video Understanding Movieqa: Understanding stories in movies through question- answering
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe5cd61-36ab-42be-b37c-532afffd49a2 · outbound
VCA: Video Curious Agent for Long Video Understanding Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13a532f-41a0-4504-8bc7-e707e9c4257d · outbound
VCA: Video Curious Agent for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80008c2b-7657-4e3f-8dde-b1a662f2e23d · outbound
VCA: Video Curious Agent for Long Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training, 2022
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b5e6cc-f345-43c1-a588-10fcc4ef2960 · outbound
VCA: Video Curious Agent for Long Video Understanding Learning spatiotemporal features with 3d convolutional networks, 2015
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 63c54d6f-f1ad-4f00-875c-22a6580deafa · outbound
VCA: Video Curious Agent for Long Video Understanding Temporal segment networks: Towards good practices for deep action recogni- tion, 2016
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a4a4cbc-e6d5-4829-babc-338979a337a9 · outbound
VCA: Video Curious Agent for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b730db-99aa-4dba-accc-95a730fc01ee · outbound
VCA: Video Curious Agent for Long Video Understanding Lvbench: An extreme long video under- standing benchmark, 2024
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 242bfa80-faa5-42eb-a0d5-48c39bc8d142 · outbound
VCA: Video Curious Agent for Long Video Understanding Videoagent: Long-form video understanding with large language model as agent
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b1b3b59-05ca-4048-9817-1904fb0b4e47 · outbound
VCA: Video Curious Agent for Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f645a1-effd-44c7-9636-206d41f72c32 · outbound
VCA: Video Curious Agent for Long Video Understanding VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91936784-82be-407c-bdcf-d75ed8a71c27 · outbound
VCA: Video Curious Agent for Long Video Understanding LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f7f34a-aac9-4cf7-bd8e-2cb94f5ac74e · outbound
VCA: Video Curious Agent for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos, 2024
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2bdc60be-9dd0-4229-9bab-ff8514d67e17 · outbound
VCA: Video Curious Agent for Long Video Understanding Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fc647345-2c72-4960-b58b-e6a70b1fb2c7 · outbound
VCA: Video Curious Agent for Long Video Understanding Longvlm: Efficient long video understand- ing via large language models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b19d0b9-2fa9-4859-9796-7de444f4911e · outbound
VCA: Video Curious Agent for Long Video Understanding Towards long-form video understanding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70d9bbdf-f295-4f3c-8cfa-3b6b443c9381 · outbound
VCA: Video Curious Agent for Long Video Understanding Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a64946-a01d-4c10-a33e-cb46b526bb87 · outbound
VCA: Video Curious Agent for Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ebff8ae7-e013-4bc4-9187-025f8f21f7db · outbound
VCA: Video Curious Agent for Long Video Understanding Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification, 2018
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d5cf55d-b454-4c7f-bd27-81a0d131a94a · outbound
VCA: Video Curious Agent for Long Video Understanding Openagents: An open platform for language agents in the wild
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e1dbfeea-7d32-4be2-9840-2af4989ece2e · outbound
VCA: Video Curious Agent for Long Video Understanding OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4883504d-1107-4b07-a411-ea6c46adb681 · outbound
VCA: Video Curious Agent for Long Video Understanding Retrieval-based video language model for efficient long video question answering, 2023
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a605b3-565d-4a03-ae31-743de26d7b01 · outbound
VCA: Video Curious Agent for Long Video Understanding Webshop: Towards scalable real-world web interaction with grounded language agents, 2023
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f9b804f9-515c-4df2-9b02-ae48e13da4cf · outbound
VCA: Video Curious Agent for Long Video Understanding Self-chained image-language model for video localization and question answering
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a52fe91e-011f-4564-bbf5-24f989ed223d · outbound
VCA: Video Curious Agent for Long Video Understanding Self-chained image-language model for video localization and question answering
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a30d6094-19d9-495f-99a7-92a78f2b1a49 · outbound
VCA: Video Curious Agent for Long Video Understanding A Simple LLM Framework for Long-Range Video Question-Answering
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d885ae91-1cc3-4a58-a0a5-14228144837e · outbound
VCA: Video Curious Agent for Long Video Understanding Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 786b2f4d-aad2-41d3-8079-eb4383c4a47a · outbound
VCA: Video Curious Agent for Long Video Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9b1cb3-0644-436c-a0af-3a7c2d6e3a39 · outbound
VCA: Video Curious Agent for Long Video Understanding Long Context Transfer from Language to Vision
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e4a9ea-bc64-4391-a384-3af8d474d0fe · outbound
VCA: Video Curious Agent for Long Video Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fecf0dcf-b274-4a34-b8e0-f75cb72ceae1 · outbound
VCA: Video Curious Agent for Long Video Understanding LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fd47c0-fed7-4208-a2a6-024676ca805b · outbound
VCA: Video Curious Agent for Long Video Understanding Videoprism: A foundational visual encoder for video understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20de5e04-6ac3-42ff-b446-a0f1f4577a82 · outbound
VCA: Video Curious Agent for Long Video Understanding Actbert: Learning global-local video-text representations, 2020
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 214f7c4f-2c7f-4f29-88a3-13f88ff0975b · outbound
VCA: Video Curious Agent for Long Video Understanding We include the prompt on how the reward model (R) generates relevance scores in Fig
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8218422c-8a9a-43e3-a93c-e35cafca49b7 · outbound
VCA: Video Curious Agent for Long Video Understanding Dataset Dataset Avg
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 54a0f9f9-e39a-4c4b-87a9-1b83b3adbb29 · outbound
VCA: Video Curious Agent for Long Video Understanding Experimental Results As discussed in Sec
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a1a8f6b0-82e6-4a5c-8c12-10ee970cd76e · outbound
VCA: Video Curious Agent for Long Video Understanding 5.2, in this section, we investigate the common failure cases of our framework, aiming to provide data points and insights for the future research
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e00a398-26eb-43cb-8fe3-3094f39cb71d · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VCA: Video Curious Agent for Long Video Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19217dc-6bd2-41ea-805a-31257957c3a9 · inbound
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs VCA: Video Curious Agent for Long Video Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.