Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:00:53.893326Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2506.18071.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:00:53.893326Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:29:12.184772Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T12:48:11.224098Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 942b5cc6-c1e0-4267-903d-465920c8a26b · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Revealing Single Frame Bias for Video-and-Language Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4371d76a-03a6-4234-86ec-cb89e8140eee · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e78f3d-95ca-403d-98ea-5f93fb05274a · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Rethinking Multi-Modal Alignment in Video Question Answering from Feature and Sample Perspectives
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf141f48-2def-40b8-ab0c-419cec439151 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Can I trust your answer? Visually grounded video question answering,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0fce756f-5b13-4b99-8dab-6464bc2e4262 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TVQA: Localized, Compositional Video Question Answering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a142e82-9404-4b2c-8701-4934220fbc1e · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Zero-shot video question answering via frozen bidirectional language models,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ad3b221-e083-4d99-8ef1-4e2d65c5a961 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34393161-975c-44b4-a476-336271239f40 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Question-Answering Dense Video Events
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490d0bf9-db27-4924-9463-5b39c8645118 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Grounding action descrip- tions in videos,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66f1ab7a-3e27-4924-a21c-034eb1d2399d · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486cf951-d285-4c6e-b262-02403b9c0613 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoagent: Long-form video understanding with large language model as agent,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4275111b-f150-48d2-b68e-a4e5ff29c538 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Morevqa: Exploring modular reason- ing models for video question answering,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ca7c569-43f1-4401-8869-754d63aac396 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoagent: A memory-augmented multimodal agent for video understanding,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a42b439b-c627-43ab-9266-5bc606dd7ebe · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Retrieval-based video language model for efficient long video question answering,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0f2aa3-d150-427e-8eb8-043a64b0eab1 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Hierarchical video-moment retrieval and step-captioning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 579fb63d-d934-4edd-baad-39d9d9769205 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Tgif-qa: Toward spatio-temporal reasoning in visual question answering,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 29a55d55-1b52-47d1-ae58-4fe1058fcea3 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video Question Answering via Gradually Refined Attention over Appearance and Motion,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a50117b2-d411-4d59-b4b3-0321affd932a · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Activitynet-qa: A dataset for understanding complex web videos via question answering,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 883e909c-3ab7-4b22-835c-c41a0327fb11 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Next-qa: Next phase of question- answering to explaining temporal actions,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a224856a-f3a5-423a-be12-2c0f7b0bb0ff · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c933d8-c36a-440b-af01-670e17a2033b · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoqa in the era of llms: An empirical study,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 874431d9-8e44-4c9c-b80f-5fc8a6dcd538 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c68b6f-8159-423e-891e-37fa0df608fe · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1b9624-74c4-4b86-a858-ad85889528ba · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering The surprising effectiveness of multimodal large language models for video moment retrieval,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation befc7197-8c2a-4874-8216-32df8b207f42 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Neptune: The Long Orbit to Benchmarking Long Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a76298-fcd8-4d55-9888-91c570d78e00 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90003f2-4a8c-4125-a83e-90d3dedde6a5 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Chain-of-thought prompting elicits reasoning in large language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0a1ed0f3-7b84-4780-a6fc-9d7abe54ff1e · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Tree of thoughts: Deliberate problem solving with large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73066226-61b1-409f-9adc-a292341aec1c · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b2cec0-88c3-43ed-9a78-cbf0295ed019 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering React: Synergizing reasoning and acting in language models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 691c59fc-a183-447f-9d1b-227cd61f2eac · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Reflexion: Language agents with verbal reinforcement learning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 80b74b4e-4832-473f-8d13-9486d172d2ea · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Improving Multi-Agent Debate with Sparse Communication Topology
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36baa93-5abf-457d-94bc-0bb89a440846 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering An empirical study of end-to-end video-language transformers with masked visual modeling,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0731a994-2002-4fc0-9026-e61e684b35e5 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Language Repository for Long Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c4d89a-2e30-4d24-ac7c-c243654f4fca · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Streaming long video understanding with large language models,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0cd89e7c-11c0-45cd-8bc6-4f53a320ddf1 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14bcbbf9-5db4-4c97-b06d-1947d947e93e · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe0df4f-7736-4f3a-8c85-8fe84cf4168c · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TVR: A large-scale dataset for video-subtitle moment retrieval,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 910ce7cb-ed85-437d-8b31-3a40b512272f · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering UMT: Unified multi-modal transformers for joint video moment retrieval and highlight detection,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2374e396-d699-4e0f-af41-dd23b32cb7b4 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering MomentDiff: Generative video moment retrieval from random to real,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ded4ff59-0737-408c-bec3-9f2ffcb38396 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Query-dependent video representation for moment retrieval and highlight detection,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98945433-9869-4053-b1fd-5ad38e202e5f · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering UnivTG: Towards unified video- language temporal grounding,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d622ec3-cbfb-4999-8861-bdd96a132ab1 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3b1b74d-b7c2-4d27-8c98-d3915b8a67fa · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Learning 2D temporal adjacent networks for moment localization with natural language,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8fa31332-25ce-4b36-930f-88143aa3d190 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Span-based Localizing Network for Natural Language Video Localization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9753c3ad-6f30-43bc-91bc-1c2a6873d5ff · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Dense- captioning events in videos,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f972f6ef-90ea-4ee1-aaea-583067b4cf6f · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Negative sample matters: A renaissance of metric learning for temporal grounding,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bc3d9d4-aae2-4a43-94a2-edeae21d62ae · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cc5d28cb-ac01-4903-b3b0-4cb656dca795 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoChat: Chat-Centric Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4d5ceb-99c8-4215-9611-00b53748ee86 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaa9ef30-e5fc-4914-88d2-77452ecc3a11 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf37d3a-1bec-4b1b-b90a-134b86d2267c · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering ChatVTG: Video temporal grounding via chat with video dialogue large language models,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8be569d4-d419-4088-9988-07dad5799294 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279424c7-390e-4244-a812-fa0606c0b5b0 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering ET Bench: Towards open-ended event-level video-language understanding,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6032c6eb-f3bc-41de-8e72-97f1af116b77 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering LITA: Language instructed temporal-localization assistant,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6fa56daa-10e0-4b4e-b521-292f6c61a455 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering RexTime: A benchmark suite for reasoning- across-time in videos,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a62dc9f2-9263-44d4-a011-a31ae332ecd6 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VTimeLLM: Empower LLM to grasp video moments,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cdb0f847-a985-4c2b-9961-23bfd09a57a2 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TimeChat: A time-sensitive mul- timodal large language model for long video understanding,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d1c0ed6-f37d-4e71-a1b4-b9afbcb2cd9d · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering LoRA: Low-rank adaptation of large language models,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cc93a4c8-9b42-4e89-abc0-2e71454545d3 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Self-chained image-language model for video localization and question answering,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 478cef9a-fc6c-433b-b6e2-ab7fe14ac6f3 · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering GPT-4o System Card
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0116188-38b5-4aea-b3ea-07233c71b8ed · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5846c80-5fad-47c0-a4f5-ee604d4987cf · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Reinforcing Video Reasoning with Focused Thinking
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d294098-dead-4bd6-b080-4485f8f9914d · outbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering SynPO: Synergizing descriptiveness and preference optimization for video detailed captioning,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43ad810-ce2b-4674-9050-c89d21b9c3d3 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Reference 299
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f805b6e5-edd3-4b9e-bb4f-aad1b3ec0003 · inbound
AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.