Pith. sign in

Paper Citation Record · LEDGER

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2506.18071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18071 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:00:53.893326Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:29:12.184772Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:48:11.224098Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 942b5cc6-c1e0-4267-903d-465920c8a26b · outbound

This paper cites Revealing Single Frame Bias for Video-and-Language Learning.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Revealing Single Frame Bias for Video-and-Language Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.580027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.580027Z digest=sha256:3776a0e35db427cea35e7a40b19ef85ca86b728170a03f54836ac757bb9fd2d1

Observation 4371d76a-03a6-4234-86ec-cb89e8140eee · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.586090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.586090Z digest=sha256:2fadcf09f77702292ed098a0258c93c3b5d116eb51a89edd8a30f18358148759

Observation 49e78f3d-95ca-403d-98ea-5f93fb05274a · outbound

This paper cites Rethinking Multi-Modal Alignment in Video Question Answering from Feature and Sample Perspectives.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Rethinking Multi-Modal Alignment in Video Question Answering from Feature and Sample Perspectives

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:29:32.849926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.591633Z digest=sha256:5be90463919246cff5e59589321437a39eb5e1898d332d78b572cfc20ff792fc

Observation bf141f48-2def-40b8-ab0c-419cec439151 · outbound

This paper cites Can I trust your answer? Visually grounded video question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Can I trust your answer? Visually grounded video question answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.248445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.597435Z digest=sha256:f2007023b0f2bb9c9c3b0cb2715002ca36db7c64a9f80c2c6de885bf0068d380

Observation 0fce756f-5b13-4b99-8dab-6464bc2e4262 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TVQA: Localized, Compositional Video Question Answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.602487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.602487Z digest=sha256:1e5cb0cb99755f438483bba6f08ee4fd6dbad3830c0bdfb6600e99c76a39f1a9

Observation 0a142e82-9404-4b2c-8701-4934220fbc1e · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Zero-shot video question answering via frozen bidirectional language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.233212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.607566Z digest=sha256:f9fe4bee5dd821b01d4445ccd5d4a0100f4050bb4405cfb0558f986518b13bf5

Observation 2ad3b221-e083-4d99-8ef1-4e2d65c5a961 · outbound

This paper cites Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.613167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.613167Z digest=sha256:c8b4330c39d6752b49626c801f9a0ee736e14957460c93a56c9cb5d9092a13c4

Observation 34393161-975c-44b4-a476-336271239f40 · outbound

This paper cites Question-Answering Dense Video Events.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Question-Answering Dense Video Events

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.618558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.618558Z digest=sha256:75719579e503bfd24ab9b65a9232d4e76fd5ab2a229b73849f4c4026975bb82e

Observation 490d0bf9-db27-4924-9463-5b39c8645118 · outbound

This paper cites Grounding action descrip- tions in videos,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Grounding action descrip- tions in videos,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.188067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.623593Z digest=sha256:4677b339f7d4536e3e3f7958a664cc28c4e48b5ea24404ed06e459bb788b4dde

Observation 66f1ab7a-3e27-4924-a21c-034eb1d2399d · outbound

This paper cites QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.628143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.628143Z digest=sha256:45ce39241006650723c0377ba089bdbcfd69f3cd22909665a2e877b7af0b4d70

Observation 486cf951-d285-4c6e-b262-02403b9c0613 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoagent: Long-form video understanding with large language model as agent,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.173111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.633884Z digest=sha256:ce3c5863cb644e5371e27b7148ccc11a8522b0d66ca6cf3405753b6ecd48494d

Observation 4275111b-f150-48d2-b68e-a4e5ff29c538 · outbound

This paper cites Morevqa: Exploring modular reason- ing models for video question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Morevqa: Exploring modular reason- ing models for video question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.157058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.638742Z digest=sha256:db44846f378d9409f797032447e3f2efdcdeda2586ef7b08918c99203e687d11

Observation 7ca7c569-43f1-4401-8869-754d63aac396 · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoagent: A memory-augmented multimodal agent for video understanding,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.138961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.643743Z digest=sha256:2d77235aeb7377bbfa01ad03dd7ccd127badf752b547675c0125cb3876bcf566

Observation a42b439b-c627-43ab-9266-5bc606dd7ebe · outbound

This paper cites Retrieval-based video language model for efficient long video question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Retrieval-based video language model for efficient long video question answering,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.648253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.648253Z digest=sha256:3999ad53995c54fe4f44e6f979335ad72fee39dffc7163eb7173045eb38f3b0e

Observation 6b0f2aa3-d150-427e-8eb8-043a64b0eab1 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Hierarchical video-moment retrieval and step-captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.122910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.653074Z digest=sha256:ac21080d580cf8c7285d45819a67b0be3b96f6cfcaaa5811124238d48cd72c75

Observation 579fb63d-d934-4edd-baad-39d9d9769205 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Tgif-qa: Toward spatio-temporal reasoning in visual question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.107800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.657697Z digest=sha256:75ff974eaf4141a509eea2a903efc7b47d0ac1fb66efa2869a336187f94e35f3

Observation 29a55d55-1b52-47d1-ae58-4fe1058fcea3 · outbound

This paper cites Video Question Answering via Gradually Refined Attention over Appearance and Motion,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video Question Answering via Gradually Refined Attention over Appearance and Motion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.091516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.662516Z digest=sha256:ca35aafa8af51fd6c32c5e1d406560579589a2a70927727e337bb457c219f8ea

Observation a50117b2-d411-4d59-b4b3-0321affd932a · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.074807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.666947Z digest=sha256:8acd39486bcfd5e306693b502bb519eeb0da1f01f5276fe3a6f4aef39329a865

Observation 883e909c-3ab7-4b22-835c-c41a0327fb11 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Next-qa: Next phase of question- answering to explaining temporal actions,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.058415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.672141Z digest=sha256:b2af3cbd1cb79724c80a6359f529121d320570e0025bdbd01bd3f4b5ee8affe2

Observation a224856a-f3a5-423a-be12-2c0f7b0bb0ff · outbound

This paper cites VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.676692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.676692Z digest=sha256:54c9ac1f564495c34d6c538c9bdd4d1fa006f6fb9c347969368ae1d93ffc92f9

Observation 10c933d8-c36a-440b-af01-670e17a2033b · outbound

This paper cites Videoqa in the era of llms: An empirical study,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Videoqa in the era of llms: An empirical study,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.040034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.681419Z digest=sha256:781d7411c2a610b9a0a645fdc8ad962b78af6cfd259806e5bfb66861d90733f9

Observation 874431d9-8e44-4c9c-b80f-5fc8a6dcd538 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.686027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.686027Z digest=sha256:d76d3bb3a5d4d1dafd311cb7223c13de69fba807e8dcf08c721c5b6b093cc61c

Observation 69c68b6f-8159-423e-891e-37fa0df608fe · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.690487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.690487Z digest=sha256:af9baa8c4b67c50436285bb78e1a23794acbd27fd5be05683a4853710af85066

Observation ab1b9624-74c4-4b86-a858-ad85889528ba · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering The surprising effectiveness of multimodal large language models for video moment retrieval,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.696992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.696992Z digest=sha256:76ba2a415cd888391a0ed6954070666c19ccfcc4a2192923ce55db565c41653e

Observation befc7197-8c2a-4874-8216-32df8b207f42 · outbound

This paper cites Neptune: The Long Orbit to Benchmarking Long Video Understanding.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Neptune: The Long Orbit to Benchmarking Long Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.701216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.701216Z digest=sha256:5d6e33491c2a54c24832c05a863adc06badc7126c9f8f65671c36b17a554125f

Observation c6a76298-fcd8-4d55-9888-91c570d78e00 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.705612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.705612Z digest=sha256:fa80596920cc506104a60ad043bc77b6eb711d0708731b320f011dc3374e5081

Observation c90003f2-4a8c-4125-a83e-90d3dedde6a5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Chain-of-thought prompting elicits reasoning in large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.022186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.710534Z digest=sha256:3422154f12afe90b7eb7ce0420b9b0869bf373250b2df2578faeee40fc0e6904

Observation 0a1ed0f3-7b84-4780-a6fc-9d7abe54ff1e · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Tree of thoughts: Deliberate problem solving with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:55.006531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.715152Z digest=sha256:d9146faced05367b2f1e6ce1ae5c547f5ffbbd8def4477400f37c4471821545d

Observation 73066226-61b1-409f-9adc-a292341aec1c · outbound

This paper cites Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.720119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.720119Z digest=sha256:879bf9c4571c97e40795e2cfa994cb2c8879a9a621d0b6bc33f76ed9244f50b5

Observation b4b2cec0-88c3-43ed-9a78-cbf0295ed019 · outbound

This paper cites React: Synergizing reasoning and acting in language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering React: Synergizing reasoning and acting in language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.991169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.725282Z digest=sha256:84ffb97c9e5f6d273ef7c43b9b4a4ce578f9e81ab0f0a0f9e9a691d51d40a1c7

Observation 691c59fc-a183-447f-9d1b-227cd61f2eac · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Reflexion: Language agents with verbal reinforcement learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.975804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.730311Z digest=sha256:b94059a41ddb96486ac8b4254c5a3a91cd2ea4ed4dc203cf9e210fe5f90d5aca

Observation 80b74b4e-4832-473f-8d13-9486d172d2ea · outbound

This paper cites Improving Multi-Agent Debate with Sparse Communication Topology.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Improving Multi-Agent Debate with Sparse Communication Topology

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.735081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.735081Z digest=sha256:98d09356d6f45936f35d7f6e530cb2ad2f85c159637f69a59f78fdbc21687dd2

Observation e36baa93-5abf-457d-94bc-0bb89a440846 · outbound

This paper cites An empirical study of end-to-end video-language transformers with masked visual modeling,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering An empirical study of end-to-end video-language transformers with masked visual modeling,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.960413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.739820Z digest=sha256:f33a592a7309f6540a91ceefcccac00810e292a09af674fb7d227f8e33f7fe6d

Observation 0731a994-2002-4fc0-9026-e61e684b35e5 · outbound

This paper cites Language Repository for Long Video Understanding.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Language Repository for Long Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.744283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.744283Z digest=sha256:9300775de50a7d9e408e62d302dc62d29b60da674b7adb0bd9c5d9af44b0a208

Observation 37c4d89a-2e30-4d24-ac7c-c243654f4fca · outbound

This paper cites Streaming long video understanding with large language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Streaming long video understanding with large language models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.945221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.749453Z digest=sha256:0a7387811e3c398d3dda035340b74f97623c872d801439a5d54a9e4241fa1b82

Observation 0cd89e7c-11c0-45cd-8bc6-4f53a320ddf1 · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.754001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.754001Z digest=sha256:68593757b832d85da7bf06e6f4a81bd120adb4d9348166bfe1ababe75a13f119

Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.758638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.758638Z digest=sha256:b87c91dd74cecd352fe28c766f9229ebfcc18afae11593a5be12eb20c5010d99

Observation 14bcbbf9-5db4-4c97-b06d-1947d947e93e · outbound

This paper cites Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.763494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.763494Z digest=sha256:4a671eca8ea5bff89099e4792f6b1a605a2b2210ad16b3327d73af67490d2680

Observation 4fe0df4f-7736-4f3a-8c85-8fe84cf4168c · outbound

This paper cites TVR: A large-scale dataset for video-subtitle moment retrieval,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TVR: A large-scale dataset for video-subtitle moment retrieval,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.929247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.768041Z digest=sha256:f7f737793fd08bddef9ea91dde0f5d2330bd887707f5ccc06154bb644cd37a4a

Observation 910ce7cb-ed85-437d-8b31-3a40b512272f · outbound

This paper cites UMT: Unified multi-modal transformers for joint video moment retrieval and highlight detection,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering UMT: Unified multi-modal transformers for joint video moment retrieval and highlight detection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.914079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.772748Z digest=sha256:22b6c80795c1acd6980b9902f6ad97c868ffcbf57b16805fe4a523be2f0f1d26

Observation 2374e396-d699-4e0f-af41-dd23b32cb7b4 · outbound

This paper cites MomentDiff: Generative video moment retrieval from random to real,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering MomentDiff: Generative video moment retrieval from random to real,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.896056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.777175Z digest=sha256:c60ec8765e746a2fbad41a9962aba98115aa5cf15ad4bf0e85db7e9a2c2664f1

Observation ded4ff59-0737-408c-bec3-9f2ffcb38396 · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Query-dependent video representation for moment retrieval and highlight detection,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.876372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.782460Z digest=sha256:a7f6e802f17960bad3b5e940b8090cc097e688945f59aedb38ad4e0140e93a5e

Observation 98945433-9869-4053-b1fd-5ad38e202e5f · outbound

This paper cites UnivTG: Towards unified video- language temporal grounding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering UnivTG: Towards unified video- language temporal grounding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.860298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.787534Z digest=sha256:e256f80fc7e6ef9e6fb0f7a92b197b6083ef44be3d3c941fb160a8a2019c8b0d

Observation 7d622ec3-cbfb-4999-8861-bdd96a132ab1 · outbound

This paper cites R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.845143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.791921Z digest=sha256:c8978b24582ab8045be6a60434f0e390038e8f9d2fa181569449ea494bbb2a34

Observation a3b1b74d-b7c2-4d27-8c98-d3915b8a67fa · outbound

This paper cites Learning 2D temporal adjacent networks for moment localization with natural language,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Learning 2D temporal adjacent networks for moment localization with natural language,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.827568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.796584Z digest=sha256:280f8267fa7c18ec77288886ca856dddede351cfafc78467c77696b722b1d1ac

Observation 8fa31332-25ce-4b36-930f-88143aa3d190 · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Span-based Localizing Network for Natural Language Video Localization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.801367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.801367Z digest=sha256:3dcee32a26b4d22d7ecf38589e0f0ebe92f9d29ed3e3f634a9f40ccb4d8b6522

Observation 9753c3ad-6f30-43bc-91bc-1c2a6873d5ff · outbound

This paper cites Dense- captioning events in videos,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Dense- captioning events in videos,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.810505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.806272Z digest=sha256:5d9aeda6c84b0bcb3064469ddfa5712a7e7c02b33e66a48591416bf6af423a75

Observation f972f6ef-90ea-4ee1-aaea-583067b4cf6f · outbound

This paper cites Negative sample matters: A renaissance of metric learning for temporal grounding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Negative sample matters: A renaissance of metric learning for temporal grounding,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.793195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.811120Z digest=sha256:d303bd7679c8128878284301889fc81fdaa60dde48995d59b0aaed7d01f99062

Observation 6bc3d9d4-aae2-4a43-94a2-edeae21d62ae · outbound

This paper cites Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.776043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.815679Z digest=sha256:5711b2e2cfe042f0ad27fff9ce84f502fad08e27e63503ea4608efdb53575a69

Observation cc5d28cb-ac01-4903-b3b0-4cb656dca795 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoChat: Chat-Centric Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.820562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.820562Z digest=sha256:fc6d40fd8e339cb5c2e3b33283388813349754fc0f2b0a04bcb092fe8731cfd9

Observation 8d4d5ceb-99c8-4215-9611-00b53748ee86 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.825598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.825598Z digest=sha256:fbc0f33e09915d35295bfa16fa9208ce053132c39b239227d4460b77d484c174

Observation aaa9ef30-e5fc-4914-88d2-77452ecc3a11 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.830018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.830018Z digest=sha256:4c05bf2b4b61305efecfae063a18ce37f3c481612e1937cdb23e5f72bdae8089

Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.834507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.834507Z digest=sha256:6ee8e1bcdb6fadb6ee96c9f9308836b1d478c1772efc37ccd2baf6ab38dceecb

Observation 0cf37d3a-1bec-4b1b-b90a-134b86d2267c · outbound

This paper cites ChatVTG: Video temporal grounding via chat with video dialogue large language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering ChatVTG: Video temporal grounding via chat with video dialogue large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.756324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.838938Z digest=sha256:da9a3e3106df493a189f6e20ebe371cddc79ee093e7c446de2f8105fb53646b0

Observation 8be569d4-d419-4088-9988-07dad5799294 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.843205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.843205Z digest=sha256:5ef84104966f52e49facb8dfff0d019e8b873caeba45a57d6d11a56aab743400

Observation 279424c7-390e-4244-a812-fa0606c0b5b0 · outbound

This paper cites ET Bench: Towards open-ended event-level video-language understanding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering ET Bench: Towards open-ended event-level video-language understanding,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.738634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.847643Z digest=sha256:0bdfb099f409af1ec47f0b0d4c53713b7d954cf8b0eeda1cee56d1c57da92c29

Observation 6032c6eb-f3bc-41de-8e72-97f1af116b77 · outbound

This paper cites LITA: Language instructed temporal-localization assistant,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering LITA: Language instructed temporal-localization assistant,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.721063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.852202Z digest=sha256:894256192763fe13cbdb62f2fdd073b6822fe97c5827a323ccd1ed4fe990c58e

Observation 6fa56daa-10e0-4b4e-b521-292f6c61a455 · outbound

This paper cites RexTime: A benchmark suite for reasoning- across-time in videos,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering RexTime: A benchmark suite for reasoning- across-time in videos,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.702056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.856811Z digest=sha256:08a9535feb58928a1803b0bc8889678afcdf989c26acb322f2a5b4c0c60d0daa

Observation a62dc9f2-9263-44d4-a011-a31ae332ecd6 · outbound

This paper cites VTimeLLM: Empower LLM to grasp video moments,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VTimeLLM: Empower LLM to grasp video moments,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:00:54.685218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.861293Z digest=sha256:f0bdd7b049974b825172713bfe495ca2ca98d60cc12dcdc76f73ccd138e74022

Observation cdb0f847-a985-4c2b-9961-23bfd09a57a2 · outbound

This paper cites TimeChat: A time-sensitive mul- timodal large language model for long video understanding,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering TimeChat: A time-sensitive mul- timodal large language model for long video understanding,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:29:33.770677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.865550Z digest=sha256:b09d33d5dbb37d34053cf1144b3d53075a5f5da9669d912ef832eff9e15421bf

Observation 3d1c0ed6-f37d-4e71-a1b4-b9afbcb2cd9d · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering LoRA: Low-rank adaptation of large language models,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:29:33.429952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.870517Z digest=sha256:20fe1a078f665c15c03faa38e407b28cf856c6e6768c7515216b3acd8bccc527

Observation cc93a4c8-9b42-4e89-abc0-2e71454545d3 · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Self-chained image-language model for video localization and question answering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:29:33.179156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:00:53.874850Z digest=sha256:c96593f03b79df59b99cbc81e71a8a014a9b92ccfd3a680684dd74a9badd6e0b

Observation 478cef9a-fc6c-433b-b6e2-ab7fe14ac6f3 · outbound

This paper cites GPT-4o System Card.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering GPT-4o System Card

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.879350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.879350Z digest=sha256:097857708c6cfd36c9d1db8ff49dce0503cbc145159805a32eb536e4879c6da4

Observation c0116188-38b5-4aea-b3ea-07233c71b8ed · outbound

This paper cites COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.883615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.883615Z digest=sha256:76191e261defa764ab68c802195c25b882d0fc5c8749d0b1b1f7a8ad9dd6a29e

Observation f5846c80-5fad-47c0-a4f5-ee604d4987cf · outbound

This paper cites Reinforcing Video Reasoning with Focused Thinking.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Reinforcing Video Reasoning with Focused Thinking

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.888143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.888143Z digest=sha256:8f751e8de755266413873117dcdc322bf4f34dc8b04dc12155b0f07e6ca1154d

Observation 4d294098-dead-4bd6-b080-4485f8f9914d · outbound

This paper cites SynPO: Synergizing descriptiveness and preference optimization for video detailed captioning,.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering SynPO: Synergizing descriptiveness and preference optimization for video detailed captioning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.893326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.893326Z digest=sha256:6a26981dd02f5c1b0a62149415d3c2129e5eb9aaf300630fde5f2cc5c4466962

Pith citing papers

Observation d43ad810-ce2b-4674-9050-c89d21b9c3d3 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

Reference 299

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:12.184772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:12.184772Z digest=sha256:2fce0defa60989c1f9f493238312ba2e5f05069ea1e3d364989a3553b19d1810

Observation f805b6e5-edd3-4b9e-bb4f-aad1b3ec0003 · inbound

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding cites this paper.

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:48:11.229386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T12:48:10.950796Z digest=sha256:b56a965787dc7348bf2e0d524ae6b245210c0d12d456c4924f747955d00f1b4d