Pith. sign in

Paper Citation Record · LEDGER

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2504.20091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20091 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:15:50.652098Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:12.739322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:20:11.800532Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa1e6425-295d-4105-908b-86faefb6c1c4 · outbound

This paper cites ENTER: Event Based Interpretable Reasoning for VideoQA.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering ENTER: Event Based Interpretable Reasoning for VideoQA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.482442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.482442Z digest=sha256:e3bdf2c0b7922b0cfcf5aeed2100d12a6b2596453f9c57f363e3a27c225a6486

Observation 5551cb9e-2eff-4e80-b4dc-475bf1407573 · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering HourVideo: 1-Hour Video-Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.488285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.488285Z digest=sha256:857f1f12756b3d49f0c78eeae2b2d56767c06139a1df35ed71a1667ac5650d48

Observation 3e0ccdbc-dc61-4c8e-9ba9-47889b54753e · outbound

This paper cites Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.493392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.493392Z digest=sha256:8d2f14b6d8bb46a55802c02540a7bc014fe82981a0547a66027cfd533d0ce16b

Observation 2c2dd7f3-ec03-476e-906b-b993d97fa18c · outbound

This paper cites Adaptive video relationship detection via graph matching and attention.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Adaptive video relationship detection via graph matching and attention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.278269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.498350Z digest=sha256:084c3c3126bdee900be6fdf33a8bfc48d05361ac6f5c95a32325e9b43874a638

Observation 6a83fc47-71a2-43f2-9405-f98f961ab9e5 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering EVA-02: A Visual Representation for Neon Genesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.503049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.503049Z digest=sha256:834f738fc3ac36ca85f011dfd5be41898ed065e30f4be27f8c8384f444bb1e6c

Observation d32cc318-6bac-40e6-9653-2b7e2686b168 · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.508081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.508081Z digest=sha256:6d51099dedde9cb6c9222e836890a5eb5b5682d437ad9b323a203da6ce6630b4

Observation c5d50635-7b9b-4544-8983-b4e1ef5550e8 · outbound

This paper cites Linvt: Empower your image- level large language model to understand videos, 2024.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Linvt: Empower your image- level large language model to understand videos, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.261938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.513182Z digest=sha256:8b691c64ecb0d9ff7a5c7c6b3419889b624de075b6f07b90274663770edfc4ee

Observation 676aa8b4-368e-4b74-9eaf-265bf3b8d0cf · outbound

This paper cites Learning situation hyper-graphs for video question answering.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Learning situation hyper-graphs for video question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.245868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.518252Z digest=sha256:7079e1e9118ee54065d0960a4cd7d6a190fd3c509a79ce731b093904fd97f269

Observation 5836d3ee-1455-4a25-a072-05b14e1848f2 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering An image grid can be worth a video: Zero-shot video question answering using a vlm

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.229551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.522952Z digest=sha256:d8f7712acee799de244eab44cf86e724ef387b41d571fee7a43d528892344941

Observation acb824ff-54de-4b32-ae30-a082cbfa9acc · outbound

This paper cites VDMA: Video Question Answering with Dynamically Generated Multi-Agents.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VDMA: Video Question Answering with Dynamically Generated Multi-Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.527617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.527617Z digest=sha256:f96ad3dd72090e784d85e9d10cd54d36b6a91f12156652c71ae5d0f40ca2ef6c

Observation 3416796f-bb40-4e78-b057-3022daf27627 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.532725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.532725Z digest=sha256:c1289184db37bb66378ef11025c8d36d70e9b2a33942aa9936aaaec23e7d52ae

Observation 1e40d79c-2319-4eef-bd6c-4144028489e6 · outbound

This paper cites In- tentqa: Context-aware video intent reasoning.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering In- tentqa: Context-aware video intent reasoning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.203549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.537796Z digest=sha256:1461dcc82dd71b167fde0e5e7616dc02d1f7c90cb38d36ead214c1cd082a8da7

Observation bc3043f9-0cf1-473e-be77-879fd79acd5a · outbound

This paper cites Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.188693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.542328Z digest=sha256:ae67346068a792cd37155c05fb03723e267b64d6e6259d9561ba8984f22d80f7

Observation 80967dbf-8a0f-4626-a629-ed2604852816 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llava: Learning united visual representation by alignment before projection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.173758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.546971Z digest=sha256:f51ff512ceaf24a526437aa3742f9cb50633ea55f1b8fe6921ab1fc98e76d0c5

Observation f11da199-0d8f-475b-bc97-8fe9002f8627 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llava: Learning united visual representation by alignment before projection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.158095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.551514Z digest=sha256:6caebcfcb5179f854413b1d017a4c02c9c433ef4d2d6e85469d4159bdfec26e7

Observation 4d65ce78-8e5d-41de-bee1-00065871fba9 · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.555962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.555962Z digest=sha256:6c76fc96112ecc37669125c04ea84048dda9d6028f59c71a9ff0d6f029a505ca

Observation 6707995e-12e0-4b8b-a491-f2475c8c986c · outbound

This paper cites Khan, and Fahad Shahbaz Khan.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Khan, and Fahad Shahbaz Khan

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.142683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.561041Z digest=sha256:4a43aad41170bc2278305ded56bcd25d83ac2ab75e27c9e494da06672e2fd371

Observation 45917096-e0dd-4428-bb9f-d77a939518b0 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.126501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.565450Z digest=sha256:033e02d582624aea0c736db4ed6950679aaf0900ffbd318f3eb7e07e988f8217

Observation fa273361-0a44-4342-80b3-1be3f7b0f28d · outbound

This paper cites Hello GPT-4o: OpenAI’s Multimodal GPT-4 Omni Announcement.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Hello GPT-4o: OpenAI’s Multimodal GPT-4 Omni Announcement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.110811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.569840Z digest=sha256:d15701cad7c0f792c2cf56186dd48202b54a1225e51775c71d71d0e489c428f8

Observation f791700f-3d40-4d0e-a349-76e78e40ad0b · outbound

This paper cites an unresolved cited work.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.574935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.574935Z digest=sha256:62e49e04da5435133b79238a81996c40e5cb9bc67c6a54d550102b7304b235ba

Observation cb6589cf-d9c5-43c5-b877-59a6d59dbd4c · outbound

This paper cites Ts-llava: Constructing visual tokens through thumbnail-and-sampling for training-free video large language models, 2024.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Ts-llava: Constructing visual tokens through thumbnail-and-sampling for training-free video large language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.095293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.580148Z digest=sha256:fd2e53f6056d5949b3c61d26e6c8caa8a5fc2153070cd5515f273f8695898d4f

Observation 97a32367-ab6e-45cc-b227-af4ff9e9ec34 · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos, 2023.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Action scene graphs for long- form understanding of egocentric videos, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.079501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.584750Z digest=sha256:ed6cafcc0ed8be4f22cfb96ab36d2f6d68c23a089a98657d8900c29d3ee199d3

Observation 84acb3fe-1c1f-4021-9b55-da674452dbdf · outbound

This paper cites Kim, Bilge Soran, Raghuraman Krishnamoor- thi, Mohamed Elhoseiny, and Vikas Chandra.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Kim, Bilge Soran, Raghuraman Krishnamoor- thi, Mohamed Elhoseiny, and Vikas Chandra

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.062980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.589237Z digest=sha256:9bb517ba60ab463b0eacf0b5bff614fefa92479785918356260294fafdc9c412

Observation 2c8a76bc-621d-4092-acb6-e445bb50f265 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video under- standing.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Moviechat: From dense token to sparse memory for long video under- standing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.047294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.593711Z digest=sha256:5d86b87296aee64f9dfb03c4aa5514095f808424b9d15733cb54c175623c6b7a

Observation 7a3f4cf5-44cc-4270-b6f8-d474997add27 · outbound

This paper cites Videoagent: Long-form video un- derstanding with large language model as agent.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videoagent: Long-form video un- derstanding with large language model as agent

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.030965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.598329Z digest=sha256:ff144f73fb4788bb390c7dec1a24c0e33af6bfcd673ddabf6889af8782c5164d

Observation a3c7cb4f-0f8f-4371-84f3-041ca0843897 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:51.014535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.603064Z digest=sha256:5ab1cae88fba570625b5fa6e4350f662bcffb1fe17d5d78a849419e8b9b47183

Observation bce10fed-84dc-4364-9e42-6b80d4908e68 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.607489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.607489Z digest=sha256:aa6e1153ed22336f5bca8d8f44535f5415ebf8105559ee7064cf53e2491453d8

Observation 08c727bb-4fa1-4e23-83e3-4214c8c997a3 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.612744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.612744Z digest=sha256:6a6f10c2ea3f8b6352c6cde96c7fe5abbf27421d6fe51c49238611656d19563a

Observation 89ee353c-6709-4e50-adad-69c4ec3deb23 · outbound

This paper cites Lifelongmem- ory: Leveraging llms for answering queries in long-form egocentric videos.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Lifelongmem- ory: Leveraging llms for answering queries in long-form egocentric videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.997159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.617432Z digest=sha256:4bcf172eb31f686c4e535ac1fca674413b8babfd5db4b4133b8827814d609407

Observation 35c81d8f-7884-47de-9fd4-8725fa37067a · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.621828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.621828Z digest=sha256:9fdc6e236e18802a91a1bc62bb3016926796387e9fde5a500df8cc91a9c1cc03

Observation e6caa70b-ae80-4def-aee5-4b9b0d1a4316 · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering NExT-QA: Next phase of question-answering to explaining temporal actions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.979557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.627092Z digest=sha256:17acede4a2161ea1a48df1b1587958ff1913126fae7e26a99a6dae98acb119fa

Observation 058e9ed9-55b4-4611-9aa0-7e1994c72b8a · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:15:50.632416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:15:50.632416Z digest=sha256:8054fb699a05cf2eb43f6058a8e482cca44a0fe31c7618185f078a576366897b

Observation 5040b133-e85b-4888-86ab-2128b412cfc9 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Self-chained image-language model for video localization and question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.963748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.637968Z digest=sha256:334f09c7cf769f81b9ac89efce765cd4b9f609aa94535153992a758a570e3f2b

Observation 0e88c148-a2b3-47b4-a6da-388ddad66e52 · outbound

This paper cites A simple llm framework for long-range video question-answering.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering A simple llm framework for long-range video question-answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.947691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.642750Z digest=sha256:1fe1e2bdfa9ec8a5cc90482329d68bb2a4eff591b476b01a8d0ec6c2fe1e904e

Observation 09e99bca-4720-465d-9119-c997c76926be · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.930519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.647211Z digest=sha256:010fe6e50958ac1d3fbe3894995a55c51171d0b2efa944c433817c335772ce5c

Observation ac40e92a-6de6-4aa7-bede-adcf353767c4 · outbound

This paper cites Hcqa @ ego4d egoschema challenge 2024, 2024.

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Hcqa @ ego4d egoschema challenge 2024, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:15:50.913525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:15:50.652098Z digest=sha256:b0d41f13358ca9d5e977bb7d06684bdadec4881cd81bd3b331f792ce19806c04

Pith citing papers

Observation bc3b5004-dfc5-4e02-a6dc-af4e8fc96a92 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.739322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.739322Z digest=sha256:1aacbf18113a456c7dd7dda9d94504a0b269276895022c3192d187c6fa601918

Observation f87ec8e5-4f9f-4518-a73d-372fddc88871 · inbound

AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning cites this paper.

AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.802601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:18:19.580156Z digest=sha256:1b6045647d5177db9e5e0fd17c56319b2e3bec028d31dd6f8820110d34b2ff19

Observation b91c65b1-a8d0-4f39-9fc3-459a2d524396 · inbound

GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing cites this paper.

GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:51.052309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:13:30.400059Z digest=sha256:00a0380a3a48d73a96a03f7383b940f3b9779ae7d20fce16ebf5f7b50c2a8ea4

Observation 40b269ad-0e41-43eb-ac0d-a2e4c554f35b · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.484637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:6984c271fe12a4dff438b99f4063e34cc2b709fe5887f055632b84753a602005

Observation b6d20c69-d121-4d76-bb05-73b40e4804ad · inbound

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework cites this paper.

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T13:13:54.698349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:13:54.698349Z digest=sha256:c6a095443b44d5be2986281ba367c7fc1d437c9bfb96aa1ab6663362b9bb1fbb

Observation de8679bc-353a-45a9-b2ca-0d4f70d4bb3c · inbound

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding cites this paper.

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:48:10.936504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:48:10.936504Z digest=sha256:92852da846f247ef5c4ba96bbcf06eb7b1e73509825e5610d5f07724d67183df