Pith. sign in

Paper Citation Record · LEDGER

ENTER: Event Based Interpretable Reasoning for VideoQA

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2501.14194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14194 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T05:01:06.299758Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:41:42.061219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T20:31:13.191851Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d07d8b03-e851-4f23-96c4-c44d70652789 · outbound

This paper cites Neural module networks.

ENTER: Event Based Interpretable Reasoning for VideoQA Neural module networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.990893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:63b46a08fc682082eaa2613b5fb577c012e85cb717900a1f5383769d4158c891

Observation 8699864a-d3a2-4c52-a355-91571959b530 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings.

ENTER: Event Based Interpretable Reasoning for VideoQA Hiervl: Learning hierarchical video-language embeddings

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.994488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ef4841335e4a768ecb744ce73251c09e58b760d3a7527ef1b794df6b11a07fc8

Observation de589788-3fb3-4e9c-bb5b-9447450f977e · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.985455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:4842ff5c5bc38b8315cfefb73adce1621c441ac995e48c4f12c7489b1d1c93dd

Observation 6a054b05-a283-419b-bd53-51d98f322e95 · outbound

This paper cites Grounded multi- hop videoqa in long-form egocentric videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Grounded multi- hop videoqa in long-form egocentric videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:02:37.270670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:2b243131411def8d61dbdba0fd3c43374a19a8799a134b0ecf979eccf8f73b7c

Observation d32f3b1f-6ffa-46ce-91fd-e4dae113264e · outbound

This paper cites Kitani, and László A.

ENTER: Event Based Interpretable Reasoning for VideoQA Kitani, and László A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.988331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:340db84a98f9fba3cdc3c4cc8a9c90c0c11215bf438a23f4b20a004cee669a2a

Observation 6e6bd0ec-03a1-4329-b41a-038edf6a51d9 · outbound

This paper cites The llama 3 herd of models.

ENTER: Event Based Interpretable Reasoning for VideoQA The llama 3 herd of models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.979794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:b8d134af1956e527e759205e20b4cd0feffe7c04eadbae1cc8fde6825279c554

Observation 15544bb3-8e15-4064-aece-eaee78c0082f · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Videoagent: A memory-augmented multi- modal agent for video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.997794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:c4f528af79db68e2e4c2f06231592933da4d84e2402a6d8300d76ef42b72ba49

Observation ddef7a3a-d5c7-4f5c-b674-2fa461a00acb · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:02:37.274890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:97a24eff522d31e4a74e5ef714b97f751f1c8f21017ea9e009eda8b1891daa29

Observation dc1f27ce-957e-4b37-bedc-182de65930f0 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

ENTER: Event Based Interpretable Reasoning for VideoQA Visual program- ming: Compositional visual reasoning without training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.982429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fbd070755a1cfba3102a25912b00c786bd3bdc589381f2b9b705049860049f99

Observation e990e8d3-d1ea-418f-b7ea-f5a5329c3250 · outbound

This paper cites Free video-llm: Prompt-guided visual perception for efficient training-free video llms.

ENTER: Event Based Interpretable Reasoning for VideoQA Free video-llm: Prompt-guided visual perception for efficient training-free video llms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.949958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:0a50162a122ffe39605b65ac687a2b0d1dd4c6a7ed3632a84febb03bb7f8a8fb

Observation 896c9d79-5826-414d-8021-02c1fa92a8e5 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Video recap: Recursive captioning of hour-long videos

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.982252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:287bb9e0d852b072a30bd23ad92371bdaa257528c639c8e1b39a08831f236160

Observation cd80d71d-2e03-4afb-b129-522ffd695730 · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.911485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fd51c0bcd4e5065a7ee29e79098aa1f91e783a5216eb4717efecaf2bc8f1a521

Observation 90b8005a-a565-4389-80dc-f0838b7a2a14 · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Intentqa: Context-aware video intent reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.946657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:e20cb4c75d4999d8b07af646612ef1514d0fdd6599f1621dd1052b1060104a29

Observation 3171786e-4afe-40f0-815d-4b0b656c054d · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Intentqa: Context-aware video intent reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.952755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:832b825d59dc929b2187acc2bb294db8e4cc0039e6a497dd78b807a1141de753

Observation b15977f6-c395-45a1-8392-193ebd2a4982 · outbound

This paper cites Videochat: Chat-centric video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Videochat: Chat-centric video understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.936590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:74cb4164777f735d3903ecae119ee7f9f72805f8db600a5985c70b837a2860e0

Observation 9440fa02-5773-45f2-b853-e3c76e922f59 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

ENTER: Event Based Interpretable Reasoning for VideoQA Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.956512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:9bf27a2bb0e5c7bf55a3654de29875124ea42c392e55b6be17b4de91ecd28caa

Observation 58dc5d70-4ba1-4ab3-ae60-dec492c3812e · outbound

This paper cites End-to-end video question answering with frame scoring mechanisms and adaptive sam- pling.

ENTER: Event Based Interpretable Reasoning for VideoQA End-to-end video question answering with frame scoring mechanisms and adaptive sam- pling

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.959343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:639b9da69a9e4de38f6ccd8d225fb731924964c19d4cf33e24b62a4eab29fe1c

Observation c6ea5862-fb7a-4aef-a489-18ff8209eaef · outbound

This paper cites VideoIN- STA: Zero-shot long video understanding via informative spatial-temporal reasoning with LLMs.

ENTER: Event Based Interpretable Reasoning for VideoQA VideoIN- STA: Zero-shot long video understanding via informative spatial-temporal reasoning with LLMs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.937177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:9fd6db3a36408a623051bb78ff3d50527e25e21094ed0ba3d60d014dcd80c728

Observation c5a1ddf9-00b0-4bae-a3aa-e0dfa638d972 · outbound

This paper cites Vx2text: End-to-end learning of video-based text generation from multimodal in- puts.

ENTER: Event Based Interpretable Reasoning for VideoQA Vx2text: End-to-end learning of video-based text generation from multimodal in- puts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.939673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:781503a84f50fe37a4e5033668bdc98bbaa79cee83503ec38e2871cfa2d40dc6

Observation a5be5e2c-e8a5-4970-88e5-a648e50a850d · outbound

This paper cites Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval.

ENTER: Event Based Interpretable Reasoning for VideoQA Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.873754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:2265668a1dc5094037cd8cdbbc9fc77d4ede4187702eb7122b3dfb301e76d7d1

Observation f7d45b12-2098-4069-a197-646927d4d0bd · outbound

This paper cites Training-free deep concept injection enables lan- guage models for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Training-free deep concept injection enables lan- guage models for video question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.030597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:c2f09f7129b2977d7f45109f9ea3fa9a72b77db3c8b89a58d28e1b12fed9c8c1

Observation be98f818-536d-411e-b3a8-5d232fa30cdb · outbound

This paper cites Hair: Hierarchical visual-semantic relational reasoning for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Hair: Hierarchical visual-semantic relational reasoning for video question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.922613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:66751ba1a38b8c9e06512044d28001a65659f35105bb55813d18aee1e7040ece

Observation 44b67fe6-a40f-441f-9a58-bcf85655c5dd · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.926521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:9f6766d198d8eaef0310319be3f21555b9e6e3993d3f54a4279a190f46851fbc

Observation 65fa6e4c-3285-4513-b444-d4ea0482cf19 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.919876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:d888bb117199c59d2aa327cb16116ff79d8621cdd5eba7c1a3c97dd5350b6b4f

Observation 00c61d5e-95d3-4662-9a73-c563c5fe94b4 · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Morevqa: Exploring modular reasoning models for video question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.033628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fe8f39efad5e87485ae0db06cbca4b38da713bef2cc309a0f2110e7468735a65

Observation 7f61adff-ff87-4656-99b4-fc396cff212e · outbound

This paper cites Question-instructed visual descriptions for zero-shot video answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Question-instructed visual descriptions for zero-shot video answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.988970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:771ed9ac36c77490c441d5e864ff2063c8c2024f75176c2150221cc3b9cfdfb2

Observation d4485737-3411-4d71-bbf1-95e07bf34cb8 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

ENTER: Event Based Interpretable Reasoning for VideoQA Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.938663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:3a87a8a2987da4b74efc81dd0a4abe150947b66ebdb2b67af3183207f125aa7c

Observation af6a8a11-0812-44d9-a385-0462a68822e6 · outbound

This paper cites Gpt-4 technical report.

ENTER: Event Based Interpretable Reasoning for VideoQA Gpt-4 technical report

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.962422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ebe37a80e7aa547c5694c191587343c2d91a198ab3f347c101a384859cb474fb

Observation 594ff0a8-6fc3-4906-ad49-cd492f120dad · outbound

This paper cites Retrieving-to-answer: Zero-shot video question answering with frozen large lan- guage models.

ENTER: Event Based Interpretable Reasoning for VideoQA Retrieving-to-answer: Zero-shot video question answering with frozen large lan- guage models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.012879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:636cad4772c464054ef304b65103b5046ef23fde05c63b665c8ae785cbc6315f

Observation be66a808-9b32-4fbf-9d6e-ec488c361027 · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.949307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:641bfd391361ec6ee65bf75c6408901f31998e4efd87794e0fe55a5b65524f5b

Observation 07ac8de7-76db-47f6-a70a-6b9171a03db6 · outbound

This paper cites Micap: A unified model for identity- aware movie descriptions.

ENTER: Event Based Interpretable Reasoning for VideoQA Micap: A unified model for identity- aware movie descriptions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.979491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:d0d416f268da65e37d8b03d70f4e743b42cec8dbdf654ad6a77f8a6f2faabb7a

Observation b9e2e6c4-83a0-4678-9b8e-e55ca3d93f79 · outbound

This paper cites Traveler: A modular multi-lmm agent framework for video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Traveler: A modular multi-lmm agent framework for video question-answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.027244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:e54b702ba6f4af7b7328d572e25f7b0bcda2b1aa5e4e709e5288e61ac54934ea

Observation dbd6406f-eede-4283-87aa-2c0473cbc8fe · outbound

This paper cites Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra.

ENTER: Event Based Interpretable Reasoning for VideoQA Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.929695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:cb06b519d4265dfb5e6d6b5fa1aea2d8e99fd872f6425c51af16d8188454ab41

Observation a905d258-c2e8-4213-91e6-089eb8141f0e · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Progprompt: Generating situated robot task plans using large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.992579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:4807cccd9983c016537c98de803569c6059d619693854ff7a3d556836f2da30b

Observation 635f295a-feae-430a-b6c7-081446c85ff8 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.009653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:383d222624658fdde1fe4450baffbdb2a8d44affaa5d9fe2918862a1f297cffd

Observation dab82b8d-8f7f-4d33-b091-7e7b9e46447d · outbound

This paper cites Modular visual question answering via code generation.

ENTER: Event Based Interpretable Reasoning for VideoQA Modular visual question answering via code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.870947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:89fcabe8fe0eb02a26ef24f64e4e08a1fc7616057d597df6c4ffdbb5f1330f4a

Observation 1facce90-297f-455a-a708-7ea961c17ae1 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Vipergpt: Visual inference via python execution for reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.965623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fd9e3def9a42196744823a2fa48df4f9a98841ea6fff97925e3bd490182e7a1f

Observation 9099eb23-eb5e-4ff5-9e9e-b47ebb119794 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

ENTER: Event Based Interpretable Reasoning for VideoQA Videoagent: Long-form video understanding with large language model as agent

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.006658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:4c6b1832a3cfa7325cfb33bd8218f705f764cf736a004b9fe8ebaf0a8263194d

Observation 9a661515-dfdb-41d3-9f13-36c6a28d6340 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning.

ENTER: Event Based Interpretable Reasoning for VideoQA Internvideo: General video foundation models via generative and discriminative learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.016989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:3483063664f8c06c1becf39ca1fc793991ad1b6762eb4f82eb4f84e8125da02c

Observation 49b4b98c-63ed-4160-b8e6-65a74784b721 · outbound

This paper cites Stair: Spatial-temporal reasoning with auditable intermediate results for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Stair: Spatial-temporal reasoning with auditable intermediate results for video question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.881287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fb14290e97c599ed17e519b1cfc3427002372fcd29456010ce0b7bbccd0dc851

Observation 8771a3d3-480f-4c70-b78c-9842f3196bb3 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.023611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:e040d3f2b60af5d9a96809a6b4b615681a79881d5f3064d4ab329c161bbee94b

Observation be423fa0-f946-4500-907f-1795e19430c9 · outbound

This paper cites Freeva: Offline mllm as training-free video assistant.

ENTER: Event Based Interpretable Reasoning for VideoQA Freeva: Offline mllm as training-free video assistant

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.995907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:70a05fc83200dc5775f3310f1dc825660ce0cd4dbd18345b04ca62f1573ba2fa

Observation b2282c50-4b72-46c9-9da6-e55757e63cff · outbound

This paper cites Next-qa:next phase of question-answering to explaining tem- poral actions.

ENTER: Event Based Interpretable Reasoning for VideoQA Next-qa:next phase of question-answering to explaining tem- poral actions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.976858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:a32bf7b169826af3b3488d33b800b87b64cb5451ac354fb46684bfacf1a22def

Observation 775fe7ea-d826-41cc-933c-ec4acb7e6195 · outbound

This paper cites Next-qa:next phase of question-answering to explaining tem- poral actions.

ENTER: Event Based Interpretable Reasoning for VideoQA Next-qa:next phase of question-answering to explaining tem- poral actions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.955730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:7d19e737a41aea184472686e0222174e55c6f7a02ad6e055b31d012b3b042683

Observation 660226c7-617c-45c3-a678-7215f34516a9 · outbound

This paper cites Slowfast-llava: A strong training-free baseline for video large language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Slowfast-llava: A strong training-free baseline for video large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.969341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:eb553d57b5b8c29b0e4448decdcb5ad504769272f68ff3e90e9cb8099c78e497

Observation 5e76f76e-502a-4372-948e-95ed9fa1725a · outbound

This paper cites Panoptic video scene graph generation.

ENTER: Event Based Interpretable Reasoning for VideoQA Panoptic video scene graph generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.999054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:008256243040e3d579c01e0cf51ca76973add7ed59538d03b709aeded8dcb448

Observation 31b81c97-78c9-4879-9f44-5ce2586c6d0d · outbound

This paper cites Mm-react: Prompting chatgpt for multimodal reasoning and action.

ENTER: Event Based Interpretable Reasoning for VideoQA Mm-react: Prompting chatgpt for multimodal reasoning and action

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.959141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:36f43527a6748f5161241ce79fc7bfd5a64ea083c440000645be72d7844e09b0

Observation 54c34af3-7868-4e7b-9ba1-0e3e9e8c4dbf · outbound

This paper cites Ayyubi, Kai-Wei Chang, and Shih-Fu Chang.

ENTER: Event Based Interpretable Reasoning for VideoQA Ayyubi, Kai-Wei Chang, and Shih-Fu Chang

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.003171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:512853013fe814f924a35e7ca2695580a227368ccb8dfe7baa1ff0ae60fb8c3f

Observation 9bf4f48a-f555-4c4a-bbe0-fea2236f7e89 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Self-chained image-language model for video localization and question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.020206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:859554485732e686022b294c14a8c005b066b42c0a7838f381062ad93c0b9cfb

Observation c1a98543-ef4a-4ec6-be37-0ba32ee7884c · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.975881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:30d0565430c4af8e6426134c94aecc71f9bd101e5fba2b3f73929a938164b6e6

Observation 0566ea3d-888b-49b9-9fe6-484a693fc0b0 · outbound

This paper cites A simple llm framework for long-range video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA A simple llm framework for long-range video question-answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.962199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:c6880196ac5a16e955b973d6b4d61eefc886d64d3a27e3a79e0893200148c305

Observation 50ebef7d-9e3f-4eae-bf1b-6ce8e5466c5d · outbound

This paper cites A simple llm framework for long-range video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA A simple llm framework for long-range video question-answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.968679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:363ee79c65e4f304534fdaecd7f5f09b0e4adf21f808b41e072a3e0c94eafef8

Observation 1dc9fa09-9994-4065-b641-d050ffdbabe1 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.985768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:6ca3c8a6174b1e5058ad451a1ac0d5e2f0daf38153e44eecac2dd2b7ca93d19f

Observation a4353d16-04fb-4e7c-81b0-313e5c1399cf · outbound

This paper cites Base" component refers to cases where the generated code operates without triggering addi- tional modules like the.

ENTER: Event Based Interpretable Reasoning for VideoQA Base" component refers to cases where the generated code operates without triggering addi- tional modules like the

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.972387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f25ef04b05dfa20124f390b5c65f912c19ad96a5cc57aa4de00c31444d25d7cb

Pith citing papers

Observation e091819e-610d-4270-a066-990a1831f0e7 · inbound

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks cites this paper.

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks ENTER: Event Based Interpretable Reasoning for VideoQA

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:31:13.198639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:41:42.061219Z digest=sha256:b0dfad255c2e6efae96e0b4e0f768bba76a5e85ffec3443ca824aced8ded6da1