Pith. sign in

Paper Citation Record · LEDGER

VCA: Video Curious Agent for Long Video Understanding

As of 15 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2412.10471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10471 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:51:15.431541Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:05.146191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:34:19.511206Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d6d84b9-b5e6-4fe4-8014-55e0c34eba9f · outbound

This paper cites GPT-4 Technical Report.

VCA: Video Curious Agent for Long Video Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.041267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.041267Z digest=sha256:37b9042ecea4247f926739dc03bbaff84a98a872976a3747bf6fc66da3b14c1e

Observation 51523699-844e-4da4-85e4-bc474baadc22 · outbound

This paper cites Pixtral 12B.

VCA: Video Curious Agent for Long Video Understanding Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.046293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.046293Z digest=sha256:18dc93fb79a31c3bb7112160765fb88e58b7ba0a8789d8de16c8a832117aaff4

Observation 261889dc-d7b7-4bc8-bffe-c72d02f478d1 · outbound

This paper cites Vivit: A video vi- sion transformer.

VCA: Video Curious Agent for Long Video Understanding Vivit: A video vi- sion transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.051066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.051066Z digest=sha256:a4ce9a74b8ea927348d95ea13b13b5a0642deb524ff373c2bbec5dd0bff95521

Observation 6865a22d-37c3-4bdd-838f-4988f968c1b8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VCA: Video Curious Agent for Long Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.055398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.055398Z digest=sha256:368350aa3fbbb811e1ebc6a79db028808005eaa0b7ec8c40e0be32d8835a57ea

Observation e3fc42dc-693d-4ecd-b0dd-fd17260f792f · outbound

This paper cites Memory consolidation enables long-context video understanding.

VCA: Video Curious Agent for Long Video Understanding Memory consolidation enables long-context video understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.060605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.060605Z digest=sha256:2cfd6fc0b5c0b18673662b2bb354eaa36d9173521d3966d0902411cf3e8abb5e

Observation 8515d0a1-12a8-4620-a969-eb49c4c311b9 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents, 2023.

VCA: Video Curious Agent for Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.064543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.064543Z digest=sha256:bb1f34d1a4863d4b032766c1ffd54fd102973c0a6b45dea47fd60231f7c564f6

Observation af3abac2-5cce-4289-af18-bfc1dc724b49 · outbound

This paper cites Is space-time attention all you need for video understanding?,.

VCA: Video Curious Agent for Long Video Understanding Is space-time attention all you need for video understanding?,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.069059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.069059Z digest=sha256:363f2fc49185a9133f925463d461980ad5768cf133fca188ec875551448c6a3d

Observation d14f7899-0ab9-4bfd-a0ea-b9a7ff6bda41 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset, 2018.

VCA: Video Curious Agent for Long Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset, 2018

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.072865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.072865Z digest=sha256:d0d7cb2b492ba56ec5291058919a5bb0f7eeca4cac2241789e18c9f4cec571ea

Observation ef648ba7-9177-4563-a7ce-4c4bc552e631 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

VCA: Video Curious Agent for Long Video Understanding PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.076893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.076893Z digest=sha256:79abd8db1f36034264274c82d5f747ecc76bedef8d6e69817698a58d5e59b1c9

Observation cd7a1a4a-20e4-407a-8323-d6ebf5831f0d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VCA: Video Curious Agent for Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.082494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.082494Z digest=sha256:86c660d2a4eb9a927188c85eb26fed2c4da85a5f3d2cf88279b14c781eb5810d

Observation 3732499b-b4b4-46c0-94be-ea5a3901ea72 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VCA: Video Curious Agent for Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.087106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.087106Z digest=sha256:e2a27e350afe49dddf44fa6f2de60bb5a933394b43b46b43e9480904c2307456

Observation 0fd94c65-4c7c-4883-a74b-5ef4939ffc3d · outbound

This paper cites Long Story Short: a Summarize-then-Search Method for Long Video Question Answering.

VCA: Video Curious Agent for Long Video Understanding Long Story Short: a Summarize-then-Search Method for Long Video Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.091438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.091438Z digest=sha256:6948a9108be088623d7bef5560e1a7d85dec687c4f4fe228d22e74edadb64f6a

Observation 3b5b0cb3-e987-447c-9c90-fb5c27115ec2 · outbound

This paper cites Neural mechanisms of selective visual attention.

VCA: Video Curious Agent for Long Video Understanding Neural mechanisms of selective visual attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.096386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.096386Z digest=sha256:6e32c5e517e863869e1032b0398c8b945a12f8236e1623378f20710d65304613

Observation 15b5829a-899c-4632-be9b-04fe5d604886 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, 2021.

VCA: Video Curious Agent for Long Video Understanding An image is worth 16x16 words: Transformers for image recognition at scale, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.100305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.100305Z digest=sha256:75a4670bb467ecf6b47a57a691aef9afd023ebfcf421e5565d2593e60cc0cf31

Observation 5577ee1c-92ed-47d2-ac91-7265309d36e2 · outbound

This paper cites Agent ai: Surveying the hori- zons of multimodal interaction, 2024.

VCA: Video Curious Agent for Long Video Understanding Agent ai: Surveying the hori- zons of multimodal interaction, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.104068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.104068Z digest=sha256:795746c307c52b6d73bb51429795383132c31b332658d66a71fc8b11a99ec4b8

Observation a47980ef-9da4-4686-8445-036e384d9c6f · outbound

This paper cites Multiscale vision transformers, 2021.

VCA: Video Curious Agent for Long Video Understanding Multiscale vision transformers, 2021

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.108418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.108418Z digest=sha256:450a18d67a8ef5f024b04567220581f3eca2a62a22a8022976e5db986e064fe8

Observation dd4d0e99-9438-4a05-ad98-7580bec6bd84 · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding.

VCA: Video Curious Agent for Long Video Understanding Videoagent: A memory-augmented mul- timodal agent for video understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.112260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.112260Z digest=sha256:33bf31c42c351252291afc32067e9a73b0238e525271f206b149a4b4e74590bd

Observation 216eef99-e433-4e15-ba2d-d43f10802772 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

VCA: Video Curious Agent for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.116123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.116123Z digest=sha256:61aa5e477bc6c0028e53d8ceae7f5d1c41f4c0873e32d26ecaff2bb4ecd5df72

Observation 8f7901a5-8f93-4ca5-b274-028bec8a148d · outbound

This paper cites Convolutional two-stream network fusion for video action recognition, 2016.

VCA: Video Curious Agent for Long Video Understanding Convolutional two-stream network fusion for video action recognition, 2016

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.120225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.120225Z digest=sha256:d5277b05dec16f994720739c3d45b3129da6d214870b118e51f401ce6f956140

Observation e9006ca0-3711-47a8-a612-5c9a2078e9b1 · outbound

This paper cites Slowfast networks for video recognition.

VCA: Video Curious Agent for Long Video Understanding Slowfast networks for video recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.123969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.123969Z digest=sha256:b67586cf0852057120d17d5850b5efb25064859817ffd52001749c3ac80860a3

Observation 7ef14b41-08f2-42f8-ad88-f8edb94d5787 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VCA: Video Curious Agent for Long Video Understanding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.127510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.127510Z digest=sha256:29f196c5b75ce63b7d125423f34b41467866d2ac51e8b8bc2b9c7b4c204277c8

Observation e3389d05-c3d7-4a46-96bb-45eda648ba62 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VCA: Video Curious Agent for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.132599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.132599Z digest=sha256:f2945afdf3d9a91bc5088e242b6c276e7e925052f36be14979e5a9eb863e9eb9

Observation b5ba67c1-aebc-4b9b-bec4-a041a6acb94f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VCA: Video Curious Agent for Long Video Understanding Ego4d: Around the world in 3,000 hours of egocentric video

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.586010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.137095Z digest=sha256:c98f2d34401c5474f738f58f4f3a51c2a398bcbc69b4a99102de2057bed10fc1

Observation c2a2dc18-7077-474d-9a6b-e4de37afa7f2 · outbound

This paper cites Cogvlm2: Visual language models for image and video understanding, 2024.

VCA: Video Curious Agent for Long Video Understanding Cogvlm2: Visual language models for image and video understanding, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.573907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.142023Z digest=sha256:0b7f6498d11a7dbae5cdc976f98312660387d7e74f6a637c3d78158a837f8ca6

Observation e8fa3bc7-3dbc-485e-8264-4a12f7377600 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

VCA: Video Curious Agent for Long Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.146078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.146078Z digest=sha256:efde1eb5a30c670849bdce9f191e24670090f1d239d6004f098cd5a8684d6910

Observation 4f6678cd-d257-4b71-af64-7cb5170f03a5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

VCA: Video Curious Agent for Long Video Understanding Cogagent: A visual language model for gui agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.150921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.150921Z digest=sha256:ed095d0258c3682ec312840e1827644d419841a81c9d7f7b9b2d4a1a4bd9b965

Observation 2d351d7f-b698-4c05-853f-767d1c889ed0 · outbound

This paper cites GPT-4o System Card.

VCA: Video Curious Agent for Long Video Understanding GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.154598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.154598Z digest=sha256:ff572504a78eb0ad476ea70e6ef3be2ed612b70af46306dc4753f502b8e1c30e

Observation 477c62d8-1208-4a0a-9f42-4e31cb5725ec · outbound

This paper cites Videowebarena: Evaluating long context multi- modal agents with video understanding web tasks, 2024.

VCA: Video Curious Agent for Long Video Understanding Videowebarena: Evaluating long context multi- modal agents with video understanding web tasks, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.555595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.159220Z digest=sha256:e09364e5636a7556e48c6d2b91c7ceb70c285fc0dad7cfe9700b3b7230ad3794

Observation f44c37ae-b272-4ebd-8bbd-5200b06ef533 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing, 2023.

VCA: Video Curious Agent for Long Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video under- standing, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.544086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.164463Z digest=sha256:9ae3f0f6e7e5b7a896d1da3821c8e8c4691f4b9d0411cd16fd00667e258a6258

Observation d3457cfb-8e41-482b-8dd7-fa7ba192d3bf · outbound

This paper cites Visualwe- barena: Evaluating multimodal agents on realistic visual web tasks, 2024.

VCA: Video Curious Agent for Long Video Understanding Visualwe- barena: Evaluating multimodal agents on realistic visual web tasks, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.532383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.168510Z digest=sha256:2c1901a00ab93898623ad2f606e59f2e04b0ba839335ba1564505442d4ffe7fb

Observation 5005d86a-8d58-41a3-8e5d-9c944600af86 · outbound

This paper cites Text-conditioned resampler for long form video understanding.

VCA: Video Curious Agent for Long Video Understanding Text-conditioned resampler for long form video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.519538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.172707Z digest=sha256:ea0fe51720e5539421754dbb20cf3a0791de165b63d7a9487beca071597e6cde

Observation 880f3b83-c31c-426d-8fcb-fba47ec844e0 · outbound

This paper cites Mmctagent: Multi-modal critical thinking agent framework for complex visual reasoning, 2024.

VCA: Video Curious Agent for Long Video Understanding Mmctagent: Multi-modal critical thinking agent framework for complex visual reasoning, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.508220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.176449Z digest=sha256:cecb81a24e0da8929a989be2698108dc9a9ab99b6cdc3a6e5b3b3f24a88e3342

Observation f21a43c9-f4a2-4950-abe8-2db653c866ff · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VCA: Video Curious Agent for Long Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.496358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.180432Z digest=sha256:1301c54fac8d49b7d3336aa2fd104515110a500179314597c864c579e6bbbb07

Observation 35849f4b-8d40-4b19-9e9c-ed9b54d7e59d · outbound

This paper cites Llms meet long video: Advancing long video comprehension with an interactive visual adapter in llms, 2024.

VCA: Video Curious Agent for Long Video Understanding Llms meet long video: Advancing long video comprehension with an interactive visual adapter in llms, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.484433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.185606Z digest=sha256:c078c173b1cd338b3cc2c415880dfa3426ff3e5155efa4f9150660c073730d36

Observation 2c05279b-68fe-4e6a-a474-16ab5fe9c366 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

VCA: Video Curious Agent for Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.472714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.189991Z digest=sha256:ef642c887f90a001a396d168ca47216f82fba6a822d4d46bcb1f1a20c7026327

Observation 3665cbd0-7ea9-40dc-ad99-aa7113cadbbe · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VCA: Video Curious Agent for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.193730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.193730Z digest=sha256:ed13a0a24f6d25f0c3a5664e524b7aeaa5b9c693fd4df23609c7ca11892fdab0

Observation 3bbff455-264e-455a-a1b6-0bf79aa81699 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

VCA: Video Curious Agent for Long Video Understanding Vila: On pre-training for visual language models, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.198407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.198407Z digest=sha256:f04d5616e931a25e39eb31329327df51d1b72a0bb8c7c8faead6284241ca3ba1

Observation 46a8d58a-6029-4d54-b466-009ec2bf1a2b · outbound

This paper cites Vila: Efficient video-language alignment for video question answering.

VCA: Video Curious Agent for Long Video Understanding Vila: Efficient video-language alignment for video question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.453454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.202695Z digest=sha256:405d4246007f84c05ceb61d3eeb50d42802c37faf824dea87c802900972f6699

Observation 6a1e798b-10dc-479c-86fd-cc36d57bd13f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

VCA: Video Curious Agent for Long Video Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.206935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.206935Z digest=sha256:7373c9fa9c63c2eda324cf33bfa1ca581475c833cd67172cba9cde4f2a108e44

Observation 612e2c43-93f8-4c0b-811d-b46df35b197b · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

VCA: Video Curious Agent for Long Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.210641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.210641Z digest=sha256:0b716f462a3b0e557034b5ed9e2e86b9ebe704ab4c31bd9babc827ad60888326

Observation 3c35b66f-85ab-40d0-b766-f318258234e7 · outbound

This paper cites DrVideo: Document Retrieval Based Long Video Understanding.

VCA: Video Curious Agent for Long Video Understanding DrVideo: Document Retrieval Based Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.215134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.215134Z digest=sha256:130457e1646695c4e58c1fb0744e94d5a6e3202fe5b32eea7c002ed384245e26

Observation 52fb01ea-7a3a-45ac-9755-4104a306c77d · outbound

This paper cites Drvideo: Document retrieval based long video understanding, 2024.

VCA: Video Curious Agent for Long Video Understanding Drvideo: Document retrieval based long video understanding, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.433418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.220446Z digest=sha256:2c16276448efc8d637f20b8d45a6fc4a2d4bb18af1673e7c817af0ed0af554d4

Observation 46e6c3fb-224b-45ad-9dca-67673151045b · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

VCA: Video Curious Agent for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.224560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.224560Z digest=sha256:042d5fb15578e8634fa3db75f7613786a5d44733a1f93a55042dccc5c9621c06

Observation c4cb7f4b-97b4-4be9-b505-2b0c4a993db4 · outbound

This paper cites Beyond short snippets: Deep networks for video classification, 2015.

VCA: Video Curious Agent for Long Video Understanding Beyond short snippets: Deep networks for video classification, 2015

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.422055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.228945Z digest=sha256:5f06e037f2f8ad4d7e1801628ac46f41156831f841cadec4282ec58c811b34dd

Observation 47a4c27d-d36b-4cc9-8205-b49832ef4134 · outbound

This paper cites A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames.

VCA: Video Curious Agent for Long Video Understanding A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.410532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.232693Z digest=sha256:776343939cf87950f9e5b9516c4800cbd66d21e9428163255d59369d9057a640

Observation de5b2391-1995-4716-975e-ba8516b10142 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long- form video qa.

VCA: Video Curious Agent for Long Video Understanding Too many frames, not all useful: Efficient strategies for long- form video qa

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.236923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.236923Z digest=sha256:b82137d0b8c479a2d08e3c1c6f8187739feb92f4fd1bce1b601be71c5018a6e2

Observation 5955f83d-b136-40fb-8dac-d4d57c812536 · outbound

This paper cites Cinepile: A long video question answering dataset and benchmark.

VCA: Video Curious Agent for Long Video Understanding Cinepile: A long video question answering dataset and benchmark

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.398826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.241180Z digest=sha256:d7b883fac79404bcb7eefe678c6afadab54acdc04220045d2552196745254226

Observation 00f39dc4-3c86-4e6e-8c63-c06ad75c3438 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

VCA: Video Curious Agent for Long Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.245087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.245087Z digest=sha256:905b0fabe167ec529efdc84944455e62e12013614394488f4a4dceefe8c04b8a

Observation 60d0596a-2ccc-4e89-ae0e-6d3cfe265492 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

VCA: Video Curious Agent for Long Video Understanding Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.387086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.249940Z digest=sha256:6e715114c8c8ffb89a8da138b6c1d97fa59a58212dc03e289d8594a508cac8b1

Observation bb21bf32-36d8-4c41-87be-b037535dab71 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

VCA: Video Curious Agent for Long Video Understanding Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.375042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.253919Z digest=sha256:bafbede061154e2f12b1faf0f41fdffa17328bf2da24926ab2c0b722ae2d84ce

Observation 4f4958a5-1ca7-450c-be0e-eb9382c7b21a · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

VCA: Video Curious Agent for Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.258067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.258067Z digest=sha256:cdcc20d9bbae8bb9ebfc45b81f6d8d08635267604d15ee298f8ac2bc43df3206

Observation 8b4a6df2-4941-45a3-b2ea-c2d322bc254d · outbound

This paper cites Videobert: A joint model for video and language representation learning, 2019.

VCA: Video Curious Agent for Long Video Understanding Videobert: A joint model for video and language representation learning, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.362523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.262654Z digest=sha256:7d4733946d29164e3bd27ba48efab9296fc4bdfb25f9fbdfc4979eed3cb81ed0

Observation d6b354a4-655a-47d7-95f2-dbe1febb5d04 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

VCA: Video Curious Agent for Long Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.266437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.266437Z digest=sha256:b0acfb22de071b7d9cab7f4710e31de875374e26671e818102967ead2c52d0e1

Observation a894cc53-7abf-4b75-adbb-d886e7773247 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

VCA: Video Curious Agent for Long Video Understanding Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.350897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.270982Z digest=sha256:bbb7ba17cfb883a44963a3405c3fc0439e2f9352c2c6a78da3ea51d940c0026e

Observation 814c598b-edbf-4d3c-9d89-4f4317151126 · outbound

This paper cites Video understand- ing with large language models: A survey, 2024.

VCA: Video Curious Agent for Long Video Understanding Video understand- ing with large language models: A survey, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.275208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.275208Z digest=sha256:a159cf02d254d1cd9037a6df67fa53cbafdc49c369fd9c51bb169ed0ccbc9f5f

Observation de543687-321a-44ee-b080-cbf3b9d5466b · outbound

This paper cites Movieqa: Understanding stories in movies through question- answering.

VCA: Video Curious Agent for Long Video Understanding Movieqa: Understanding stories in movies through question- answering

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.279070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.279070Z digest=sha256:1b5a154ff0ed9a64158f21b9136b1f2a42f338ef86aa4f232236f3fc68b8fc11

Observation 3fe5cd61-36ab-42be-b37c-532afffd49a2 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VCA: Video Curious Agent for Long Video Understanding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.283036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.283036Z digest=sha256:deaed8ebfd729a6cfec2557724ed34f7cbad1acb1904fba704ab6bc83a145cf5

Observation c13a532f-41a0-4504-8bc7-e707e9c4257d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VCA: Video Curious Agent for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.287162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.287162Z digest=sha256:884a79acded3eddba69d745e25f56e54b49d4e9b6cd235389035259bc73c2129

Observation 80008c2b-7657-4e3f-8dde-b1a662f2e23d · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training, 2022.

VCA: Video Curious Agent for Long Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training, 2022

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.291100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.291100Z digest=sha256:c726df0889df85d6bf0c4c98bfdb2661f3006183668f42e6cdce934c34f51ce2

Observation 55b5e6cc-f345-43c1-a588-10fcc4ef2960 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks, 2015.

VCA: Video Curious Agent for Long Video Understanding Learning spatiotemporal features with 3d convolutional networks, 2015

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.317405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.294799Z digest=sha256:50fe11358687407775f3cabe753d5fefff6456ff5c944c49c92b25d9d2796a72

Observation 63c54d6f-f1ad-4f00-875c-22a6580deafa · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recogni- tion, 2016.

VCA: Video Curious Agent for Long Video Understanding Temporal segment networks: Towards good practices for deep action recogni- tion, 2016

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.305767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.298494Z digest=sha256:c55463f455c074770590e99f59afe59f1acea78762cb6ede478bae33455edc05

Observation 3a4a4cbc-e6d5-4829-babc-338979a337a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VCA: Video Curious Agent for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.306717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.306717Z digest=sha256:31ae0903603eb53539d29f964be2d03925b60f28c037c5f4942a91fc413c045b

Observation 76b730db-99aa-4dba-accc-95a730fc01ee · outbound

This paper cites Lvbench: An extreme long video under- standing benchmark, 2024.

VCA: Video Curious Agent for Long Video Understanding Lvbench: An extreme long video under- standing benchmark, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.294283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.310302Z digest=sha256:c2083b9ea8063f86e0006e32823f0053ea2e7e6f85dbbbc80069b0d5d72aeac1

Observation 242bfa80-faa5-42eb-a0d5-48c39bc8d142 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

VCA: Video Curious Agent for Long Video Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.282230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.314328Z digest=sha256:4d9b8fe7431defe1e4c00d15fd4de75ba62e0d0360e40afb44a05945a030fc6d

Observation 8b1b3b59-05ca-4048-9817-1904fb0b4e47 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VCA: Video Curious Agent for Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.318360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.318360Z digest=sha256:b8e7b129114b0129d4761390223481eeb0a85fccf7a88ed1809c918a910635ce

Observation 69f645a1-effd-44c7-9636-206d41f72c32 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

VCA: Video Curious Agent for Long Video Understanding VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.322474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.322474Z digest=sha256:400e7a60b08414908121bcc29d6e4de97dcd5c56dd902f47ad5d790ca1f3c6d6

Observation 91936784-82be-407c-bdcf-d75ed8a71c27 · outbound

This paper cites LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos.

VCA: Video Curious Agent for Long Video Understanding LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.326991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.326991Z digest=sha256:3d202fdc881dfe86f20fa8bd92710dc8c301c2c2bc899812c49085baa80e1fb3

Observation 02f7f34a-aac9-4cf7-bd8e-2cb94f5ac74e · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos, 2024.

VCA: Video Curious Agent for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.270861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.331488Z digest=sha256:4bf93180a36164b95a02e2effb568f667bc0978f677c4a6b2069c5740b9b96c2

Observation 2bdc60be-9dd0-4229-9bab-ff8514d67e17 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

VCA: Video Curious Agent for Long Video Understanding Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.259206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.336085Z digest=sha256:b6ad55e1fe57f4933c05579862443a5373d8807df7ac7ea4ee03d99aab7f5d61

Observation fc647345-2c72-4960-b58b-e6a70b1fb2c7 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

VCA: Video Curious Agent for Long Video Understanding Longvlm: Efficient long video understand- ing via large language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.248534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.340529Z digest=sha256:baa971575cc4f8a2ec43a956aa1c584084503925e096fe1a40668011bde45f38

Observation 1b19d0b9-2fa9-4859-9796-7de444f4911e · outbound

This paper cites Towards long-form video understanding.

VCA: Video Curious Agent for Long Video Understanding Towards long-form video understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.236838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.344278Z digest=sha256:95d63bf2545f06ed6edf5bd47e38f74c0fbc0445b2b8797c77e80fc2408c91ba

Observation 70d9bbdf-f295-4f3c-8cfa-3b6b443c9381 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

VCA: Video Curious Agent for Long Video Understanding Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.348677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.348677Z digest=sha256:1a32591917d043c52b757ff51ce57863333acfe46d24cec50573f9e0272be85e

Observation 34a64946-a01d-4c10-a33e-cb46b526bb87 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

VCA: Video Curious Agent for Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.217809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.352825Z digest=sha256:2b1b94ca146c26e999e70315d518ff23d5e1ff41f7e80c4ae1723272707db5c4

Observation ebff8ae7-e013-4bc4-9187-025f8f21f7db · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification, 2018.

VCA: Video Curious Agent for Long Video Understanding Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification, 2018

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.206124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.356838Z digest=sha256:0f92a7774bfda9547da1f788a8bcd310628b7ec63998a0f8693fa4f2a1e9677c

Observation 9d5cf55d-b454-4c7f-bd27-81a0d131a94a · outbound

This paper cites Openagents: An open platform for language agents in the wild.

VCA: Video Curious Agent for Long Video Understanding Openagents: An open platform for language agents in the wild

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.194348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.360768Z digest=sha256:3f2ed9034db7da56204e1befd96a70c196f6fb386c4a870358c0490a57fba7f2

Observation e1dbfeea-7d32-4be2-9840-2af4989ece2e · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

VCA: Video Curious Agent for Long Video Understanding OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.364479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.364479Z digest=sha256:67f8a6c6f0d2ab659b0061242b5c027fff7adb7bec4f81a05aa4a1d66e094131

Observation 4883504d-1107-4b07-a411-ea6c46adb681 · outbound

This paper cites Retrieval-based video language model for efficient long video question answering, 2023.

VCA: Video Curious Agent for Long Video Understanding Retrieval-based video language model for efficient long video question answering, 2023

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.368788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.368788Z digest=sha256:0bb63c5b06a4e09ce9545417cb665e4fe001ae73c1b09287c760ea3ef2343e96

Observation 54a605b3-565d-4a03-ae31-743de26d7b01 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.

VCA: Video Curious Agent for Long Video Understanding Webshop: Towards scalable real-world web interaction with grounded language agents, 2023

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.176026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.372527Z digest=sha256:eea8cdc67c5020a4ca6422703422dc681b089a0cc720ff32ddaeb9df5e12fb4c

Observation f9b804f9-515c-4df2-9b02-ae48e13da4cf · outbound

This paper cites Self-chained image-language model for video localization and question answering.

VCA: Video Curious Agent for Long Video Understanding Self-chained image-language model for video localization and question answering

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.164614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.376675Z digest=sha256:a1b1f56a86244e7bc42f501c739d54b9136318ea50448ad6e27ed6502dda4480

Observation a52fe91e-011f-4564-bbf5-24f989ed223d · outbound

This paper cites Self-chained image-language model for video localization and question answering.

VCA: Video Curious Agent for Long Video Understanding Self-chained image-language model for video localization and question answering

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.151829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.380287Z digest=sha256:c9744af5058b3865a6e7049f9bb2f8b8cfdfb299bbe91f053370dd32dc6ffd97

Observation a30d6094-19d9-495f-99a7-92a78f2b1a49 · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VCA: Video Curious Agent for Long Video Understanding A Simple LLM Framework for Long-Range Video Question-Answering

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.384737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.384737Z digest=sha256:93a40683c544d52b8fa916f7d8ae5810f3c662729cf02ee9d91b04e16f0b1a3b

Observation d885ae91-1cc3-4a58-a0a5-14228144837e · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

VCA: Video Curious Agent for Long Video Understanding Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.139526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.389321Z digest=sha256:d16e9c95f96ea97d98fc8ba7df9fa100be3e198c669ef3b26d4528aa0bbb7fbb

Observation 786b2f4d-aad2-41d3-8079-eb4383c4a47a · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

VCA: Video Curious Agent for Long Video Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.394379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.394379Z digest=sha256:87c3aed6ad21e765055b671403860ea134d844b7519161da387e07ca4921290c

Observation 0e9b1cb3-0644-436c-a0af-3a7c2d6e3a39 · outbound

This paper cites Long Context Transfer from Language to Vision.

VCA: Video Curious Agent for Long Video Understanding Long Context Transfer from Language to Vision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.398268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.398268Z digest=sha256:28a6fc8c5d5807937c526706eaf178cb7b14194c6fa9ccf8a85990793e8ca8cb

Observation 82e4a9ea-bc64-4391-a384-3af8d474d0fe · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VCA: Video Curious Agent for Long Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.127777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.402285Z digest=sha256:7981a0c7287ec7c3a4d7b418c484a6c300b7468144820d34dc9c79b262a7f965

Observation fecf0dcf-b274-4a34-b8e0-f75cb72ceae1 · outbound

This paper cites LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration.

VCA: Video Curious Agent for Long Video Understanding LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.406260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.406260Z digest=sha256:a4f6c53034771926b2132669d165df6b571f996cdc39403bf803c9a8c97d5dd5

Observation 91fd47c0-fed7-4208-a2a6-024676ca805b · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

VCA: Video Curious Agent for Long Video Understanding Videoprism: A foundational visual encoder for video understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.115324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.410133Z digest=sha256:3d14e6ccc0a99fc00a9059446e940e0d9d27b1be080e43c1c85010403e67bb09

Observation 20de5e04-6ac3-42ff-b446-a0f1f4577a82 · outbound

This paper cites Actbert: Learning global-local video-text representations, 2020.

VCA: Video Curious Agent for Long Video Understanding Actbert: Learning global-local video-text representations, 2020

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.102719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.413804Z digest=sha256:fd665c37df285acaa1fb80deed4d4573b21d83690091a6809191ea58fd85bbd1

Observation 214f7c4f-2c7f-4f29-88a3-13f88ff0975b · outbound

This paper cites We include the prompt on how the reward model (R) generates relevance scores in Fig.

VCA: Video Curious Agent for Long Video Understanding We include the prompt on how the reward model (R) generates relevance scores in Fig

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.090015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.418290Z digest=sha256:580e35f6c14375208d350c4e98bd31a47a9db0ff9768d04771adc15d63082fd0

Observation 8218422c-8a9a-43e3-a93c-e35cafca49b7 · outbound

This paper cites Dataset Dataset Avg.

VCA: Video Curious Agent for Long Video Understanding Dataset Dataset Avg

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.077560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.422812Z digest=sha256:94aa15be68f45b99ce0a3e596b0d9ff71ffdc2289ebe1794fe3aed930d568e48

Observation 54a0f9f9-e39a-4c4b-87a9-1b83b3adbb29 · outbound

This paper cites Experimental Results As discussed in Sec.

VCA: Video Curious Agent for Long Video Understanding Experimental Results As discussed in Sec

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:51:16.063874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.427174Z digest=sha256:e9f2597ee29ef208d424f45bd11465f2a0cfb5a6e145452b38b20e7bfaf6f5b9

Observation a1a8f6b0-82e6-4a5c-8c12-10ee970cd76e · outbound

This paper cites 5.2, in this section, we investigate the common failure cases of our framework, aiming to provide data points and insights for the future research.

VCA: Video Curious Agent for Long Video Understanding 5.2, in this section, we investigate the common failure cases of our framework, aiming to provide data points and insights for the future research

Reference 93

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T16:51:16.050757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:51:15.431541Z digest=sha256:c8285ba7c5a3013b249ab807913555b6e152c1f33c50cb407478028dab389a40

Pith citing papers

Observation 1e00a398-26eb-43cb-8fe3-3094f39cb71d · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VCA: Video Curious Agent for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:05.146191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:05.146191Z digest=sha256:74a423ff00f1dcbcfd562847e55adbe9b64361d4a22665ef1cf9ad04c9d60073

Observation d19217dc-6bd2-41ea-805a-31257957c3a9 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs VCA: Video Curious Agent for Long Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:19.512782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:94c4dd0156ca3a8fdb023dffad6e88d3e5a96c0f0d9558334c414bf7a27474e4