Pith. sign in

Paper Citation Record · LEDGER

Understanding Complexity in VideoQA via Visual Program Generation

As of 17 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 2 inbound Pith citation observations for arXiv:2505.13429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13429 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:18:53.143452Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:32:22.239838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:34.105509Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89d10c2e-1c25-4f44-848f-5fe60e5a602c · outbound

This paper cites write newline.

Understanding Complexity in VideoQA via Visual Program Generation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:50.991012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:50.991012Z digest=sha256:a9c2c4461dd289d02fd6d060898d714cdcecd018524c7c9fd20f87231c7a9b37

Observation 0bd956ba-abd7-406e-9e0a-4b1cdbb95458 · outbound

This paper cites A shared neural substrate for action verbs and observed actions in human posterior parietal cortex.

Understanding Complexity in VideoQA via Visual Program Generation A shared neural substrate for action verbs and observed actions in human posterior parietal cortex

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.059253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.059253Z digest=sha256:ead96061ae7b1b9c39438b70577d5a36aec9a07a2526e44d82d7796d7db3094c

Observation 4e4bd6e0-9009-433b-9681-1475a3df77d4 · outbound

This paper cites Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications.

Understanding Complexity in VideoQA via Visual Program Generation Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.064334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.064334Z digest=sha256:b268bc94f3b9a6377aa247354104d0b7e7cd1d6de8e0a6f6e1ada55e486a3f18

Observation cc8c7136-bec3-4ceb-bc97-11ae44aaad58 · outbound

This paper cites Neural module networks.

Understanding Complexity in VideoQA via Visual Program Generation Neural module networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.068935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.068935Z digest=sha256:7fe51cb477e52a2eae1157fb6989f60ceb012525dbaf06b8ef0675b714ca6501

Observation 8a7bc1fc-7307-4c53-8d3b-808d616c1fc6 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Understanding Complexity in VideoQA via Visual Program Generation Is space-time attention all you need for video understanding? In ICML, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.072802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.072802Z digest=sha256:fb1d71e2f9e95434508c230b7c83e9a5912916cd6cf5b8ca9cc06c73e5eaf810

Observation a0cf77f5-ca21-45f3-b652-b62bcca0071f · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.076952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.076952Z digest=sha256:fa179b12d87d40408b810619e5bf2fb9d90e58e71f9c17c4e131bdad1d828da4

Observation f7092f22-b78f-4178-9d96-c073b9fd1a44 · outbound

This paper cites H., Lu, P., Nocedal, J., and Zhu, C.

Understanding Complexity in VideoQA via Visual Program Generation H., Lu, P., Nocedal, J., and Zhu, C

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.081807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.081807Z digest=sha256:0cd09f36a2225eab0232c6e7174ff2270fa5d87cd0f52d1844fb705f2f9a4345

Observation 27f51f1a-24da-4516-addf-e93ff3112c37 · outbound

This paper cites ActivityNet : A large-scale video benchmark for human activity understanding.

Understanding Complexity in VideoQA via Visual Program Generation ActivityNet : A large-scale video benchmark for human activity understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.196727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.196727Z digest=sha256:99b1adbc672ce48b41479a7c5a577f3bc6c457902aa83096c734a71075f10f04

Observation 7d088cd8-c151-44b2-b710-7a69da522442 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Understanding Complexity in VideoQA via Visual Program Generation Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.223294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.223294Z digest=sha256:b87e11d60305a987c6c7477e82ec3913a76666f451da86a8b3994a870d7f40ab

Observation 60cc9b1c-45ba-4a8b-b171-89c597e44e29 · outbound

This paper cites E., et al.

Understanding Complexity in VideoQA via Visual Program Generation E., et al

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.227980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.227980Z digest=sha256:0d8516de554545ca10a758a3266260d031545deb3b386a4c2568fb04c03d4665

Observation 40d5f1a4-85a6-4567-9246-d9a0ef99a0b7 · outbound

This paper cites and Gr \`e zes, J.

Understanding Complexity in VideoQA via Visual Program Generation and Gr \`e zes, J

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.231760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.231760Z digest=sha256:2c09254e0a709b9a7ea2c5ef95e4bb8513bde38b7bc2bca4bf191fca46234e98

Observation 1b166968-310d-4ddb-a329-fb8689f7c3ea · outbound

This paper cites Brain activity during observation of actions.

Understanding Complexity in VideoQA via Visual Program Generation Brain activity during observation of actions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.566769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.235391Z digest=sha256:19f0537c52ee79b6e54a017e2a00c309803b5e259deb356874603912c017e7d1

Observation 5887e24f-ebc2-42d2-ac0c-16ec36401406 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:56.553800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.240031Z digest=sha256:8a93c52c17cb9ccf7941ac264ef6d4bc48c2fb089a31946cba3047841a9a7b2c

Observation ca03a4fd-f291-4ce6-b059-8afaa07ecb69 · outbound

This paper cites K., Winn, J., and Zisserman, A.

Understanding Complexity in VideoQA via Visual Program Generation K., Winn, J., and Zisserman, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.540693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.244319Z digest=sha256:5288122cc0a01ba6e2e99904794cb7fc1dfb48cb7c6f5cd41ce558caa919b29f

Observation 7130db4c-09a2-4529-ae70-5129474e11da · outbound

This paper cites and Soto, A.

Understanding Complexity in VideoQA via Visual Program Generation and Soto, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.413268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.310511Z digest=sha256:681622c31c90227b9c18f4333a7177a8fa3be03c3cb52f0448efc7bfd099b7da

Observation d2e8376d-b64f-42d8-919d-f2117f9d7be6 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Understanding Complexity in VideoQA via Visual Program Generation Masked autoencoders as spatiotemporal learners

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.329880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.313786Z digest=sha256:c2f823bc485b74395a20d10fea09ae18c6da98f2cc66c57ad669596be969039a

Observation ecacacbb-e0c7-4e8b-921e-e8b30b1338f9 · outbound

This paper cites GPTScore: Evaluate as You Desire.

Understanding Complexity in VideoQA via Visual Program Generation GPTScore: Evaluate as You Desire

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.317153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.317153Z digest=sha256:d4af5c4a5649cec2ae9ee62bde093d3cf54edbbdf1b66529c69f84ba32884f3d

Observation 0fa37ad3-6d7a-43d0-bcc8-f59a618acda0 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Understanding Complexity in VideoQA via Visual Program Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.322345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.322345Z digest=sha256:b24e61c93034c915b808834db08532ac6c72f84a3c07fbbf2a8552fb2f18cdb1

Observation c108a262-3a8f-4caf-8cd2-0b48b7b7fc61 · outbound

This paper cites Recursive visual programming.

Understanding Complexity in VideoQA via Visual Program Generation Recursive visual programming

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.318898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.392721Z digest=sha256:4f9d4e0d2ac0fe7c9e0cbd3fe3c8f4b9331f3c419dab36711545967a7bd9e4e0

Observation b0e85dfe-50ca-49f7-af29-7c23e729e380 · outbound

This paper cites Adaptive computation time for recurrent neural networks.

Understanding Complexity in VideoQA via Visual Program Generation Adaptive computation time for recurrent neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.308405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.477662Z digest=sha256:0a3b9a00e926923a0478f38a0d62207f897cac5e4e070af0d0b4d1048f5201c4

Observation 1fd69439-8776-4337-9730-9ad50c48a07b · outbound

This paper cites AgQA : A benchmark for compositional spatio-temporal reasoning.

Understanding Complexity in VideoQA via Visual Program Generation AgQA : A benchmark for compositional spatio-temporal reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.183194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.544947Z digest=sha256:78daee3b30fbd20f54207af8d2146091997ad06185e7987521370695559aa5e0

Observation 0765632e-91f3-44ac-b645-9ed3a96339d4 · outbound

This paper cites and Kembhavi, A.

Understanding Complexity in VideoQA via Visual Program Generation and Kembhavi, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.174661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.549995Z digest=sha256:7324e2e5b7e1ae19b67f68422b42168ab6ffb01886e9a5bd7eb4a35b0d0cf4c4

Observation 77b9ce6b-182f-48ed-83b7-62ea30d7dc28 · outbound

This paper cites V., Sethi, R., and Ullman, J.

Understanding Complexity in VideoQA via Visual Program Generation V., Sethi, R., and Ullman, J

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.140553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.553834Z digest=sha256:77236ab51a438e6d327fb8c0434f11a7bc9628cddd1183437bad781e6f46076a

Observation bb2e49d6-2f87-4124-8aa8-6bdd98dfe28a · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

Understanding Complexity in VideoQA via Visual Program Generation Learning to reason: End-to-end module networks for visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.022028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.557803Z digest=sha256:752e7a91e02895f7a4dfb74aea840f0f67b659dcea5351c546d6c91f72c755ca

Observation db603616-1975-4174-a68f-6b4da6f4e151 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:56.011441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.562304Z digest=sha256:35befa73fa4a8d55d65d68d492fa4ccfbfeb11980e4711dfa966c54002fc7868

Observation b11e4c7e-41b0-48a1-acb0-8711aed44e0e · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.999582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.566362Z digest=sha256:df8f31e5a080a8573fb42835af4c42955b58954c57cb38dc53d11dd53813ef84

Observation bad9a63c-f821-409f-9078-d7a0fd2d04a1 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.990558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.571755Z digest=sha256:53482df22d8c9dd21621874d3454e76ffcc9e5be32b5bfbec8dce7bd57f9d0ff

Observation 7be27510-0989-4a80-a577-fae7d25c3287 · outbound

This paper cites and Han, Y.

Understanding Complexity in VideoQA via Visual Program Generation and Han, Y

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.831349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.636026Z digest=sha256:c9190d600ea315c38e9e6b79d2b57c6e34bfdd9ae7c3220cc2549960c38956d7

Observation 52d0685e-e369-426c-8407-dc37a16de43f · outbound

This paper cites CLEVR : A diagnostic dataset for compositional language and elementary visual reasoning.

Understanding Complexity in VideoQA via Visual Program Generation CLEVR : A diagnostic dataset for compositional language and elementary visual reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.787570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.682106Z digest=sha256:18fd92dcba4223b767ec7c25abd4f31233324a59eaa96ed4032c06ad49ad37b9

Observation a4a7cd14-7886-4e6f-bf97-dabeafb351df · outbound

This paper cites Inferring and executing programs for visual reasoning.

Understanding Complexity in VideoQA via Visual Program Generation Inferring and executing programs for visual reasoning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.777869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.780729Z digest=sha256:1c265f3593d4ebe13fe7544f9a4edd6dd72f22498b1900bc4ecc69b264b8f442

Observation b0c07739-aa8c-4d9d-8cd4-07464fbab446 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.766798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.784668Z digest=sha256:9502fd3911cfd0fd5bdd77b7d87dc5bef1af2450cdd528007b594768ca51b70e

Observation cf7058e5-aedf-4cb1-9b87-09d8d8c7e458 · outbound

This paper cites W., Tapaswi, M., and Fidler, S.

Understanding Complexity in VideoQA via Visual Program Generation W., Tapaswi, M., and Fidler, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.688625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.788366Z digest=sha256:f2d0cd5de45e6e1c4be78583f22938b718ee192926537fb03726a875aadb1da8

Observation c3beb552-4555-4324-a1cd-0a9299354cdb · outbound

This paper cites and Bojar, O.

Understanding Complexity in VideoQA via Visual Program Generation and Bojar, O

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.677424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.791691Z digest=sha256:fe46e44207b348d3ba76cccdb950e8423ba736f8237ba526e09731b3e8862f20

Observation 99390403-e02d-41ab-9bb4-aa4900e56ad2 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.664823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.796153Z digest=sha256:73a7e58557571bc6179aae782337d7bb32ec6d0d6b9970c8f9bac95a27622060

Observation 59170970-a60e-427d-b6ea-254d76a67255 · outbound

This paper cites and Kramer, O.

Understanding Complexity in VideoQA via Visual Program Generation and Kramer, O

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.586622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.801724Z digest=sha256:08b87866b1437251b2fd47e7483f204741456ce0e71b2dafa3d78ae2c7457e6b

Observation eb2bc160-0bf3-4a8e-94f7-99ef0991233a · outbound

This paper cites Dense-captioning events in videos.

Understanding Complexity in VideoQA via Visual Program Generation Dense-captioning events in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.505228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.806322Z digest=sha256:6e84be0cbc36aea1e7027eb397c0edb7e7762225f0d063997bf5e610b53afcab

Observation b345012c-50fb-41b0-bbc4-56890447f830 · outbound

This paper cites BLIP-2 : Bootstrapping language-image pre-training with frozen image encoders and large language models.

Understanding Complexity in VideoQA via Visual Program Generation BLIP-2 : Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.491312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.897852Z digest=sha256:852ba210588216f1f76c3699e94048fef82e2a701e6164675cbe69bb5f655cb1

Observation 6004c26b-7436-4087-bf83-ba831b93b451 · outbound

This paper cites MVBench : A comprehensive multi-modal video understanding benchmark.

Understanding Complexity in VideoQA via Visual Program Generation MVBench : A comprehensive multi-modal video understanding benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.440557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.973337Z digest=sha256:e8818df16513b6b1ed45269bc30c031a79475a498d2eb71d17bf0f5ca95cfa84

Observation 08bbdb7b-4e67-43fd-9767-878955cf28ac · outbound

This paper cites A technique for the measurement of attitudes.

Understanding Complexity in VideoQA via Visual Program Generation A technique for the measurement of attitudes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.976804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.976804Z digest=sha256:bce4390a4a224a95f2ab23004a79d93c042fe7616273f4133510eb61e7f80c66

Observation d852b290-6d99-4c5b-aa41-734f0a6094bc · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.359443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.980311Z digest=sha256:81f1b3659c8341cad94ee8d54c832ab838c62203921fb3b4de63cb023c2ab592

Observation 7e17724d-cf08-4d74-af08-42ae8fdd86df · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.984196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.984196Z digest=sha256:159c7a13b504c4661d7cbb606db10d6f80042c353437b481ca6f6d96ffbe5606

Observation 0ccd3f8a-0eff-47ba-b4fb-0de746ebedaf · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.341479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.988101Z digest=sha256:98801efcb88a984aa05fd6f399d8dde612e94c7be61ae5faaaa8e5d10724329e

Observation ad5bd453-9680-40d6-b2c8-8c243b968d21 · outbound

This paper cites L., Nejadasl, F.

Understanding Complexity in VideoQA via Visual Program Generation L., Nejadasl, F

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.328942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.992410Z digest=sha256:a8de71185cc90e4942872f30c1c9db2d2f4a3bcc4be764306bc3a4631786ec51

Observation 1a2eef59-ab74-4526-aee2-f79c3c8ecddd · outbound

This paper cites C., Adeli, E., and Li, F.-F.

Understanding Complexity in VideoQA via Visual Program Generation C., Adeli, E., and Li, F.-F

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.294349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.996538Z digest=sha256:41b2c80e6ba392711a56de9e81474f4744dbaf938b0c1f47611c68dc1310073f

Observation 4e3f0a8a-6d99-4add-8081-45b11def5ea2 · outbound

This paper cites Y., Wu, J., Niebles, J.

Understanding Complexity in VideoQA via Visual Program Generation Y., Wu, J., Niebles, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.251931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.154088Z digest=sha256:740cb8f64ee6cea41b5c74718a2b3add9c5927e560845db213d6d406bda5707b

Observation 2956f30d-8cff-437f-bb74-40007ba61fad · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Understanding Complexity in VideoQA via Visual Program Generation VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.268106Z digest=sha256:2854a5b153cf92a7eb99d92f2dfb1220bf1b92ba01cfeb508905b4777a34daf6

Observation 04dd940f-c62f-4b38-98a0-001d23921706 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Understanding Complexity in VideoQA via Visual Program Generation Self-refine: Iterative refinement with self-feedback

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.241956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.272442Z digest=sha256:7ec35b686b2be5689351a589717b77a2963297d15c20d456fadaaa59ff91cda5

Observation 3e63888b-157e-4a7b-8945-445c70b22ef1 · outbound

This paper cites EgoSchema : A diagnostic benchmark for very long-form video language understanding.

Understanding Complexity in VideoQA via Visual Program Generation EgoSchema : A diagnostic benchmark for very long-form video language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.229170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.276215Z digest=sha256:39eb003d5fb73c5fac5540b505458313b6c7c343c3ea97a093cdcf9d9b158fbe

Observation 157573ae-15cc-4564-a454-05c2a6117481 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.218861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.281674Z digest=sha256:042b34584d2486ea122005cbaef4ff9ec59aef2bf833d552986d7191049e8a31

Observation bd40ece6-92b4-4b2a-b863-56dd25c2bbb5 · outbound

This paper cites and Vondrick, C.

Understanding Complexity in VideoQA via Visual Program Generation and Vondrick, C

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.179229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.285183Z digest=sha256:3982e349b4ff1e50dc76fc6fb3bd8f5e4f96b11202d42c8214a3d6efe3027515

Observation 0f745b9c-c4a7-4e3a-936d-d8e0b923dc59 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.084828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.290047Z digest=sha256:b5fc0779b98fd170f90eeb485c1bec3c7564dc8550d38221978678037002051c

Observation 4abd0291-46ce-4fce-bd37-46f94e1b3f70 · outbound

This paper cites GPT-4 technical report, 2023 b.

Understanding Complexity in VideoQA via Visual Program Generation GPT-4 technical report, 2023 b

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.072753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.294290Z digest=sha256:5cb0cc4a381c6bc8a6a74263f281b5b37477179678aa75ada778c91e9cab8385

Observation c016c4a7-fb51-4f2e-83c6-d55dffdf232f · outbound

This paper cites and Schitter, C.

Understanding Complexity in VideoQA via Visual Program Generation and Schitter, C

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.035573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.355937Z digest=sha256:3bad9d8a5dbc4abc3ae06046f8ad5eb49edb9894ecbf77a66d093185640d3e32

Observation ca391c39-d642-4a39-be9c-1360b4519c2f · outbound

This paper cites A., Stretcu, O., Neubig, G., Pocz \'o s, B., and Mitchell, T.

Understanding Complexity in VideoQA via Visual Program Generation A., Stretcu, O., Neubig, G., Pocz \'o s, B., and Mitchell, T

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.935382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.439614Z digest=sha256:7b6de3f7c90d714fc604d8fee73447c8098c6b015920e292633683f1d95faea9

Observation 9bb79496-4021-4868-be98-b342926e1830 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Understanding Complexity in VideoQA via Visual Program Generation W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.505588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.505588Z digest=sha256:417ee49ab382b3f7f09277a244ccbcc2dd0780398de8e2af293208a5c4210df8

Observation 66c93eda-5306-4770-b782-9afe71995b4d · outbound

This paper cites D., Ermon, S., and Finn, C.

Understanding Complexity in VideoQA via Visual Program Generation D., Ermon, S., and Finn, C

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.510472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.510472Z digest=sha256:193aed9fff59bd2e083cd1e5ce5f183adb087d612c85fe22d2e69f7c895dc430

Observation 42d4c87c-6a6a-44b6-827e-3703364546a6 · outbound

This paper cites Annotating objects and relations in user-generated videos.

Understanding Complexity in VideoQA via Visual Program Generation Annotating objects and relations in user-generated videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.892020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.513979Z digest=sha256:39563f5e9d6d73f86c2fea5f96372cac4a710be7976666775d0da13e0d3db3f8

Observation 2d485fa0-c56d-454b-b223-ea5a89d6485d · outbound

This paper cites HuggingGPT : Solving AI tasks with ChatGPT and its friends in hugging face.

Understanding Complexity in VideoQA via Visual Program Generation HuggingGPT : Solving AI tasks with ChatGPT and its friends in hugging face

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.835259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.517327Z digest=sha256:6a39f887f4dbc100f479e9aa86e5eab7e72c8ffbb24e0b658046922df5164631

Observation fb75f1b2-6544-412a-8f2e-20e9e56f825f · outbound

This paper cites A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A.

Understanding Complexity in VideoQA via Visual Program Generation A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.821822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.522238Z digest=sha256:a82aedc706759c08e8d09f0620e181550ee069b42920c61bf6a6bd8af1042a6b

Observation bf749fed-8985-4223-b74e-56a9e968d7a8 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:54.729547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.526964Z digest=sha256:bd6d74d18513cdb201fbd9d337058a214f84ca74adbceff5562fa42a0359e5bf

Observation f403c41a-248b-4a1e-a39c-ce172b6e01e8 · outbound

This paper cites T., and Leordeanu, M.

Understanding Complexity in VideoQA via Visual Program Generation T., and Leordeanu, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.687545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.531242Z digest=sha256:3d95767d9407b1ce9a6af581860e8339f1fbefa0de2de1adc00f28c57828af08

Observation 86f92031-0c26-4206-bfcd-adab7d323854 · outbound

This paper cites less is more.

Understanding Complexity in VideoQA via Visual Program Generation less is more

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.616302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.535344Z digest=sha256:dde3c7e7dd84190cfeb0790e56fe77be0ab555c265bb5f0ed3bf98e224f3dc9a

Observation 6719b62a-8ec8-449d-ab37-3ca6793abaa9 · outbound

This paper cites Modular visual question answering via code generation.

Understanding Complexity in VideoQA via Visual Program Generation Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.548302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.616141Z digest=sha256:653ee8eb4d3ac6e8edd692158400fba6d9b5977c7eed0b6ce0c40e0925212ece

Observation ca8dff22-2529-4b59-8ff5-755e67f4c2ac · outbound

This paper cites ViperGPT : Visual inference via python execution for reasoning.

Understanding Complexity in VideoQA via Visual Program Generation ViperGPT : Visual inference via python execution for reasoning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.536876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.677195Z digest=sha256:f4d70a64fde12d7b6042d8ff176d1ce21ab01f72d3117a9f1b96710b080405e9

Observation fcb49ab8-56ad-40a6-a2ba-042cc377d8ee · outbound

This paper cites T., Fu, J., Phan, M.

Understanding Complexity in VideoQA via Visual Program Generation T., Fu, J., Phan, M

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.472736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.680634Z digest=sha256:3eb53a5eaa9eb32143720aad99fa22050811e80e4d8014ffbf6e7890bac94104

Observation 9b720306-7b7f-4986-94b3-e57a2bb6c320 · outbound

This paper cites A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J.

Understanding Complexity in VideoQA via Visual Program Generation A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.441325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.684219Z digest=sha256:e28a4c8c159a522a05664e216af15eebb6fa19bb03a208cc84883be25fd09ff8

Observation f9125a1f-35b7-4ff1-83ea-0b0b3f591949 · outbound

This paper cites Learning the curriculum with bayesian optimization for task-specific word representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Learning the curriculum with bayesian optimization for task-specific word representation learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.430869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.689888Z digest=sha256:c1eb8e8735da21df3dd5c30210a757c87f35907c43647e8259480cd1b1d4a795

Observation 16537687-d975-4425-be25-714d1e1780e8 · outbound

This paper cites Learning the curriculum with bayesian optimization for task-specific word representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Learning the curriculum with bayesian optimization for task-specific word representation learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.419868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.693265Z digest=sha256:2b198873b4cefdb012ac350a1b4166eaed3a9844dc97390a63cfe2bbd18975de

Observation d4511928-59ec-49e5-b35c-d7b93c0d0c51 · outbound

This paper cites P., and Ferrari, V.

Understanding Complexity in VideoQA via Visual Program Generation P., and Ferrari, V

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.395810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.697978Z digest=sha256:28555678941619675440dc3a198e1746dcb9ca53947591687f7498d2f295f819

Observation bb1fa786-6272-40f2-8ba2-23a4e27b8b15 · outbound

This paper cites Neural discrete representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Neural discrete representation learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.365097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.737006Z digest=sha256:e51b51ff5185319423c0c9cc4d14240dbfa9f588051a35a7ee67aadfaa594e54

Observation 2c1be96e-0b4a-4d08-ac31-3363eb4e34f0 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Understanding Complexity in VideoQA via Visual Program Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.797005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.797005Z digest=sha256:2bc08c62c4f322c2102ed4716679d9ddc1e8fb1fc9b2c12bfb2f40657ba6b409

Observation 9d6faa57-5ac2-4c06-9732-2142efba93d2 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Understanding Complexity in VideoQA via Visual Program Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.824535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.824535Z digest=sha256:bec63c7b7f68e5ccdbe55d8ea8a909e947fd9d8241e40530d0c8bb355e1a3711

Observation d3c51b07-1823-4f27-bf79-d91550cafb26 · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners.

Understanding Complexity in VideoQA via Visual Program Generation Language models with image descriptors are strong few-shot video-language learners

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.251351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.829436Z digest=sha256:55dcc244ce7397167fc193338e2936f6d0045adbc83d005bbfe2006219fa0e4c

Observation 024758c1-057c-4296-af9d-9cd960e4c43e · outbound

This paper cites STC : A simple to complex framework for weakly-supervised semantic segmentation.

Understanding Complexity in VideoQA via Visual Program Generation STC : A simple to complex framework for weakly-supervised semantic segmentation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.240302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.833774Z digest=sha256:dfe1848e5fb9077d8f2f9b1b4b56a845582d83f62bd5ead4be20851402e736ec

Observation 5a5aa68f-7ab6-46df-96c9-f8cad68acceb · outbound

This paper cites B., and Gan, C.

Understanding Complexity in VideoQA via Visual Program Generation B., and Gan, C

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.170867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.838175Z digest=sha256:8fc9d0724401ab311dea8cf633ebe597268727ec8c143d8415182e11166dd52f

Observation 6c4c1dfd-2fd1-44bd-bff1-c6f74c7d7b0d · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:54.091843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.842464Z digest=sha256:18168f1c9973fbce5b3d860d199a104d2a99dd74121e38161dfbb6b57bff8bd9

Observation e52093e4-2b9e-4fb0-ad5c-b5dd892ec27d · outbound

This paper cites Next-QA : Next phase of question-answering to explaining temporal actions.

Understanding Complexity in VideoQA via Visual Program Generation Next-QA : Next phase of question-answering to explaining temporal actions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.081918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.875308Z digest=sha256:8a07ead9a0f6466bbb0ab13ce8d836d556442f596e052b78e4ed9f3879260049

Observation cdf6bca2-8547-46d6-9c05-8ee636428641 · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

Understanding Complexity in VideoQA via Visual Program Generation mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.964977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.964977Z digest=sha256:e792a2be5d959ba47c85ad21190a0406fe3fa1c391ddc986f92b2569d4de8160

Observation 0303ce46-2eb6-4197-9eed-bb9a81059ffd · outbound

This paper cites W., Salakhutdinov, R., and Manning, C.

Understanding Complexity in VideoQA via Visual Program Generation W., Salakhutdinov, R., and Manning, C

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.002089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.969649Z digest=sha256:ff0bb1524e214ea426be7b7a70af16c248cf42cbc0ff41b4ab95f2d2288fb7ab

Observation 7a2ab9c0-d6f5-4a9e-83a0-62150c5bf623 · outbound

This paper cites Neural-symbolic VQA : Disentangling reasoning from vision and language understanding.

Understanding Complexity in VideoQA via Visual Program Generation Neural-symbolic VQA : Disentangling reasoning from vision and language understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.931880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.972655Z digest=sha256:fd754b30e6af6a0804097b2fcad2972189dad0418500d37563af03ad0745145a

Observation 40c08bc2-3a66-42bc-b321-6d9ddfb3a8a1 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Understanding Complexity in VideoQA via Visual Program Generation Self-chained image-language model for video localization and question answering

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.920673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.977650Z digest=sha256:bfa657363156e83be1054e781f5b02b1c8bc7eebb95f44d0d1f2da83ba41e544

Observation 1b879339-726b-4b7f-8ee6-5397d7ae8b1b · outbound

This paper cites ANetQA : A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos.

Understanding Complexity in VideoQA via Visual Program Generation ANetQA : A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.740461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.982250Z digest=sha256:c3652a5f0860adc6d62b5750d72cb1c8bcc739ffcb7829d81d7b2900d5c0f49a

Observation 662ab958-dd81-444f-abc3-bf481bfcf122 · outbound

This paper cites S., Cao, J., Farhadi, A., and Choi, Y.

Understanding Complexity in VideoQA via Visual Program Generation S., Cao, J., Farhadi, A., and Choi, Y

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.728692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.987034Z digest=sha256:f06abcd980dc37d97e11ca0862233aff29a339fb3525aeaa5003b5eee2ca9482

Observation 4902b6f6-d0f6-4c1f-8293-e79c098a773d · outbound

This paper cites Socratic models: Composing zero-shot multimodal reasoning with language.

Understanding Complexity in VideoQA via Visual Program Generation Socratic models: Composing zero-shot multimodal reasoning with language

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.662002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.039709Z digest=sha256:4a9883e81b7633e2edcabfbe8192a4b6f8d9f9f569e73501e96c43bef658f81b

Observation a80565bb-cabc-40ef-b32b-3474db711920 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Understanding Complexity in VideoQA via Visual Program Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:53.128703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:53.128703Z digest=sha256:923bcbe1dc6c4ac93ba48549fd96348c47fb804e2b20d1e891ee827e8d83b3c0

Observation 3f09dded-6a1b-4309-98b5-1a3fe53b8741 · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Understanding Complexity in VideoQA via Visual Program Generation Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.493343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.132142Z digest=sha256:be7c3a7f682623c120236fd138898fffaef9dc05522cf4b5f7ce036e74045425

Observation cd2bf090-c967-4181-93b4-2b902cba8d16 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges.

Understanding Complexity in VideoQA via Visual Program Generation Video question answering: Datasets, algorithms and challenges

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.481138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.135905Z digest=sha256:7a6df3ec5d1f3d147f57f28934a25c4209593e61df3ef9beb71da786783d943b

Observation ed530913-ea58-4888-b9e1-a821b2d8c0ef · outbound

This paper cites J., and Rohrbach, M.

Understanding Complexity in VideoQA via Visual Program Generation J., and Rohrbach, M

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.467634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.139787Z digest=sha256:af9e427f353e043d6c800182f43c5e2d5cac671fbfbdc9abdaa30aabc06e3f4b

Observation 5c21cffe-8e23-465d-901e-3412ac9b505c · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:53.276370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.143452Z digest=sha256:d799b753c94da3537862daebdc98b2d194c917a1bc75f8e16209182a93e2524b

Pith citing papers

Observation 8a371f5e-a9cb-4c5e-89fe-76a533985358 · inbound

An Attribute-Based Measure of Video Complexity cites this paper.

An Attribute-Based Measure of Video Complexity Understanding Complexity in VideoQA via Visual Program Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.107311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T19:00:54.718177Z digest=sha256:906ab528776c631d1081523f87bbf3d89448a202a59ff0d9aeb27596fb3fdcc0

Observation e4a04b15-1e58-4a35-83c1-cd3e39cf05e4 · inbound

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding cites this paper.

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding Understanding Complexity in VideoQA via Visual Program Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:22.239838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:32:22.239838Z digest=sha256:ef7d0e4b8c37e0c96934fe3c988f48db015ffbe78b65b2ca82fc3aaa5e6e4f65