Pith. sign in

Paper Citation Record · LEDGER

Admitting Ignorance Helps the Video Question Answering Models to Answer

As of 18 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2501.08771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08771 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:22:29.090022Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy55
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9532b078-bd9d-472e-8895-c0dcacd07458 · outbound

This paper cites Bilinear attention networks,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilinear attention networks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.083794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.815676Z digest=sha256:824a7a86bfdd9f13a5ee0680035b9dbbc5bfd4590cd51b8e87c0c44ac3332239

Observation 78a8b23e-b5f6-4ecc-8f34-c9b7493bd253 · outbound

This paper cites Attend what you need: Motion-appearance synergistic networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Attend what you need: Motion-appearance synergistic networks for video question answering,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.072790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.820087Z digest=sha256:3cd872f3003f44fafbc6c3a4d8ee2b0a5f5e1011edf42db2cb2da506344681cd

Observation 58cab2bd-070b-4bfe-807c-6733d6ab7df8 · outbound

This paper cites Video as conditional graph hierarchy for multi-granular question answering.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video as conditional graph hierarchy for multi-granular question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.061788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.824049Z digest=sha256:748fdc6d828fc7f187acf23f92834952efbc801179047421ffc343a385b99447

Observation 852a8b05-c6fc-4fb9-8212-9d3c23425d1b · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Merlot: Multimodal neural script knowledge models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.052125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.828251Z digest=sha256:f76002d7f3713c5e8a70c3de861e03f82eb8837eb24bdc9a31c2cfb9e45cc45c

Observation 21caaec5-68af-4a53-bed9-af6c1a4b862a · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Admitting Ignorance Helps the Video Question Answering Models to Answer VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.832209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.832209Z digest=sha256:911d85ba0a6dee4689b5c04e841b61660cd372c8f3d744f6bff2540532111757

Observation 803b90c5-5350-4593-ad3e-b099f4f80ad1 · outbound

This paper cites X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks.

Admitting Ignorance Helps the Video Question Answering Models to Answer X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.836606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.836606Z digest=sha256:8e859367a1d90106136318535e9df59efb14a13c6645daf8c5f6639b073b9327

Observation a1bf4bc0-51da-4e06-92df-b148982abc98 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Admitting Ignorance Helps the Video Question Answering Models to Answer InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.840944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.840944Z digest=sha256:f8baff1b6742c9648a411c0701c0f6e21e4c04bf536a5d4098ae2631d31bed4c

Observation 9b75616c-ba59-4a81-b431-8e48192bd12a · outbound

This paper cites All in one: Exploring unified video-language pre-training,.

Admitting Ignorance Helps the Video Question Answering Models to Answer All in one: Exploring unified video-language pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.041510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.845487Z digest=sha256:4fb8ca5c79b3df2de6170378f1e1c4b821fd0894fbde77b636fbaa93e3960460

Observation 1eb77c64-db60-464f-bb67-2227cdef559f · outbound

This paper cites Invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Invariant grounding for video question answering,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.029231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.849971Z digest=sha256:03a7f18c3814f38e1f4c02eb6099f7f083890fd4eafd3b2d4744f5c83bf28675

Observation bd64f643-ca09-41ba-a6c7-a19c30af739e · outbound

This paper cites Equivariant and invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Equivariant and invariant grounding for video question answering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.016176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.853927Z digest=sha256:7968c7cfeac9efda56fe8c2c574acd3a4dfec4ca5a68501ee43b7c19fa8e9484

Observation 31bda1b2-d1e2-41f1-8610-04c2a0ef86fb · outbound

This paper cites Transformer-empowered invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Transformer-empowered invariant grounding for video question answering,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.000451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.858221Z digest=sha256:8fdf8e9e5b2a27bb8588cd12fd34645ee5a0fa866093233e5530494a4e821893

Observation dcbdd53f-f305-4ab8-926c-31a3dd24f80a · outbound

This paper cites Discovering spatio- temporal rationales for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Discovering spatio- temporal rationales for video question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.988421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.862463Z digest=sha256:8bee707bf55387e355b435d91aef13c336a1edab1b5413d559d6bdb39d8474bf

Observation 4819a8d9-4faa-4d3d-9909-24a637d781d7 · outbound

This paper cites Adversarial vqa: A new benchmark for evaluating the robustness of vqa models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Adversarial vqa: A new benchmark for evaluating the robustness of vqa models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.976532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.866406Z digest=sha256:e244b56482ed16c5c1bc94222c45d89311b239c2e4f77a51930f7cfc4654f245

Observation 51a192e9-c3b7-48e5-acff-21624f1775ce · outbound

This paper cites Discovering the real association: Multimodal causal reasoning in video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Discovering the real association: Multimodal causal reasoning in video question answering,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.964326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.869197Z digest=sha256:dc935b01463d36e56db0d14c6503f49f99be0d27964e71b3cae7466a1e00c715

Observation a49e0de7-7419-402d-a33b-e09219841169 · outbound

This paper cites Coun- terfactual vqa: A cause-effect look at language bias,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Coun- terfactual vqa: A cause-effect look at language bias,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.951666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.872228Z digest=sha256:0ed293552d54c79dbf2aa4883dcaceea6d2d6b387c72e414585b4dfc74abf0ea

Observation e8864a2e-5808-40c6-b518-77ab0c08d48f · outbound

This paper cites Beyond question- based biases: Assessing multimodal shortcut learning in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Beyond question- based biases: Assessing multimodal shortcut learning in visual question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.938674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.875299Z digest=sha256:978d433e45a3329ad9a08623a05e076c2e34fb6366094a283add570caa251527

Observation d575a6f5-4b9f-4d92-b8cc-39f2a7f3baa7 · outbound

This paper cites Roses are red, violets are blue... but should vqa expect them to?.

Admitting Ignorance Helps the Video Question Answering Models to Answer Roses are red, violets are blue... but should vqa expect them to?

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.926218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.878700Z digest=sha256:36b1d60fea6a0da0127c2579b7f61b70f4e2bd7aa27aaff885615bfd9d96be7e

Observation b2019ef5-2c84-4f92-a84e-c247a7e3a4d3 · outbound

This paper cites Human-adversarial visual question answer- ing,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Human-adversarial visual question answer- ing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.912538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.882176Z digest=sha256:e931c4995dea48b69132bebca4126a4bf2c8d2f7e0adba2a48d1c21ac832e970

Observation 4e026506-d2ad-4521-ae3e-6cad9a0e5ec7 · outbound

This paper cites A survey on curriculum learning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer A survey on curriculum learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.898588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.886366Z digest=sha256:d02fd2853df3c89bcf7f517990da3cfe29f426c92433d76753296881f0fb93c4

Observation 3e8ded24-e28a-4311-8830-2aeb40d69a6b · outbound

This paper cites Curriculum learning: A survey,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Curriculum learning: A survey,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.889818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.889818Z digest=sha256:7c6a2de16b13acfaa2dda36c36dbb8a347f751b802ad9acb8d0785c548b85049

Observation d0b6f7db-8ade-4e59-ac02-3bccd89b6f9b · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.878483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.893376Z digest=sha256:13d1e642647c51179f21ddb8ff5ad11a8a5a1850afc21bbc60571c0824a06c39

Observation bb3893a4-465a-41e6-85a1-5026b5dd85fa · outbound

This paper cites Revisiting the.

Admitting Ignorance Helps the Video Question Answering Models to Answer Revisiting the

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.867380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.897187Z digest=sha256:6a980017a418e4cf1b18a862903dfe037044a93bf71d2a54c8b9143fd0fc339f

Observation b2da382d-c2bd-44d2-9171-27d8ee531a5f · outbound

This paper cites Answering from sure to uncertain: Uncertainty-aware curriculum learning for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Answering from sure to uncertain: Uncertainty-aware curriculum learning for video question answering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.901073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.901073Z digest=sha256:57b48cb6a8d463603da574b8316d295eef869c91bd39d65dbee45ac552ca2e0d

Observation e2a152b2-a62d-42ff-86cd-e1328346f135 · outbound

This paper cites Question-guided erasing-based spatiotemporal attention learning for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Question-guided erasing-based spatiotemporal attention learning for video question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.857861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.904741Z digest=sha256:91131167816b3edbee20fed92adb0740168ebdf9da0725d3dd419b569ebd250a

Observation 06093558-5918-49fd-9669-f538d0ec5d5f · outbound

This paper cites Memory augmented deep recurrent neural network for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Memory augmented deep recurrent neural network for video question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.846939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.909220Z digest=sha256:74ea8f2bff53decc18c74e7582c67fa3bbfc7f67543d9c369ed99d1937f5327e

Observation 431db470-deaf-48f4-a4fe-75b4f3d4abc5 · outbound

This paper cites Knowledge-routed visual question reasoning: Challenges for deep representation embedding,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Knowledge-routed visual question reasoning: Challenges for deep representation embedding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.835410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.913520Z digest=sha256:01d042b901587c6de37821a6dddb1f00d420a59988e871312377dc2c650f3a3f

Observation b3be0ff1-18e9-4203-8bcf-5a412b7c57d0 · outbound

This paper cites Multitask learning for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Multitask learning for visual question answering,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.823215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.917336Z digest=sha256:015a785b7491be37af332d82b2ed0161b106fae8f3492c0f8f16c7a9db6243a1

Observation b830ec11-203f-4216-b2f6-8cabae86c334 · outbound

This paper cites Bilinear graph networks for visual ques- tion answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilinear graph networks for visual ques- tion answering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.811946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.922428Z digest=sha256:2ba538bdab6422fce63889bab23a08f1f134317f61917609c56960492cce4336

Observation a36e2dd3-5034-466a-8c1f-e7fd57e0df58 · outbound

This paper cites Bilateral cross-modality graph matching attention for feature fusion in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilateral cross-modality graph matching attention for feature fusion in visual question answering,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.800548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.925840Z digest=sha256:99b961517a09ef0abe8cc2ec465b787ff8de09c098e13a670c83d85278801ba1

Observation 45178bf5-0eb8-414d-9906-9416d950a396 · outbound

This paper cites Bridging the cross- modality semantic gap in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bridging the cross- modality semantic gap in visual question answering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.787587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.929126Z digest=sha256:a05dffb977978567798ff94a8ec6cd6b24a2455f845e0591d86506da1b890d21

Observation cc25a2d1-e088-44e7-9b65-4fd7ae5029a7 · outbound

This paper cites Latent attention network with position perception for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Latent attention network with position perception for visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.776048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.933066Z digest=sha256:7171f4c5d3399bdd8a50da1c052872f4b65ad710246786bdeb317520403aeec2

Observation aac3d65f-0e70-4238-9cc7-daabc6288571 · outbound

This paper cites Webly supervised knowledge-embedded model for visual reasoning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Webly supervised knowledge-embedded model for visual reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.764674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.936324Z digest=sha256:db6938d6a8a742f22781b9053926364b97e8ce768f1002e9433ed6665f36728a

Observation c7110ae4-0f09-434c-9d73-442c8d76815b · outbound

This paper cites Uncovering the temporal context for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Uncovering the temporal context for video question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.753036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.939590Z digest=sha256:adfcf8de0392bc2722ad614a21aa3e2feff8f42bbb44eb1329260d32e02be0c9

Observation 8d3ecdae-de28-48e3-8d52-27ad057b8c44 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video question answering via gradually refined attention over appearance and motion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.740341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.943213Z digest=sha256:bdb87bf4fb0a44a417db20ffd0a21004fee60b646e40d96d2bb9becc0f025272

Observation 863e8001-2d72-42b1-ad53-1526236563e3 · outbound

This paper cites Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.728219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.946761Z digest=sha256:f823779dedd9c75c279a75e1bd93ce918802a13a6a199c7bfbbfd0a2094f8f26

Observation 8dc85746-341a-442d-9ad1-b34a4b8fd5e1 · outbound

This paper cites Video question answering with spatio-temporal reasoning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video question answering with spatio-temporal reasoning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.716195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.950509Z digest=sha256:7d0b36457e0e96d10d8cd66bc6372a0ca603a982cc102360b701ddb1e82ec3af

Observation 69e3270d-7c37-4141-9d1a-424baa1332d6 · outbound

This paper cites Dualvgr: A dual-visual graph reasoning unit for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Dualvgr: A dual-visual graph reasoning unit for video question answering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.704012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.953955Z digest=sha256:fe2855b91cb099619c83665a4ac75cfca7811ce24251d518bfe55377108dc436

Observation 3cd45be9-6a74-4205-af87-22150a4fa54e · outbound

This paper cites Bridge to answer: Structure-aware graph interaction network for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bridge to answer: Structure-aware graph interaction network for video question answering,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.691371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.957975Z digest=sha256:3b5525140b58093f32421b2012382c6644f38a97638ab52bcba41abb9dbd3967

Observation bd0de087-f9af-47aa-a621-98dddf70774b · outbound

This paper cites Motion-appearance co-memory networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Motion-appearance co-memory networks for video question answering,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.679267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.961511Z digest=sha256:7ada9ad82ecfc5e89cb8b31f8c9643d476d5525e2bc78b2d4a1ecbab63d6e03c

Observation c1149717-3af7-4511-80d4-ea9312ec6448 · outbound

This paper cites Heterogeneous memory enhanced multimodal attention model for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Heterogeneous memory enhanced multimodal attention model for video question answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.668248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.964910Z digest=sha256:a2da32e680893df0ca26be0fa2f3b9730922eecd953c13f47657c25bf5e7765f

Observation 0fc7a196-c301-474e-a175-c42904a5edbf · outbound

This paper cites Hierarchical conditional relation networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Hierarchical conditional relation networks for video question answering,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.656871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.968357Z digest=sha256:4fb90e5c2da13403fbac521609a3ee7c55315c737105d8ee68394cf3cd1cd638

Observation e3cfb448-3188-4f39-a469-c80dc57ae182 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Verbs in action: Improving verb understanding in video-language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.645191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.971679Z digest=sha256:1a013920ed0d5b8309be5148bbfdc7fd4d2349204da4cbdc3161f3fe62d2aec4

Observation 42900346-0920-434b-ad7a-25087c9f9ff9 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Admitting Ignorance Helps the Video Question Answering Models to Answer When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.633656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.974963Z digest=sha256:8297498f0b076bb1d9c9509b63148b398f7ed4a10fcf4aa02f998484fb44eedc

Observation aebf50f1-0815-48f5-981b-127bb796945f · outbound

This paper cites Teaching structured vision & language concepts to vision & language models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Teaching structured vision & language concepts to vision & language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.978593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.978593Z digest=sha256:d6d447950159aac1cdead9d2165857934eb5deb98432a9642fcd6d4e3e6353c3

Observation 3a7f542e-94e2-4463-b39b-ce6ffd6f613d · outbound

This paper cites Rubi: Reducing unimodal biases for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Rubi: Reducing unimodal biases for visual question answering,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.614069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.981805Z digest=sha256:e0e4e7da36399cf351a50be90c9cfb1bc7a5f15e35f9498894b0e8bf19fe0907

Observation 0390a37c-da9f-41b4-b8a0-51f726c37a2b · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Don’t just assume; look and answer: Overcoming priors for visual question answering,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.600848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.985272Z digest=sha256:0a631633dd7bdaf88916ae2d03ef0024893e2feb4faa4214ff25e8d9bcb3c814

Observation 0ea48d93-d1da-411f-9aa2-eb29c500de50 · outbound

This paper cites Overcoming language priors in visual question answering with adversarial regularization,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Overcoming language priors in visual question answering with adversarial regularization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.588740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.988924Z digest=sha256:470046676aa202b0f3540446ec47045f51b4121e7ed6db838bc007f3135166ba

Observation 2297cfd9-6abf-4134-9808-5ac885c7dc23 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Reliable visual question answering: Abstain rather than answer incorrectly,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.577485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:28.992385Z digest=sha256:b73e9981d22ea990d1411c3394bda50263ba84be3231c0dcac44f267096ef0f7

Observation 889e1f67-fff4-4e59-aada-db3fe4670164 · outbound

This paper cites Addressing failure prediction by learning model confidence,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Addressing failure prediction by learning model confidence,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.996874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.996874Z digest=sha256:748abe15e9484667c2c5a8f42873930b38ed658c0adf1a06e56b20c7a88953e8

Observation 4b57a7cf-c70b-4472-909e-31cc9f10b7f6 · outbound

This paper cites Combating Label Noise in Deep Learning Using Abstention.

Admitting Ignorance Helps the Video Question Answering Models to Answer Combating Label Noise in Deep Learning Using Abstention

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:22:29.182883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.001019Z digest=sha256:642c7769202d3ac5f4dd5444ad65490174466240cf6c535b756062a97e9c2f41

Observation ddf71bdb-f7d6-42df-ae1f-a2794b0077d5 · outbound

This paper cites The art of abstention: Selective prediction and error regularization for natural language processing,.

Admitting Ignorance Helps the Video Question Answering Models to Answer The art of abstention: Selective prediction and error regularization for natural language processing,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.559720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.004953Z digest=sha256:f9e35392d27a99a6514edfce1a62fd55d4227b13e095e63328de3fc593d80f9e

Observation 7a613141-c7e7-461b-b55f-889d64a619b3 · outbound

This paper cites On the foundations of noise-free selective classifi- cation.

Admitting Ignorance Helps the Video Question Answering Models to Answer On the foundations of noise-free selective classifi- cation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.547982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.008397Z digest=sha256:3cba8e6ec7bb03f53d063e81c9af3a50d2d7e866898877b2ee20423f25972487

Observation 1379861a-8ac1-4286-ba1b-c9a08c13a3e8 · outbound

This paper cites Investigating Selective Prediction Approaches Across Several Tasks in IID, OOD, and Adversarial Settings.

Admitting Ignorance Helps the Video Question Answering Models to Answer Investigating Selective Prediction Approaches Across Several Tasks in IID, OOD, and Adversarial Settings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.011509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.011509Z digest=sha256:7ca521ae5a43d153bc7f2230647949ab2614f687fd35be279f56240ab4e8c126

Observation 3341b454-363a-46cb-ab49-62b3041d7861 · outbound

This paper cites Selectivenet: A deep neural network with an integrated reject option,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Selectivenet: A deep neural network with an integrated reject option,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.535581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.015510Z digest=sha256:d3c1661013c9e7b7ca7112e726117b222f888771a519917287546a5264f49509

Observation e151b0e8-61a6-4d4a-87b5-20442f4a437d · outbound

This paper cites Selective question answering under domain shift,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Selective question answering under domain shift,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.523427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.018882Z digest=sha256:8076d8eda353691af479e7f734fbf0ba543134a48a65b17e4cf5810f70dadb18

Observation 1325e5ac-d15d-4ada-9316-a8365ef8b309 · outbound

This paper cites Attention is all you need,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Attention is all you need,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.022646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.022646Z digest=sha256:dcd032f71b54eefbbf600b5af5e0f51e4fdf121b543ea9bd959bf02eaca3ee2d

Observation 0f449ef2-86c9-4a81-9fda-b7d70ecf11fb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Learning transferable visual models from natural language supervision,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.026236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.026236Z digest=sha256:421b97ed00a9bec95ffb0c44dabe25fffb9734f7c69ee385a7a3bb8f0f360902

Observation 2f469043-f51d-4d8e-8e30-0a21a2fb13e7 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.496535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.029742Z digest=sha256:d266b96dbf7af06716274e873d01d552f0bec978a3a9a6f24e35c6e52b529397

Observation 99d0fe4a-e38b-4ac0-8e79-f35b8e9f8117 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.484728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.033198Z digest=sha256:3d87c9f79f6955acbfbff41f6a83c2a7c809a29c9811f0fa2c7c33d9233c1b1d

Observation 094aa94b-d13c-4553-bd4a-22f3d69d4f8f · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Ava: A video dataset of spatio-temporally localized atomic visual actions,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.473291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.036935Z digest=sha256:f86f91ef9f22a39d35d1e02b0aef8e850ca70ca27a566fd26a974b215e09c6a3

Observation 999faf4b-2dce-45c4-9888-1b0b1c780e61 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.461656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.040850Z digest=sha256:ad1ef48e404a352dfb40ccff97c91e465ed42969228575fe45e3fbec47cc238c

Observation edf1e6c2-5324-4e8d-a7e5-b820685d25e4 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.450128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.044658Z digest=sha256:5f606acc1529bebb83841cf2532f56dbf2ff3e4165241b738838db3ce120fa29

Observation 21650f3e-d479-4a7f-9edc-e1f891b68af4 · outbound

This paper cites Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering.

Admitting Ignorance Helps the Video Question Answering Models to Answer Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.048239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.048239Z digest=sha256:d19a9d881bb0b25cae74854fdae58f4c1c113fefc2a83dbd33ea1259f5f729f9

Observation ad5ab970-1c3d-4c30-b94a-cfb9e4e2cd4f · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.438477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.052228Z digest=sha256:65d63bff1ccc0302e4eb3bcc24efcd576f9246b19c5561973452946b967ec2b7

Observation adbb3e57-2533-42ff-9de2-c5893e946466 · outbound

This paper cites Causal inference in natural language processing: Estimation, prediction, interpretation and beyond,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Causal inference in natural language processing: Estimation, prediction, interpretation and beyond,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.055862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.055862Z digest=sha256:1bca264eef55119cf0ae7c5d30e2dfe8f321737a6179d29a51ccbc03a928ba25

Observation 758ebd2b-997d-491a-9a43-aa119f008327 · outbound

This paper cites Docogen: Domain counterfactual generation for low resource domain adaptation,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Docogen: Domain counterfactual generation for low resource domain adaptation,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.419757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.059587Z digest=sha256:c0879519999854f83eb539352168522edf35c5922e5943b926b3a7670608da3f

Observation 4967e378-4854-41ca-bcab-a0c2b723200c · outbound

This paper cites Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models.

Admitting Ignorance Helps the Video Question Answering Models to Answer Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.063214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.063214Z digest=sha256:54ced2f932f11d0c4c924d353d94f2ebeab894cf9c7806a72d392f958b1e894c

Observation 3c9abe9d-38ec-438e-a2fc-2ae6b0974985 · outbound

This paper cites Language models are unsupervised multitask learners,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Language models are unsupervised multitask learners,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.067137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.067137Z digest=sha256:865f46a5ebba4ce98b1a59a371932d25df31b9f2979aac5b040924f0dcfda5ee

Observation e7017573-60f0-4102-aeb3-ff0fc7b0940b · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Admitting Ignorance Helps the Video Question Answering Models to Answer Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.070599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.070599Z digest=sha256:09d6f9371ac6469f11dd4b4ebd4e9b2045da05bc2889af3025b257716099f859

Observation 3c5a5043-d24b-441b-9a3e-7133004cc6ac · outbound

This paper cites Stacked attention networks for image question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Stacked attention networks for image question answering,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.074601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.074601Z digest=sha256:3c0a69202277c22bc2950c803ff2b3b95598b5503913a5149e451dc66e379f09

Observation f693dfea-2e83-4e8b-9452-b39d8a32a65d · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Multimodal compact bilinear pooling for visual question answering and visual grounding,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.392912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.078064Z digest=sha256:684bec8fd31bc99843a38833409e9a031bb8b7645d7994d14c0aa5175c3e1e86

Observation d2d1af64-df40-4420-9a50-aa37acd053a1 · outbound

This paper cites Vqa: Visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Vqa: Visual question answering,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.081042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.081042Z digest=sha256:f07424403c86dfb786715acd66cacac504ba491945ceab91bb5c52ccfbd491d2

Observation eceb8198-17c8-4dd0-993a-e8d068f6e578 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.083891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.083891Z digest=sha256:5603978e29aac4cdf7f8cc2ca36db16e3705edb7289bc8cfdf37032869c85b09

Observation 2abc7e2f-be50-408b-a452-3a954810dc27 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.368636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:22:29.087051Z digest=sha256:c053c91b3d2056da4614cb586f9fe33de2bdb7204cfc9454a6907e9d8af51da6

Observation 2938132f-fb81-4bf3-b4bd-7aa0e143015e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Admitting Ignorance Helps the Video Question Answering Models to Answer VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.090022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.090022Z digest=sha256:10da6fc66d50b8733ac0893d69e8deb1c35c2b8b790e1bb2d36b268c8587ad91

Pith citing papers

No inbound Pith citation observations are available.