Pith. sign in

Paper Citation Record · LEDGER

Moment Sampling in Video LLMs for Long-Form Video QA

As of 15 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 4 inbound Pith citation observations for arXiv:2507.00033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00033 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:49:14.041691Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:25:25.260193Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8a7cd85e-a487-4551-9a63-9dc64925b8fa · outbound

This paper cites GPT-4 Technical Report.

Moment Sampling in Video LLMs for Long-Form Video QA GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.781573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.781573Z digest=sha256:80ae25d402e4df378e5910ebf8d101862bd7e7e04ad9d7df747a1e2619d6bb1d

Observation 23936293-6c31-4d1f-88fe-5720b91b9cee · outbound

This paper cites Combining global and local attention with positional encoding for video summarization.

Moment Sampling in Video LLMs for Long-Form Video QA Combining global and local attention with positional encoding for video summarization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.890698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.787217Z digest=sha256:e041760676004f4d26add2311f19e963e7fb7f88f59770ec6658389f04603049

Observation e4a9f1be-fcea-436b-a530-01cb420116af · outbound

This paper cites Vivit: A video vision transformer.

Moment Sampling in Video LLMs for Long-Form Video QA Vivit: A video vision transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.791273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.791273Z digest=sha256:5f2520598be7fd6a2c27e1f00389eb1fd5924a11545c1b3df96019db192dce1a

Observation 5dd2bd53-860c-46aa-a80d-23a5f3ce19ec · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Moment Sampling in Video LLMs for Long-Form Video QA OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.795458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.795458Z digest=sha256:316338985c7f3b13a2ab8692d2465afed171abdf1f9e6d0a09909b716d5776e1

Observation 34a87b98-988e-412e-83f0-2636b3696bef · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Moment Sampling in Video LLMs for Long-Form Video QA Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.799624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.799624Z digest=sha256:3d2c31f21aa66c23b916b351233eab8aac03fcfcd2402bcee363ec8651eb318b

Observation a6b7bf85-dd4f-4434-a698-7f8861a27c1e · outbound

This paper cites Memory consolidation enables long-context video understanding.

Moment Sampling in Video LLMs for Long-Form Video QA Memory consolidation enables long-context video understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.868051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.803766Z digest=sha256:b8a7b680502fea42bb178191d4c53e862bcd8640ff6ad8632a3b862ef5444051

Observation d921aeaa-0d76-477c-a8b7-e45be23c970b · outbound

This paper cites End-to- end object detection with transformers.

Moment Sampling in Video LLMs for Long-Form Video QA End-to- end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.807674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.807674Z digest=sha256:19b8570462c47bf821740a2a63a2aef82ee47b53dfd1e90bfffcc3b8db66768a

Observation a324b76c-e902-46ba-b997-561e3f0b9365 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

Moment Sampling in Video LLMs for Long-Form Video QA Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.844881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.811423Z digest=sha256:1536a7e2d881dacd839f2cbdc91e2153aff8bb15e2867e587d071aed59a82b60

Observation 214f83fd-d1d8-4a77-a32c-e6ad694f31b4 · outbound

This paper cites Cosa: Concatenated sample pretrained vision-language foundation model.

Moment Sampling in Video LLMs for Long-Form Video QA Cosa: Concatenated sample pretrained vision-language foundation model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.831045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.816004Z digest=sha256:a2bbf59c08b54e41d8d0c4d314664c268987e595f4b625c53e4d292ac00116d1

Observation 9dc9d8a0-7fbc-4a34-b545-83c8775bd0a5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Moment Sampling in Video LLMs for Long-Form Video QA Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.817307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.820959Z digest=sha256:7a41c2b456b946031d1b4e03419db7238638151e16e8c08997893aba69a99374

Observation 2005cc7a-ba8c-4396-b0f2-d634013dcfe5 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Moment Sampling in Video LLMs for Long-Form Video QA VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.826315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.826315Z digest=sha256:ca4c0e764965d62dc87255d9741a6c6e93fa2758b67c218cfd578f091bc13880

Observation 6ab27103-2808-4dd0-b1b6-54589ecd8f80 · outbound

This paper cites The Llama 3 Herd of Models.

Moment Sampling in Video LLMs for Long-Form Video QA The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.830989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.830989Z digest=sha256:736d3011c828d8c02fb824ba36104be31ac90a6259522d97380b528c869697de

Observation 9654396f-d8ef-43e0-b316-33d23be3021b · outbound

This paper cites Coot: Cooperative hierarchical trans- former for video-text representation learning.

Moment Sampling in Video LLMs for Long-Form Video QA Coot: Cooperative hierarchical trans- former for video-text representation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.801394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.835657Z digest=sha256:2f5fc6706b7d29daf266c7ccbcb21439d0fa4623c17d2e128f5372fa0e8a1293

Observation 6eef0292-7efd-43ef-a241-e8eb2a9395e0 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Moment Sampling in Video LLMs for Long-Form Video QA Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.840061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.840061Z digest=sha256:e36f5afa8dd98fb4b8a94afd7ff1650fe34ecec3788a34bbaba3851128cbb172

Observation 466fa2bd-4ed7-4c96-9546-c53016ae58ef · outbound

This paper cites Efficiently mod- eling long sequences with structured state spaces.

Moment Sampling in Video LLMs for Long-Form Video QA Efficiently mod- eling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.787043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.844372Z digest=sha256:c999155190cd7fd55a1697ed5f56bbb9689b79b6a61767e35d8c04eeb086f213

Observation a9309774-7048-4215-bc1e-30cd7ef99747 · outbound

This paper cites Pidro: Parallel isomeric attention with dynamic routing for text-video retrieval.

Moment Sampling in Video LLMs for Long-Form Video QA Pidro: Parallel isomeric attention with dynamic routing for text-video retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.773861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.848550Z digest=sha256:f4721e9b9345d57d99d158d060bfe00b33009493df010f9bf9616fd0a20562e5

Observation 13ada487-7703-4192-8db8-70963d61915a · outbound

This paper cites Creating summaries from user videos.

Moment Sampling in Video LLMs for Long-Form Video QA Creating summaries from user videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.760027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.852846Z digest=sha256:217d79eb4e5a80e40ded9f8bb4647966c9e497d2491658c505e32928336132a2

Observation ad0e74a5-14fa-4a27-8072-591d0094f9c0 · outbound

This paper cites Object-region video transformers.

Moment Sampling in Video LLMs for Long-Form Video QA Object-region video transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.745466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.857191Z digest=sha256:8809d2dd44e4e313c30c9b827f32ea513cb6cd40b6255e3ae07927074af8febc

Observation 33b56590-b2bf-4394-802e-9494323429b4 · outbound

This paper cites Video re- cap: Recursive captioning of hour-long videos.

Moment Sampling in Video LLMs for Long-Form Video QA Video re- cap: Recursive captioning of hour-long videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.731146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.861750Z digest=sha256:6ea918b04a85dd205099e1c1bca6438ff591d13190d06d69de1473567ffdf94d

Observation 1dac3981-d347-429a-97dc-379efe6db7dc · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Moment Sampling in Video LLMs for Long-Form Video QA Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.865630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.865630Z digest=sha256:ee1edc20b56d01a1c0d349118a994e807be7ccc55b57cc51274db716fdbe64ba

Observation ea832f9d-264c-44d1-94ab-21da59aac086 · outbound

This paper cites Mistral 7B.

Moment Sampling in Video LLMs for Long-Form Video QA Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.869339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.869339Z digest=sha256:95f9eb3995a08dd8550d01a6a72f33c26aef3a5110cd1788f9fb9e40d086a385

Observation efeea2b1-fdcd-4b61-b500-5f39926634eb · outbound

This paper cites Language repository for long video understanding.

Moment Sampling in Video LLMs for Long-Form Video QA Language repository for long video understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.707433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.873969Z digest=sha256:768a6a7d75976211de256d6b14a0b90dd6b15f286dfdbe3f75bb19e7afa29974

Observation 1e718c09-8f90-43f4-8640-eb935888ee78 · outbound

This paper cites Large language models are tempo- ral and causal reasoners for video question answering.

Moment Sampling in Video LLMs for Long-Form Video QA Large language models are tempo- ral and causal reasoners for video question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.694405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.877726Z digest=sha256:3a5bc826f76b5abfecb3f4dc624f53daeb7e00238963369888e35eed19186b63

Observation 1316294e-d040-4614-b0dc-2b48dc995cb7 · outbound

This paper cites Movinets: Mobile video networks for efficient video recog- nition.

Moment Sampling in Video LLMs for Long-Form Video QA Movinets: Mobile video networks for efficient video recog- nition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.881801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.881801Z digest=sha256:8f50500902e10cf1eebeb1e0ce4c9bb08e922c6b2308469e7fbc4c00de4abd7e

Observation ecdaeb2f-eeef-4d31-9c75-e94f3a9f34da · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

Moment Sampling in Video LLMs for Long-Form Video QA Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.669903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.885943Z digest=sha256:b361c5a7f436f8acef5cb201710d534a104172b1632f89c441e3875271b73b12

Observation fbc22641-7272-404d-bf21-54a72e19dfea · outbound

This paper cites Dai, Zhifeng Chen, Claire Cui, and Anelia An- gelova.

Moment Sampling in Video LLMs for Long-Form Video QA Dai, Zhifeng Chen, Claire Cui, and Anelia An- gelova

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.655587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.889723Z digest=sha256:2a200bce5c1688f9e38665ffd0dc967ce1dcb874e6659097435bf9e010b454fe

Observation 273140d3-8ba9-4586-99c1-e015116524c0 · outbound

This paper cites Cast: cross- attention in space and time for video action recognition.

Moment Sampling in Video LLMs for Long-Form Video QA Cast: cross- attention in space and time for video action recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.642129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.893877Z digest=sha256:46ff2247860b598cea7ff64c397957aed65aa4a1a1ff841ca43a317731188728

Observation bc5caf1c-a410-43dc-8e98-e93adb3ef1c0 · outbound

This paper cites BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos.

Moment Sampling in Video LLMs for Long-Form Video QA BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:49:14.245532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.898262Z digest=sha256:61a2be98cc9e81c8978696eefff4b9615e5c578d883374e6d2a5653256465994

Observation 91c8bbab-404c-4533-a26e-8e5a671bec3b · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Moment Sampling in Video LLMs for Long-Form Video QA Detecting mo- ments and highlights in videos via natural language queries

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.627656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.903197Z digest=sha256:9d9f38eba494cc810170e5191774ce116f077af50ab7f1d41339d0ae90ff9c54

Observation 465e99b2-37c3-4dc8-b02a-74dd3c30da5c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Moment Sampling in Video LLMs for Long-Form Video QA Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.907639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.907639Z digest=sha256:ddbcfabc9bb9976249e7811642f47ebbde71b8c9fd3e7b34eb11694854158874

Observation 8e7dd09a-aaf7-40b3-88f9-ac08053c8512 · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

Moment Sampling in Video LLMs for Long-Form Video QA Inten- tqa: Context-aware video intent reasoning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.604338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.912352Z digest=sha256:4030c92c546dee10b62150fd95fcdee26131d68f26d4c421afc8d55bc3f13e82

Observation 470f0515-2c8c-44d7-a420-31281e198412 · outbound

This paper cites VideoMamba: State Space Model for Efficient Video Understanding.

Moment Sampling in Video LLMs for Long-Form Video QA VideoMamba: State Space Model for Efficient Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.916938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.916938Z digest=sha256:aaa8495990f694bf1fb89001fed146cb96b7fcda2b8af6064f7b1cb228b0f61f

Observation bbd863f8-8223-49f4-ad73-f240bc6bc628 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Moment Sampling in Video LLMs for Long-Form Video QA Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.921793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.921793Z digest=sha256:f91a02c0e484756b0937b033ebb6b5a122e745e7839349bd887672a73ec3166a

Observation a8855266-b7ed-4f92-9cb3-7bb0aa14e3f5 · outbound

This paper cites Visual instruction tuning.

Moment Sampling in Video LLMs for Long-Form Video QA Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.926781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.926781Z digest=sha256:7ae3d95eaf94b90ec15c56182987a09ff1f84372d270ab81e4a39497e42fb058

Observation 261898cd-f898-4a34-9d86-cbe9130f2b24 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Moment Sampling in Video LLMs for Long-Form Video QA Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.581499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.930792Z digest=sha256:9782b4ee02726b9ae67679a01ce1871fa4ec19a174e3425a6275853e9eaf4996

Observation f81c324e-3a73-4427-af94-fa61f07a4d5f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Moment Sampling in Video LLMs for Long-Form Video QA Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.566355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.934919Z digest=sha256:faf3163533f409c5331c0924ae6479bb7f0e721f6be315755edac6da08cb0d59

Observation 1fe137ce-b37f-40c5-8df3-fcd07d178175 · outbound

This paper cites Morevqa: Exploring modular reason- ing models for video question answering.

Moment Sampling in Video LLMs for Long-Form Video QA Morevqa: Exploring modular reason- ing models for video question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.550422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.938773Z digest=sha256:721021d4cd8d32f4afccc194b21b5ca028ab8da33048125058f755614896f325

Observation 70f01212-f1ef-47d5-801e-0d132bbaa859 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

Moment Sampling in Video LLMs for Long-Form Video QA Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.943018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.943018Z digest=sha256:3c1df8f2a27c3d4e09b7b7006b284867c714e4a451ea89bb208f912279992d30

Observation 31330b83-b762-49fa-a97e-ee55966a5bbb · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

Moment Sampling in Video LLMs for Long-Form Video QA Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.536200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.947049Z digest=sha256:b9f5cbfa73aaf802c9f82c41279690bb345eb55f7c04f2f1c261a2d2bade964c

Observation 823d3faf-0e19-44f0-aa23-f33a464059c5 · outbound

This paper cites Gpt-4o: A language model.

Moment Sampling in Video LLMs for Long-Form Video QA Gpt-4o: A language model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.523064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.951307Z digest=sha256:b5a817818292fc5f64cf3d9a904d6fea09d117698d661da2bcfb2842b139a764

Observation 8181255d-6ba2-4a58-b2f5-53b59e8c6a1b · outbound

This paper cites A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames.

Moment Sampling in Video LLMs for Long-Form Video QA A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.510148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.955559Z digest=sha256:207d3615a105228833b22a8871b1bd1b1885ad64767e5888bdfb14ca2edab243

Observation ca045dc8-8383-476c-b9d0-4844e7fa69f4 · outbound

This paper cites VideoMamba: Spatio-Temporal Selective State Space Model.

Moment Sampling in Video LLMs for Long-Form Video QA VideoMamba: Spatio-Temporal Selective State Space Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.959963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.959963Z digest=sha256:3fc548c1d9901e493f3e24fdffb4aa958c8931ef5da08d38fe0c289aa1308979

Observation 577b636c-aeaf-4079-a9c2-1e941f30a828 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long- form video qa.

Moment Sampling in Video LLMs for Long-Form Video QA Too many frames, not all useful: Efficient strategies for long- form video qa

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.495921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.965062Z digest=sha256:03daa1c6f5029dfd6ec2dc407e8df724a102701f11082160ef0370ac918b96e9

Observation 2fed04de-bdd1-4585-84c2-de56e680f568 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Moment Sampling in Video LLMs for Long-Form Video QA Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.969397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.969397Z digest=sha256:615de0381f58c08e383debbe2c3caac5af2d104ecaae326d55afd3c1eb8c432c

Observation 5c0adb49-fb7f-4404-b2c7-75895a244b18 · outbound

This paper cites Cinepile: A long video question answering dataset and benchmark.

Moment Sampling in Video LLMs for Long-Form Video QA Cinepile: A long video question answering dataset and benchmark

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.472302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:13.973561Z digest=sha256:b856e4338a055cd4ce4d1124c3a1d4b3ea9cfb3f49a2d05383730b96daa31d33

Observation ac28df29-5b6b-4366-af50-ed3b8e31854e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Moment Sampling in Video LLMs for Long-Form Video QA Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.977936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.977936Z digest=sha256:a2d07d7766a844388d0097d84b4fa07a5aa293f18f46eca46d443cfaf6081d9b

Observation 79e68040-fd6e-4fba-ad06-fcc72405ecda · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Moment Sampling in Video LLMs for Long-Form Video QA Gemma 2: Improving Open Language Models at a Practical Size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.982480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.982480Z digest=sha256:b2ff3ba9e9e1b61516eaa2a07ad581619bd83211be8d860e9f4109753259326b

Observation 8d35a2ef-4da6-4129-862c-6155f01f3cc6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Moment Sampling in Video LLMs for Long-Form Video QA Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.987205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.987205Z digest=sha256:1a9ef607438efe02cf319a40b6351477c618af65272e22f57b56764a50183030

Observation 078344ea-6d76-4509-90e1-f29e6e851a66 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Moment Sampling in Video LLMs for Long-Form Video QA Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.991824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.991824Z digest=sha256:bbafbaf0b16aed3a44a4860ea8354c17d9fe78d070ac05104964ea368ae78030

Observation 3f38c2ae-c62c-4724-9799-691f83cdf3a3 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Moment Sampling in Video LLMs for Long-Form Video QA VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:13.996352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:13.996352Z digest=sha256:7e282a5bfed8676a3273eb22337d5eb34bebd12ea721ba4efc28c77e2b1c13ad

Observation 92e356c5-8227-412e-a364-fa5b39475ea3 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Moment Sampling in Video LLMs for Long-Form Video QA Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.457246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.000868Z digest=sha256:8a61787a5f8c89381c0c902c5b4eaaf9fa6c4cb3ea45c7b9240e1c0b2c89c80e

Observation 11b6c044-a12e-43f7-8800-da6aa44fb582 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Moment Sampling in Video LLMs for Long-Form Video QA VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:14.005166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:14.005166Z digest=sha256:636ebcd847eff5f3e3591354ddc86e784304415d85428c63cb030b4fb497de64

Observation e74df59c-1c20-4223-a6a9-cb9a49d65aa7 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Moment Sampling in Video LLMs for Long-Form Video QA Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:14.009498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:14.009498Z digest=sha256:51dc12ee9e40b80a42dbc554350bf7b9ca70251ec84273485b0e8602e0b61d25

Observation 99176eaf-ab80-4b96-8e45-ef60fb5f61ec · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Moment Sampling in Video LLMs for Long-Form Video QA Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:14.013379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:14.013379Z digest=sha256:33586238536c553eb4863899ebb1bb9b902b48b774b1fd58a69f660d9bb734cf

Observation e5c7cb81-7272-42a7-9555-94a7b3dba2f7 · outbound

This paper cites mplug-2: A modularized multi-modal foundation model across text, image and video.

Moment Sampling in Video LLMs for Long-Form Video QA mplug-2: A modularized multi-modal foundation model across text, image and video

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:14.017315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:14.017315Z digest=sha256:5aca42d83ac28421e31006b33d1bbd2dd6688991fbfef5e47d21e261b1ea4a32

Observation a7064fca-4f3f-446b-bf82-2a3d6191104a · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language alignment.

Moment Sampling in Video LLMs for Long-Form Video QA Clip-vip: Adapting pre- trained image-text model to video-language alignment

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.415406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.021257Z digest=sha256:6866fe309e5586e1cc907aafe5d6b19e1d931d6595bec3ec7cfc5c8821c1a4b4

Observation d06544b1-b752-48e8-a853-9417cb66e851 · outbound

This paper cites Multiview transformers for video recognition.

Moment Sampling in Video LLMs for Long-Form Video QA Multiview transformers for video recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.402132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.025374Z digest=sha256:d44d380fc577e4d5bab092b00aaf7492e2c7d5489654b3442ddbfcede48ac741

Observation 9028138a-c96c-4f7c-a8f4-d5afa5518c7a · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Moment Sampling in Video LLMs for Long-Form Video QA UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:49:14.029376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:49:14.029376Z digest=sha256:98a2e4cb638018869032d38f37d109ceddfdb70d3ad23fd94c1696c373c11965

Observation 39ffd586-a6c0-43d7-ba6b-99bad832ebc7 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Moment Sampling in Video LLMs for Long-Form Video QA A simple llm framework for long-range video question-answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.388344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.033745Z digest=sha256:d0ba2d259fe015ca0b1de8c486f2d3df06dc840ea4abb55bcfc2fa53c4bb87e5

Observation ca96d32f-d90e-499a-ada7-9d551cfaa18e · outbound

This paper cites Learning video representations from large lan- guage models.

Moment Sampling in Video LLMs for Long-Form Video QA Learning video representations from large lan- guage models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.374245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.037626Z digest=sha256:e2df2d20f81a249414e3fbc622b7f0376849358a23fc345f154ad31de47c0f00

Observation c1091b5b-b062-4252-a48b-813fe42e3f0b · outbound

This paper cites Rela- tional reasoning over spatial-temporal graphs for video sum- marization.

Moment Sampling in Video LLMs for Long-Form Video QA Rela- tional reasoning over spatial-temporal graphs for video sum- marization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:49:14.360106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:49:14.041691Z digest=sha256:4e657dcad849b349c8f1a82d46c3f57c0582720e3f74c6f3186bba5684e9c931

Pith citing papers

Observation b47ab44f-d0c3-4e78-afd8-b1ac796e9032 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Moment Sampling in Video LLMs for Long-Form Video QA

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.670603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:3726fd4750a58d19cfde8f1e6b1e3c22fb12377ef74cee265b46f1515ae79a7b

Observation a2d673b4-5787-45dd-ba9a-e0f2e25494bd · inbound

Answer Self-Consistency with Margin-Triggered Question Re-Arbitration for the CVPR 2026 VidLLMs Challenge cites this paper.

Answer Self-Consistency with Margin-Triggered Question Re-Arbitration for the CVPR 2026 VidLLMs Challenge Moment Sampling in Video LLMs for Long-Form Video QA

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:16:43.984219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T07:25:25.260193Z digest=sha256:4e2112a41befd7f2de324fd06aab6b2ae2506d53868c2f825239019a8b7fd843

Observation ec2ea21f-21c6-4803-b5c6-c34be3ea894f · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Moment Sampling in Video LLMs for Long-Form Video QA

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:01:29.190584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:fa554dc9b786413fc8a664cdf9efb5636e559aa06b2c28e5fa02884ab7319a1a

Observation b9625e01-e5d5-4a30-bb0c-6006fb438648 · inbound

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? cites this paper.

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? Moment Sampling in Video LLMs for Long-Form Video QA

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.988215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:30:33.428489Z digest=sha256:e780e68e87fdca3747895c11471a5ec09bb6d87adbe4a7915044173090a2eb0c