Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:49:14.041691Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 4 inbound Pith citation observations for arXiv:2507.00033.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:49:14.041691Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:25:25.260193Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
61 of 61 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 8a7cd85e-a487-4551-9a63-9dc64925b8fa · outbound
Moment Sampling in Video LLMs for Long-Form Video QA GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23936293-6c31-4d1f-88fe-5720b91b9cee · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Combining global and local attention with positional encoding for video summarization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e4a9f1be-fcea-436b-a530-01cb420116af · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Vivit: A video vision transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd2bd53-860c-46aa-a80d-23a5f3ce19ec · outbound
Moment Sampling in Video LLMs for Long-Form Video QA OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a87b98-988e-412e-83f0-2636b3696bef · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b7bf85-dd4f-4434-a698-7f8861a27c1e · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Memory consolidation enables long-context video understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d921aeaa-0d76-477c-a8b7-e45be23c970b · outbound
Moment Sampling in Video LLMs for Long-Form Video QA End-to- end object detection with transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a324b76c-e902-46ba-b997-561e3f0b9365 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 214f83fd-d1d8-4a77-a32c-e6ad694f31b4 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Cosa: Concatenated sample pretrained vision-language foundation model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9dc9d8a0-7fbc-4a34-b545-83c8775bd0a5 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2005cc7a-ba8c-4396-b0f2-d634013dcfe5 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab27103-2808-4dd0-b1b6-54589ecd8f80 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9654396f-d8ef-43e0-b316-33d23be3021b · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Coot: Cooperative hierarchical trans- former for video-text representation learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6eef0292-7efd-43ef-a241-e8eb2a9395e0 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 466fa2bd-4ed7-4c96-9546-c53016ae58ef · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Efficiently mod- eling long sequences with structured state spaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9309774-7048-4215-bc1e-30cd7ef99747 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Pidro: Parallel isomeric attention with dynamic routing for text-video retrieval
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13ada487-7703-4192-8db8-70963d61915a · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Creating summaries from user videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad0e74a5-14fa-4a27-8072-591d0094f9c0 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Object-region video transformers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 33b56590-b2bf-4394-802e-9494323429b4 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Video re- cap: Recursive captioning of hour-long videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1dac3981-d347-429a-97dc-379efe6db7dc · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea832f9d-264c-44d1-94ab-21da59aac086 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efeea2b1-fdcd-4b61-b500-5f39926634eb · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Language repository for long video understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e718c09-8f90-43f4-8640-eb935888ee78 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Large language models are tempo- ral and causal reasoners for video question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1316294e-d040-4614-b0dc-2b48dc995cb7 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Movinets: Mobile video networks for efficient video recog- nition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdaeb2f-eeef-4d31-9c75-e94f3a9f34da · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fbc22641-7272-404d-bf21-54a72e19dfea · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Dai, Zhifeng Chen, Claire Cui, and Anelia An- gelova
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 273140d3-8ba9-4586-99c1-e015116524c0 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Cast: cross- attention in space and time for video action recognition
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bc5caf1c-a410-43dc-8e98-e93adb3ef1c0 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 91c8bbab-404c-4533-a26e-8e5a671bec3b · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Detecting mo- ments and highlights in videos via natural language queries
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 465e99b2-37c3-4dc8-b02a-74dd3c30da5c · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7dd09a-aaf7-40b3-88f9-ac08053c8512 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Inten- tqa: Context-aware video intent reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 470f0515-2c8c-44d7-a420-31281e198412 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA VideoMamba: State Space Model for Efficient Video Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd863f8-8223-49f4-ad73-f240bc6bc628 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8855266-b7ed-4f92-9cb3-7bb0aa14e3f5 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261898cd-f898-4a34-9d86-cbe9130f2b24 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f81c324e-3a73-4427-af94-fa61f07a4d5f · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1fe137ce-b37f-40c5-8df3-fcd07d178175 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Morevqa: Exploring modular reason- ing models for video question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70f01212-f1ef-47d5-801e-0d132bbaa859 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31330b83-b762-49fa-a97e-ee55966a5bbb · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Query-dependent video representa- tion for moment retrieval and highlight detection
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 823d3faf-0e19-44f0-aa23-f33a464059c5 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Gpt-4o: A language model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8181255d-6ba2-4a58-b2f5-53b59e8c6a1b · outbound
Moment Sampling in Video LLMs for Long-Form Video QA A simple recipe for contrastively pre-training video-first en- coders beyond 16 frames
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca045dc8-8383-476c-b9d0-4844e7fa69f4 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA VideoMamba: Spatio-Temporal Selective State Space Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577b636c-aeaf-4079-a9c2-1e941f30a828 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Too many frames, not all useful: Efficient strategies for long- form video qa
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2fed04de-bdd1-4585-84c2-de56e680f568 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Learning transferable visual models from natural language supervi- sion
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0adb49-fb7f-4404-b2c7-75895a244b18 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Cinepile: A long video question answering dataset and benchmark
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ac28df29-5b6b-4366-af50-ed3b8e31854e · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e68040-fd6e-4fba-ad06-fcc72405ecda · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Gemma 2: Improving Open Language Models at a Practical Size
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d35a2ef-4da6-4129-862c-6155f01f3cc6 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078344ea-6d76-4509-90e1-f29e6e851a66 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f38c2ae-c62c-4724-9799-691f83cdf3a3 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e356c5-8227-412e-a364-fa5b39475ea3 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Internvideo2: Scaling video foundation models for multimodal video understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 11b6c044-a12e-43f7-8800-da6aa44fb582 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74df59c-1c20-4223-a6a9-cb9a49d65aa7 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99176eaf-ab80-4b96-8e45-ef60fb5f61ec · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Next-qa: Next phase of question-answering to explaining temporal actions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c7cb81-7272-42a7-9555-94a7b3dba2f7 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA mplug-2: A modularized multi-modal foundation model across text, image and video
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7064fca-4f3f-446b-bf82-2a3d6191104a · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Clip-vip: Adapting pre- trained image-text model to video-language alignment
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d06544b1-b752-48e8-a853-9417cb66e851 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Multiview transformers for video recognition
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9028138a-c96c-4f7c-a8f4-d5afa5518c7a · outbound
Moment Sampling in Video LLMs for Long-Form Video QA UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ffd586-a6c0-43d7-ba6b-99bad832ebc7 · outbound
Moment Sampling in Video LLMs for Long-Form Video QA A simple llm framework for long-range video question-answering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca96d32f-d90e-499a-ada7-9d551cfaa18e · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Learning video representations from large lan- guage models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c1091b5b-b062-4252-a48b-813fe42e3f0b · outbound
Moment Sampling in Video LLMs for Long-Form Video QA Rela- tional reasoning over spatial-temporal graphs for video sum- marization
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b47ab44f-d0c3-4e78-afd8-b1ac796e9032 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning Moment Sampling in Video LLMs for Long-Form Video QA
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a2d673b4-5787-45dd-ba9a-e0f2e25494bd · inbound
Answer Self-Consistency with Margin-Triggered Question Re-Arbitration for the CVPR 2026 VidLLMs Challenge Moment Sampling in Video LLMs for Long-Form Video QA
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec2ea21f-21c6-4803-b5c6-c34be3ea894f · inbound
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Moment Sampling in Video LLMs for Long-Form Video QA
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b9625e01-e5d5-4a30-bb0c-6006fb438648 · inbound
Rethinking RAG in Long Videos: What to Retrieve and How to Use It? Moment Sampling in Video LLMs for Long-Form Video QA
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.