Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:18:49.027719Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.19460.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:18:49.027719Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:26.993426Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:51:30.627840Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9db3598f-9853-47f7-92c4-d2fc54454efe · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72076b41-50c3-43a2-94c9-56dab3c8f539 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe1f85c-45f4-4158-8036-90da990e5da3 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Lan- guage models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130a81ae-75d5-47e6-afa2-7b1689c72968 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75fc56b-087f-48ca-b8c9-02b5f4d8cfaa · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Extending Context Window of Large Language Models via Positional Interpolation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032fbfcd-72d3-44fc-b461-348aa8ea2cd9 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Training Deep Nets with Sublinear Memory Cost
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3e0c6f-df85-4f9e-afcf-ca4858684ca2 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87afc0f-a902-4040-b83e-2ccdf8d83d17 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8c4342-707f-48ec-afd4-dd58f7a92bda · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing InstructBLIP: Towards general-purpose vision- language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d39ef85d-e66a-4939-93b6-be7430344979 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Carbonell, Quoc V
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef8deede-0ef0-4dda-a51f-c4880bb54a2a · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7e8b83-a8d7-49c5-af49-4116617ec714 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0254b474-b400-4efc-b41a-b0ad6926d13f · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ba5ba2-f4bf-4df5-8c52-ed6ab6216899 · outbound
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c7812791-c6a1-431b-9757-69c4bd574e66 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702511e1-4a83-4f5a-a365-e04a74162392 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26554f67-49de-4f01-982e-89072e98f455 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f403274c-bd80-484f-b457-a47c2170bd8c · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2fbcb1b-fc7c-481e-ad04-a3026aca8db2 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4ece90-3024-4099-abc0-e2a9c28d6972 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaVA-OneVision: Easy Visual Task Transfer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c10b7af-dae0-447f-8b34-e5e2ff668878 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b12c7654-61c5-473a-b9b6-a37441062a41 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Llama-vid: An image is worth 2 tokens in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72dbc329-34c5-40b9-9250-cf6f1d7c0ab2 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fa7ac8-2a5b-4a0c-be7c-00968b382655 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Vila: On pre-training for vi- sual language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f6934b-25ae-4e31-8863-b41f7f2144b4 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Improved Baselines with Visual Instruction Tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc0fa21-6e0e-4470-add8-3ec019072f18 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Visual instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 375c1fb3-0bd4-4be5-a08e-3478fb860253 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d22fb7ad-5e82-4d75-9f09-38732f86db06 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing St-llm: Large language models are effective tem- poral learners
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation df05a570-1c62-4d11-b7ad-2696d90c3d61 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a7696c-fb1a-49f3-8dae-f8e67e313c6c · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 318d6140-e45d-4adf-8505-73b2dc1ca5e3 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gpt-4 technical report, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a05b3b41-1cc9-46d1-8e1e-ac0d238245fa · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing GPT-4V(ision) System Card, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0b310332-44b2-44df-9697-3ab7728a8237 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Hello gpt-4o, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d7e09123-524b-4721-ab18-e10a330f4c1f · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Per- ception test: A diagnostic benchmark for multimodal video models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 09d38a4e-fc63-47d0-abeb-2b0ef6059c06 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7364031-9522-42c9-bd11-01cc2acc7562 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Learning transferable visual models from natural language supervi- sion
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fe7973-96b5-444b-9284-9a4716d729d0 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Zero: Memory optimizations toward training trillion parameter models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789174ea-bc8d-40b0-ae0a-b48751e13b74 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3890664-597c-4466-b583-b757d81a20a9 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Moviechat: From dense token to sparse memory for long video understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba598f2d-d3f2-4e3f-aefa-3cf25108b366 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38878d5f-5a09-4f6b-8070-c769f23e7f2a · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaMA: Open and Efficient Foundation Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def39225-2242-4f4b-bb55-28bd1a1a49ca · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b2aced-f036-4763-b2fc-8a0db721a5dd · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3a4d02-a354-4c08-bae7-3a324569d062 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Next-qa: Next phase of question-answering to explaining temporal actions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d15fe8d8-4572-4417-b33e-e6dc7cc13461 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249bb2d6-462a-4700-affa-8237952600de · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea51b055-9ed7-4a4e-8877-40140064e02c · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014dcff1-068b-4b01-aed4-565e544efa92 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Long Context Transfer from Language to Vision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2145f0f-e25a-4d12-88ab-d7495dd6e669 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e61ed9f-6618-42e0-922a-6d32942ed904 · outbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183b11b8-f18f-4977-97fe-0a930ed3a14f · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.