Pith. sign in

Paper Citation Record · LEDGER

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.19460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19460 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:18:49.027719Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:26.993426Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:51:30.627840Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved33
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9db3598f-9853-47f7-92c4-d2fc54454efe · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.804862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.804862Z digest=sha256:1521d8f2eaf8c171b9c867dd05f34572d59a796e5a9bae3bc79060f9dfb11602

Observation 72076b41-50c3-43a2-94c9-56dab3c8f539 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.810075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.810075Z digest=sha256:20263637d56700842742b3d8570dc8f26bf8515ffa066cbf33867b5185a5e164

Observation 2fe1f85c-45f4-4158-8036-90da990e5da3 · outbound

This paper cites Lan- guage models are few-shot learners.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Lan- guage models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.814804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.814804Z digest=sha256:2df16a4d483c002a053961bd0030c9da628de486a7e12ee742f236ea63ca57dc

Observation 130a81ae-75d5-47e6-afa2-7b1689c72968 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.819643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.819643Z digest=sha256:53366e4870573528f3745d7e3e7d95b8993d473e4af90f78805f2807027998d0

Observation b75fc56b-087f-48ca-b8c9-02b5f4d8cfaa · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Extending Context Window of Large Language Models via Positional Interpolation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.825125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.825125Z digest=sha256:889f614717306e97c30f6f41a061aa0de46cbcf50ebe0e4698eefca0d6cde180

Observation 032fbfcd-72d3-44fc-b461-348aa8ea2cd9 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Training Deep Nets with Sublinear Memory Cost

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.830334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.830334Z digest=sha256:84e3e8ee3385a5c7b2d1d7c1ceca5fb0a9d17aab724d38d69b1a3c5c4e81955f

Observation 4f3e0c6f-df85-4f9e-afcf-ca4858684ca2 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.835878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.835878Z digest=sha256:2c870871dc17ab70863fdd126b00a3e76a96f2f24f1f57a52b6175959435a2fd

Observation f87afc0f-a902-4040-b83e-2ccdf8d83d17 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.841017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.841017Z digest=sha256:af6a2c0f63711b9b5c5965932e4e37450e8e30064773dcff048c020501ee6f80

Observation 5d8c4342-707f-48ec-afd4-dd58f7a92bda · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.709411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.846792Z digest=sha256:7639d9c64e9b1cf9dd779031bdd28cb4422455a343225d24bd2ed7ee80e056de

Observation d39ef85d-e66a-4939-93b6-be7430344979 · outbound

This paper cites Carbonell, Quoc V.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Carbonell, Quoc V

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.696856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.852157Z digest=sha256:64dd8ea7ba27f9270c70203a3cd7710c57c196f78c6c0e8eef5fb5f927e93191

Observation ef8deede-0ef0-4dda-a51f-c4880bb54a2a · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.856760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.856760Z digest=sha256:e0f7bc03f726ceba6d44f59c988c36edf9ef04fd37362572174c2eacaaf172ca

Observation 1c7e8b83-a8d7-49c5-af49-4116617ec714 · outbound

This paper cites an unresolved cited work.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:18:49.683720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.861579Z digest=sha256:3d40c3917180e759e4e9c2f70995ebc3a76123d2a01221e4f8edf25c0d32d152

Observation 0254b474-b400-4efc-b41a-b0ad6926d13f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.866004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.866004Z digest=sha256:84841c5295f1a1d087f25d91d91fbe6611174219593d587af2b3b5a63fc7db29

Observation 59ba5ba2-f4bf-4df5-8c52-ed6ab6216899 · outbound

This paper cites Gemini, 2023.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gemini, 2023

Reference 14

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T10:18:49.668616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.870798Z digest=sha256:fa4b638b45153f1b61ab3101e857838c7c42f3bd75514d8435d5ca5c136edc36

Observation c7812791-c6a1-431b-9757-69c4bd574e66 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.874459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.874459Z digest=sha256:cddee83a087479003785fc9df628ff5c77287432d059872cade2eb5a61537744

Observation 702511e1-4a83-4f5a-a365-e04a74162392 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.653153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.879095Z digest=sha256:aa92ef62685024c7c2e7f258197231addb9d6a2a37e61af1a90c71c5dae093d7

Observation 26554f67-49de-4f01-982e-89072e98f455 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.883094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.883094Z digest=sha256:bb1aab6b47c79f56a0dafee13cc5b300011cbb65c28c6eaab498a81d3f8664c3

Observation f403274c-bd80-484f-b457-a47c2170bd8c · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.638726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.886856Z digest=sha256:ac6628b509ca4aa2100f668a1a9cca1baba2c3c0b268d59bc8d199f29b059fc2

Observation e2fbcb1b-fc7c-481e-ad04-a3026aca8db2 · outbound

This paper cites SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.890691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.890691Z digest=sha256:d1373e347b7d964be958dc35ef9626111a5a7d80f6c4c92f0b894c0396ccd294

Observation 3f4ece90-3024-4099-abc0-e2a9c28d6972 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.894912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.894912Z digest=sha256:3e2c6b242e0f9e4c491952e1a247ee96fc4934647122e62dd4cf6241530ffd7f

Observation 3c10b7af-dae0-447f-8b34-e5e2ff668878 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.623640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.899206Z digest=sha256:c0761fe3e2325bf2404cff4be3c5a7e1f412416ad8bd486f8b58e36bfdb683a1

Observation b12c7654-61c5-473a-b9b6-a37441062a41 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Llama-vid: An image is worth 2 tokens in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.608876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.903388Z digest=sha256:d3faa5cac92a06197fdeb875295761b1126880b5ca831a364f15e9b5882fdd02

Observation 72dbc329-34c5-40b9-9250-cf6f1d7c0ab2 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.907774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.907774Z digest=sha256:ee6187db2d6e660e0fe494f8cf194f8591e4791ad10c189525d8e73a16c65623

Observation 60fa7ac8-2a5b-4a0c-be7c-00968b382655 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Vila: On pre-training for vi- sual language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.912065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.912065Z digest=sha256:2e49fca625519bafde72d5de6ca40890673db1a98671572d51fcbcc5a779a165

Observation 24f6934b-25ae-4e31-8863-b41f7f2144b4 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Improved Baselines with Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.916464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.916464Z digest=sha256:a3e011b0e2e0ed0c76202d3679040d5a356ba8d317579711459afadb5c526fdf

Observation 6cc0fa21-6e0e-4470-add8-3ec019072f18 · outbound

This paper cites Visual instruction tuning.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.585250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.921543Z digest=sha256:855e6f826bd0f6eeef616439b2ea7fb16a64325edb810597c3c340db92a51cf3

Observation 375c1fb3-0bd4-4be5-a08e-3478fb860253 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.571689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.926840Z digest=sha256:f415b4ec867f2e4e7ab9fb9620a2e0d39df5647ce1c491ca3a87f58d52f1298b

Observation d22fb7ad-5e82-4d75-9f09-38732f86db06 · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing St-llm: Large language models are effective tem- poral learners

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.556981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.930923Z digest=sha256:7e6d62f9427c4fa0b7ea18de7a17455eb2d1c8a9005baa740a5d566a346d4d61

Observation df05a570-1c62-4d11-b7ad-2696d90c3d61 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.935134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.935134Z digest=sha256:a6c0f3a4954a041b4db4ead980352a5c904756b9f80c6f8e57395bcf6cc0a606

Observation 84a7696c-fb1a-49f3-8dae-f8e67e313c6c · outbound

This paper cites an unresolved cited work.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:18:49.540791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.939475Z digest=sha256:f48584ecd7d250fecb85db385fa4fd1fc3d2dde0b3f697a3daa464c6a45ca56f

Observation 318d6140-e45d-4adf-8505-73b2dc1ca5e3 · outbound

This paper cites Gpt-4 technical report, 2023.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gpt-4 technical report, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.527375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.943464Z digest=sha256:c35a3ff1efc7b9e00b45860076392b8a718de7851b570f76e98eb960c91c0197

Observation a05b3b41-1cc9-46d1-8e1e-ac0d238245fa · outbound

This paper cites GPT-4V(ision) System Card, 2023.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing GPT-4V(ision) System Card, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.514325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.947683Z digest=sha256:cf1d2a040b1c2f15e3591d82640e8e5b5fadd3cb962f8a925bcb0773cab69823

Observation 0b310332-44b2-44df-9697-3ab7728a8237 · outbound

This paper cites Hello gpt-4o, 2024.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Hello gpt-4o, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.500726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.952650Z digest=sha256:570550ba4f755103337a37a3c9c0f0298a76f895ab0c5ef5ca6b346e6984c306

Observation d7e09123-524b-4721-ab18-e10a330f4c1f · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Per- ception test: A diagnostic benchmark for multimodal video models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.486239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.957475Z digest=sha256:1c87b93eecd9ccc860aab4d0784e7cceb61b9c1d391a55bbbda41ee070dc8c84

Observation 09d38a4e-fc63-47d0-abeb-2b0ef6059c06 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.961665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.961665Z digest=sha256:87db7cf2640af032652f4248fb4cdd38c70674d7969b42fba993e9724e9899c6

Observation c7364031-9522-42c9-bd11-01cc2acc7562 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.966731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.966731Z digest=sha256:660f9366adfabebb6a15b0e73319046ad1269c20911e206000d1c6892e65458e

Observation d5fe7973-96b5-444b-9284-9a4716d729d0 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Zero: Memory optimizations toward training trillion parameter models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.970446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.970446Z digest=sha256:1e3458c19457989ced2f1216fc6014f93d80022b692c56ae0860ddb4bf3a0cba

Observation 789174ea-bc8d-40b0-ae0a-b48751e13b74 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.974203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.974203Z digest=sha256:373cc03b15900cb0c8e9589ae3b8f86cb549af6d800dc2e08076ac60ef22dd86

Observation d3890664-597c-4466-b583-b757d81a20a9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.453690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.978180Z digest=sha256:074acf7bff64c099dfd29a60cd8a23ba63460b577e0e8ee411178afeb6c06e67

Observation ba598f2d-d3f2-4e3f-aefa-3cf25108b366 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.981971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.981971Z digest=sha256:4b0f609b62f528d07a15d9593a3f1e7262f657bd814c56b43bfae11991899632

Observation 38878d5f-5a09-4f6b-8070-c769f23e7f2a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.986058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.986058Z digest=sha256:fc8cacfa83649671569954a876fb5c183f78e6def2986e538c2d7196a0c2f81f

Observation def39225-2242-4f4b-bb55-28bd1a1a49ca · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.990518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.990518Z digest=sha256:4488d1bbdfddf7f6338a0ce389ec53252a3e78922072b29285157ea47d62c677

Observation 19b2aced-f036-4763-b2fc-8a0db721a5dd · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.994759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.994759Z digest=sha256:e3b5b872f223d2904ab0f56e8dc77ec0c46a190634347b51688120209c24e075

Observation 0a3a4d02-a354-4c08-bae7-3a324569d062 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Next-qa: Next phase of question-answering to explaining temporal actions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.428747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:48.999110Z digest=sha256:cc97c88fd6cef669cb5e86e24498c96ebf0d50218374d1715b5dce04183c4bc0

Observation d15fe8d8-4572-4417-b33e-e6dc7cc13461 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.003350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.003350Z digest=sha256:f9e6cd272e7d4d0e4c2da03721dcaccd53bf5d2db6afbf7364a42eaf8d46d6e0

Observation 249bb2d6-462a-4700-affa-8237952600de · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:18:49.413523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:18:49.008889Z digest=sha256:5bc2dfa0fe3b274b8c733acfdc41ad15319011618dfaa3477d6a9658b80d6271

Observation ea51b055-9ed7-4a4e-8877-40140064e02c · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.013908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.013908Z digest=sha256:7dd9faf7d2666dc451594e81b480afb966a4e856e6496884e8242714b54afacb

Observation 014dcff1-068b-4b01-aed4-565e544efa92 · outbound

This paper cites Long Context Transfer from Language to Vision.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Long Context Transfer from Language to Vision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.018500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.018500Z digest=sha256:12d1de9260874e6b95f149571cbc9d20a49bbcf9b27dc19d92c3c026336cea6f

Observation a2145f0f-e25a-4d12-88ab-d7495dd6e669 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.023170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.023170Z digest=sha256:0b3d043cb544b25817d044a5a9919dd56713764b6712e310f529f154e9c78214

Observation 6e61ed9f-6618-42e0-922a-6d32942ed904 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.027719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.027719Z digest=sha256:c21efa95fc33be7795897089306d489413d68f99d254cb48209bd97f9b748577

Pith citing papers

Observation 183b11b8-f18f-4977-97fe-0a930ed3a14f · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:30.674736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:51:26.993426Z digest=sha256:435e54bf329727fc19d6a1761cf272b1c160b5761a3f41fe8f47be29e466b487