Pith. sign in

Paper Citation Record · LEDGER

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2404.05726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05726 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:44.831913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.683925Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10ff7284-3e28-4fa0-9804-a77c502c2499 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.539507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:8210c1415892896169644e99eff00eae0c71772a24c8d8379b07b4eb1eb82d98

Observation 1fda6659-f351-4b47-9145-27bc985f5b95 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.557379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:70e5a9cd71c9394c85a8655fe6a36fef62433ba198038b999e810f210b775355

Observation e3558315-0c00-4d59-819c-9fbc06c7dbbb · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.429141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:458fa71115e7b52cb661c8d2422cdfab45b5eb9abb6789bf2726eb6eac486403

Observation e91f77fe-e0a8-43dd-a17d-f75572f8d718 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.729626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:f3715382bb534a63e96fd613d1fb451a8a045bbfbcafcd0e24e3174492a36416

Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.831913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.831913Z digest=sha256:d45a0b49f424f8a6d920c67c0291365b543ac21560e6ef905ce1d4a2aca70bee

Observation fdc0e26e-7bbc-413b-abce-bc2b993963ca · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.772596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.772596Z digest=sha256:bcee0912f4930b735e5ca9b3114a61cca0c98048b877c89fbed9b4ebd07c0640

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:6d4022acbd9bb72c8f1e2b5b9c4d7f21edbc62ee0e3f5d3ad736edbc7df39b77

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:7674ad800e1efd63e0f381ebe2c1de98a2dd6c1aa56974b321611747ec1ead51

Observation 604b0fd2-5bef-4b08-adb0-c9b68053db71 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.327746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:54f9e995256b712a969ed58a31e4f8fe0b6d59bb4f052cd31d3433c3c78cf57c

Observation bb49de16-191a-4ae0-9658-6af5cbc4c952 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.819686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:c9b8437bcd776493bb4b7b98d944dea9cbd69f07cc6e744d1734b51de66d1a08

Observation 6a8ce2a8-7760-465a-8253-1ddea61a5315 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.685514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:548fef221b74477e41abfe128a5bda058e2a22bcab319c80537fed2ab1e6a54d