Pith. sign in

Paper Citation Record · LEDGER

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2404.05726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05726 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:44.831913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.683925Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10ff7284-3e28-4fa0-9804-a77c502c2499 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.539507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:2559beb7d04e756e247ce473da64148b872975c00766929c5089bed914eef6f5

Observation 1fda6659-f351-4b47-9145-27bc985f5b95 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.557379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:32cadab78e1a9bdb846721d817f60edb7b0af611e9473862e2d5b8795167d5fe

Observation e3558315-0c00-4d59-819c-9fbc06c7dbbb · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.429141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:68211bf9315e7094c93db4892bd4e3d1ecdae43713ba21c3099b110fe51b2696

Observation e91f77fe-e0a8-43dd-a17d-f75572f8d718 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.729626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:91ab398f6404aa9ec9c1d48fd9fb2ae2b95b59b028b46ea574b674210d454e6f

Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.831913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.831913Z digest=sha256:a696ba91090eb99bd0613fc6cb6a9c43d24fea252866f5140bbdd67552a41a9c

Observation fdc0e26e-7bbc-413b-abce-bc2b993963ca · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.772596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.772596Z digest=sha256:c4464be302fff9eb9b00f0ba6640f6f3e2a5c21bf9685b06ff812cd51331931e

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:f701ee9dae195f28316249815b0a1f5ff82b0676dfcbe6cc0d28f94d054564e2

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:e18d0120fdb33a6bde7f916617cc2b3f0605c198817f302255768ac63794a89f

Observation 604b0fd2-5bef-4b08-adb0-c9b68053db71 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.327746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:537a3170a1c3f7a38b86f7b6917f413bdd41d4853da9b413888ab1c030384cd0

Observation bb49de16-191a-4ae0-9658-6af5cbc4c952 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.819686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:9dee201dd733e46e885b15e06e3fb84aaff4c41ffaa6967b6fe888a77643b4cb

Observation 6a8ce2a8-7760-465a-8253-1ddea61a5315 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.685514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:2552939264ca6f3755910d9433f7556c392e0ca809ad3b6d045b3ad845070c09