Pith. sign in

Paper Citation Record · LEDGER

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2401.10529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.10529 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:57:07.707953Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T19:13:40.522467Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 90f2d030-8bf9-4556-ab50-675d6e5521a9 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.446521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:e358bef7dad909d1283730ae83d255ace27c861acf0b615a15ae07b7d6d1d1a7

Observation 04cca4ad-d10f-4e47-966c-7749221d48db · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.548843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8e1641340ba6efa255d26f77501141bd58af232e83741e18509edda6dfa44358

Observation f2f9a9dc-f0df-4599-93ea-7c7e563967cc · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.136717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0af7f64a04ed01b50870ab3b72988250120064e360595a9243bbae6c6426c8b3

Observation c98dafb4-293c-4733-8528-d5619c45cda9 · inbound

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models cites this paper.

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:07.707953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:07.707953Z digest=sha256:cf8215f4b071d3560898a5fa7453cc6c446d97f7bfd4f4e6356d2dd0069da5c1

Observation 0798781e-ecbc-47a2-85d0-3f8cf615c646 · inbound

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding cites this paper.

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:18.668518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:18.668518Z digest=sha256:da14379eeefb9e62b48fddcbcb7d15896df88af49f628aa399cdb995ec6aa944

Observation 990122ee-1d85-491e-8966-27c4fad6d45e · inbound

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images cites this paper.

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:34.691445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:43:34.691445Z digest=sha256:5e24fefeaee0a807f87dbaef8e3c24d33aae05ba8fe776a44051c65874585fb7

Observation 7b28090b-ee75-4c24-8f16-201a27e47f49 · inbound

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning cites this paper.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.137629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.137629Z digest=sha256:e322e01cbd72f21ad2eb53f56744fdf53a54e5a5aee2ae8412d05d56df66b17b

Observation 275553a3-2d36-4427-becc-762e4e7daf42 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.464295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.464295Z digest=sha256:019b28e0343dc793b1f97e51f9e8daae86aa1848ebe4f6998ad84c859d42a7e0

Observation 63eefdf4-10f7-41b5-a24d-b1cf97df0815 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.886467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.886467Z digest=sha256:b7fb9ebecddd06b7bbccddc61078e3c0e8079150420231b5ec0f73feebb44c56

Observation f98ebe7d-445a-4880-9b65-cad8f634549e · inbound

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models cites this paper.

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:14.066097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:32:14.066097Z digest=sha256:7868b89cc70e7a8eda4207d97f48ef695133d84fb3d265af7a1733fd5983feab

Observation 8b18b037-8b02-4f4e-86ca-244c91950496 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.926817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.926817Z digest=sha256:d9b699ef2762e108dc976ffc68de94c466da18de010020c546bed7bb170f4f54

Observation a0c0d88c-c8f5-4176-9148-54c8e93ec1ad · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.674180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:bf57f9e5a8574498fc117e89d8aa27a46439ce0ddcdc574c1864f5cb385d6a1d

Observation 2583d57f-3e94-417a-bda6-ee3a603071d1 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.524351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:db0972c550205eafa271aa0f22eabaa691a36c2115a9cb7925ef58332afcf72d

Observation f5853c1d-4920-472a-b56f-c7429762a3f2 · inbound

Beyond Retrieval: Analytic Memory for Multimodal Agents cites this paper.

Beyond Retrieval: Analytic Memory for Multimodal Agents Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T07:06:05.783875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:06:05.783875Z digest=sha256:3aa577a33174cdc000daa854936d25372a4bcc5680095ccae42cec38156a25ab