Pith. sign in

Paper Citation Record · LEDGER

A Survey on Benchmarks of Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2408.08632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08632 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:49:51.472790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.635040Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6df3758a-82c1-4555-967f-f0116940e192 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning A Survey on Benchmarks of Multimodal Large Language Models

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:33.237633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:cbc0346753f4ec0a05f69e396477777db1e62b518b25625b203b2d794e083a2c

Observation dc6c9300-5af4-4b18-aca5-a9ab9bb1ea3d · inbound

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark cites this paper.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Survey on Benchmarks of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.472790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.472790Z digest=sha256:f0a739d4e72f85b492e3b6d1f1f5cffcd68f5d95ab41c602df9b3fde7bcd87af

Observation 98842e2d-20fb-41ae-86cc-e0076730f568 · inbound

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models cites this paper.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.875661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.875661Z digest=sha256:c6b3409800266b1573c286c42924e64088d7daf25b5173ebd880d68aed089372

Observation c4541fd2-edff-4e88-8ec7-b424820fe260 · inbound

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization cites this paper.

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization A Survey on Benchmarks of Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:47:13.058347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:47:09.500229Z digest=sha256:a590ccc8e27f9ca66bbb604c80ae7d2cd792ad0c56fe111da8a7ad01cd7d2656

Observation 7e2f3ed1-726d-445c-be3e-c0fba0554758 · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias A Survey on Benchmarks of Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:22.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:22.093509Z digest=sha256:24526d528de0daa6c76bd8d170a4ebd5057fbddc7e076df473cf8830e829e9e2

Observation 792c852f-1151-4710-90d4-5a20d32e9ccb · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:13.867592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:13.867592Z digest=sha256:3b3e504a0d4cb3f1c462655073c3fa1fc1a4f1beb54ac521e74a539135f7919e

Observation 93700d1e-7cf0-4a5d-8199-fbe04144cfe1 · inbound

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests cites this paper.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests A Survey on Benchmarks of Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.018980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.018980Z digest=sha256:a01a057e42687aec8e2f5b716b2bd2935deb7f649b2e0695788dcd02ed0b3fd8

Observation f51502c7-a87f-45f3-a521-4b4b84b61990 · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:59.994525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:59.994525Z digest=sha256:c06fea2b694b8a2c74c37bb6395ecf8fd386ebf1959bbf532f5198e78332676d

Observation b3b66b49-a8e3-45ae-99c4-5c84ace87e7b · inbound

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? cites this paper.

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? A Survey on Benchmarks of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:56.960671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:36:56.960671Z digest=sha256:d81cbc0ee419a62220a72eca6e5b3affd0a1f6ba50d075085beda3741011b547

Observation 2bd0f30d-8816-4898-adaf-fc8413a2e385 · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Benchmarks of Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.436309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.436309Z digest=sha256:786e68a6bbb755043088f0b5e4f522a524c29785f386975da1570350b767bb31

Observation ab6c4cd2-5071-4449-802b-8455bbee0d98 · inbound

Re:Verse -- Can Your VLM Read a Manga? cites this paper.

Re:Verse -- Can Your VLM Read a Manga? A Survey on Benchmarks of Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:16:54.122059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T23:14:25.088197Z digest=sha256:b6ce077bda26ee68462555778dc6429afd783b96e42caaf4265e2f5d37b97d57

Observation c3afe12d-cd83-4cc8-9109-5f02e39d92c9 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM A Survey on Benchmarks of Multimodal Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:48.481837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:48.481837Z digest=sha256:4d83a9a323f6ad26f655be56e13620c081f08562a1dbc23e9cdf246ef93c67f6

Observation df52c407-f409-4ba5-a79a-a4dbcea8c987 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding A Survey on Benchmarks of Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.843114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.843114Z digest=sha256:0ceca211e510951a40597c34e7fb6445dc2b64c57510c31d35ce82f45b4602dc

Observation 9e7b8396-5a13-45bb-8dc1-0c787d59870d · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks A Survey on Benchmarks of Multimodal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:01.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:0ee494c0ae6708e9a33ff245ba1624e9e28f6b3efc11674a45379cfcf1f7f31d

Observation 50d669ae-960f-4164-a420-1135b47ce2a0 · inbound

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments cites this paper.

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments A Survey on Benchmarks of Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:00.913053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:46:15.107175Z digest=sha256:009af5d5f1b6c4c0ba9227190ee7e1161017c2aed8ee4a2027b64a92e003c986

Observation 37c82ed5-a722-435d-9946-cd0650edd23c · inbound

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models cites this paper.

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.559769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:23:46.371799Z digest=sha256:ea326b0baece429da1233a336a89b1981ad52f4ea0e3da3d9f73f64a7d66ad29

Observation b1bf0ca9-97bf-4be4-86a5-21bfcd4fa148 · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety A Survey on Benchmarks of Multimodal Large Language Models

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:05.630683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:c8eceb3fd682f56f1857474202d42642dd165f44ccf351cbef26efb8e3f4691a

Observation 2c46244f-d023-4efa-9798-d884440b9e7b · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing A Survey on Benchmarks of Multimodal Large Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:46.595108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:b40880d8fc24db0dc73fc5928b4754448759dae1aa5339305a8b9c3a0afaaea7

Observation 7d41d828-6c75-4b97-a4f0-ffa9c0d0d6a4 · inbound

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? cites this paper.

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? A Survey on Benchmarks of Multimodal Large Language Models

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:03:39.464904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T19:03:29.501299Z digest=sha256:dd6b97e01bcd78a13bf2244c7efe4cbebe09a76656a66fc4d935bd0f884f1447

Observation 601467ae-c5fe-497f-8db2-94132ff06ce8 · inbound

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models cites this paper.

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.630245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:16:45.665440Z digest=sha256:4dbfcaf9c6b7d62f02790b9fd598f5ef3025d722dc0d733bc5ebf7afa6d4ba27

Observation 3f26418d-e546-4b78-9e59-d466642eec04 · inbound

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments cites this paper.

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments A Survey on Benchmarks of Multimodal Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.944986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:36:11.721143Z digest=sha256:c622fb6a9258aae62a033482f60eac9f99d133e42e09153ee4855fb5a2193e28

Observation 80be2a3e-6e86-4a65-bfd2-480ed0f1d975 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 A Survey on Benchmarks of Multimodal Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:41.637190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:632b5a920aef94a1cdb35380ca404ed954335b1b66c6711c4352c782e1cb33fb

Observation f3148e2d-14da-4ba9-b0e5-0a7ddd658545 · inbound

PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement cites this paper.

PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement A Survey on Benchmarks of Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:18:42.939955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:18:42.939955Z digest=sha256:c294ca26114e8a49a6702a9eeee5ad50ec88712754ab40bad9c9beaf3cad302b