Pith. sign in

Paper Citation Record · LEDGER

MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.00468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00468 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:37.549174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:58:15.808350Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 924fe3f5-033c-46b1-8e68-14981efdb8c4 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.549174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.549174Z digest=sha256:0a4365fdc73c0cf15476fd7ef41c38a068a6bc03922393bf84370378171166da

Observation 0b543bb0-13b7-4da3-8b8f-97b0e785143d · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.854524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.854524Z digest=sha256:46179280916b7c6398f0a80769a1e02bd1f8ab7d22815841c2c544eef7fd31e9

Observation b20e21dd-f66b-48fc-9468-c5b5bfce7573 · inbound

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models cites this paper.

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:31.913836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:31.913836Z digest=sha256:ac8cc25f8295f67e9e6ff49c2da83419732fd1da44e56a57a9c0ffe7cdbb6e24

Observation 43819a59-7a02-4009-b212-5a7a357ef73b · inbound

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation cites this paper.

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.673244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.673244Z digest=sha256:e7fa2a0828db4dc0d82bebc7166c6388a422b0d933e39fbc8d06fbf43b8c0d65

Observation 397f078d-7416-4ac3-9335-7e91e6ff9a3a · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:09.400501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:09.400501Z digest=sha256:8492a7dfd72cf02b8ec35a51610d2556ee35d22c0f1a74a41885dc92927ff528

Observation 4af5246c-d98b-474a-b118-4a42e8a91ecb · inbound

BabyVision: Visual Reasoning Beyond Language cites this paper.

BabyVision: Visual Reasoning Beyond Language MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.655028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.655028Z digest=sha256:ffd690ceb8adb55d7a4708f684075a04d1b4e560ba743eccaa32fc1b7833fac8

Observation 051c8370-01b6-4caa-add3-a41064bfe29d · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.809699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:867e25289735bf2f62f19761529ea529344372e83f41cd7c27bf20fc7137ba27

Observation d20946ab-dcc9-478f-94be-b9b502577198 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.142910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:6442cf4f3eb136ebd406881a214bc54afc4667a40c8f6ed9115f6cfb49e4122e

Observation 9b425f9c-4190-40e3-93ca-1d4fe5f9d3d1 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:94aa1d690722101f67fcea04a244358a6b451d0acf8adb71fb5764b3e153c3c5