Pith. sign in

Paper Citation Record · LEDGER

MANBench: Is Your Multimodal Model Smarter than Human?

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2506.11080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11080 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:36.266807Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cbbc84f-f1cd-4c8e-94de-025e13f9b210 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

MANBench: Is Your Multimodal Model Smarter than Human? MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.019433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.019433Z digest=sha256:b5bb8e5f4277dbcd7442f95207da0130ff60ba81a3f9544fce004731efd0499f

Observation 6fc41757-0bb0-4ca5-9a50-5ec21542d75b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.100287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.100287Z digest=sha256:d2ad4935c134c0e64527dd7d707261539995322fadd06421e53bfc1231f9ca4a

Observation c4c0966f-75f9-43c1-94ef-c43269d723e0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MANBench: Is Your Multimodal Model Smarter than Human? Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.203398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.203398Z digest=sha256:c3f75bc3bb9a991b25b7916b4311c4d2536b30e5e8641ed4ca0b17dce542759f

Observation faf179a8-d2c9-4552-931e-acdd0bd8466a · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

MANBench: Is Your Multimodal Model Smarter than Human? GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.269754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.269754Z digest=sha256:2736cfcf972b6a2975483239963e0cd4f4cac1383a986d753cd77c99bc4aeeca

Observation e79d540c-5bd5-4e1e-b525-ad946d81fda6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:38.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.342134Z digest=sha256:1e7e66eb71693da3e4f26fddaa56a3ac4da0f5891fec17c547527696cbb2553e

Observation 46b97ec5-2d1c-424f-aaac-7a0c7fe1c17b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.922036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.413796Z digest=sha256:45cf5009448368e64329f75035874aa3c39c7d9fcd820caf910f92aef68aed75

Observation 45ff32d9-7ea2-4adc-885d-acb2aed40042 · outbound

This paper cites The Llama 3 Herd of Models.

MANBench: Is Your Multimodal Model Smarter than Human? The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.473651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.473651Z digest=sha256:352d58933abe019a9c31c0beab60465a251f26dc042fa7658433535c8f7588b0

Observation 6840f7bc-b91c-4ab6-8362-746e39cd4238 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MANBench: Is Your Multimodal Model Smarter than Human? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.576525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.576525Z digest=sha256:dbf179e0dda58c72e8f4baef9afa1d50cf6d74fa188298f52667d84afd87ce38

Observation a7295d2f-a5da-461c-bb3a-9925fd622e13 · outbound

This paper cites OpenAI o1 System Card.

MANBench: Is Your Multimodal Model Smarter than Human? OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.681836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.681836Z digest=sha256:a17dc381cd406b0ab6a460056a5b593d39dee622ee3c50ec991ba01b5cba2da2

Observation 79938242-1aaa-48e2-8de3-877f771a804b · outbound

This paper cites DeepSeek-V3 Technical Report.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-V3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.753861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.753861Z digest=sha256:2e06b43801bf17bf6d92e87107d7127c4e5e50f62b4c133d9537b5e555c6b2b2

Observation a1dfed1f-bfa6-46e5-a649-b08ab28235d8 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.852870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.852870Z digest=sha256:356c58918fa93d2719e8f6990a254454e8121bc8e7b4aa0daa761b6d9e0b1fdb

Observation f3cceeb6-bdc8-40de-8901-36b3e3619c04 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.952899Z digest=sha256:76236a2ee7a5ee644a2395ef8ae39f7efc401e57edc28680c2feeb09bbf10793

Observation ced02137-89da-4243-82e2-44d681229640 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MANBench: Is Your Multimodal Model Smarter than Human? MMBench: Is Your Multi-modal Model an All-around Player?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.025517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.025517Z digest=sha256:3a8b5191cd49ff4eaeb1e75e7e534bdd34c094278f06eef637438edb4f329751

Observation c2780544-60bf-41f8-ae47-dc9314b71fe6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.099423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.099423Z digest=sha256:a058e576987ca96b66fe8b424e25a0b1db2936617005ff0d55ad5e6d8ce8e946

Observation d94c870a-ba9f-4acc-8d8c-bd44a87f2ac6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.194730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.194730Z digest=sha256:44c7e7520baf66867c9570e218fb04edc378710cd42561f3a2dc03cde657dd75

Observation de2927ee-074a-4a77-9fda-d5f316e2214b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.275314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.275314Z digest=sha256:eaac3a8107eaa87228d9a45d6aa111f5c5268376f102cb8c9f8a2aa2ed4a3a6b

Observation 3b4515d8-05fb-458c-aa7c-fd9cf9c1bd68 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.353892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.353892Z digest=sha256:3c10a5e74e0a25105612bed20307ffcca6f024dfe361da2af88dec2271e0e4d4

Observation b38614fb-0d8e-44cd-9e28-7398a53e3386 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.549855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.481908Z digest=sha256:65262944d155665bd6d4cf8760cb2547c8294785f3f5b28948b80328a7b873a0

Observation 6a764844-966c-4a55-adf4-7d89ad29ec5f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.555811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.555811Z digest=sha256:40aa3673f7b071b803634cb0787fcfc150fa9c93e50318c4456151ff2cfe130f

Observation 84ea90b2-4610-4482-9b1e-57435bfd7e6f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.627505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.627505Z digest=sha256:8d829891558ffac89a80cf64b62678ffc38ddc8d27fbe755fe6b071c6f3a780d

Observation ce03bda5-2ee7-4d7d-b199-22debc836638 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge.

MANBench: Is Your Multimodal Model Smarter than Human? ImageNet Large Scale Visual Recognition Challenge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.736985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.736985Z digest=sha256:6722d015d4ecee88c54999b333a4a7e4e59c9e47c05927aa0b38d9c6b57d69d2

Observation aa645b74-9800-47bf-8a4a-9f9697541a3d · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.398739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.811098Z digest=sha256:b1e98351b19d1ea3afec9748ecc22d50da580a605f40a391840956b5b716eb8a

Observation 10dcd35f-3ad8-4473-9d3b-2533e9f33ab9 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.245133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.906468Z digest=sha256:2d0aee071035fd54d22ec78f89e73a3d7acbaa2d0ae1b63092c628955c2469b5

Observation 91cdb192-d7c1-4429-8c88-834a38467284 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MANBench: Is Your Multimodal Model Smarter than Human? Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.957896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.957896Z digest=sha256:b3fe8c7fe8206f95a6703a7de24a54a174366db7b620088d180fdce7d6e16597

Observation 178da340-62c3-46c5-b605-1794e37e26ea · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.077215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:35.053070Z digest=sha256:f9ff4c215cab8c7dbd7b29467cf5625cdf0409aad2072d53efa560cef0203de2

Observation 1ff12e9b-7eac-4590-aa1a-a924fec77b41 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.122944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.122944Z digest=sha256:cfc9db1336eccfcba01827bbdf833af61af0276e680288facb93c921bfe16445

Observation b143a180-e90c-4d7c-8a66-0230a099cedf · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.321699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.321699Z digest=sha256:9b1ec2aae7f562ec140e7324367e1818bf84fcdb165d824411eceb965214f3bc

Observation cbe520ed-f158-447a-9e94-ed8fce53cbd6 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

MANBench: Is Your Multimodal Model Smarter than Human? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.401166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.401166Z digest=sha256:aa03876a4b947bf79b2010b6b418ba87b9b90e28508a4d78c01f4e8bfbdb23aa

Observation 0a02c4e2-b065-46fe-8300-7b4fba8ddcc1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MANBench: Is Your Multimodal Model Smarter than Human? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.585470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.585470Z digest=sha256:c0d8feee66af854a0f26f9d73ada76a4cc3e3e092a2f91f367a4a76863ce1ce9

Observation 9697a96f-5196-47ad-9dad-7c8174cf9069 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MANBench: Is Your Multimodal Model Smarter than Human? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.660401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.660401Z digest=sha256:aa9f66242c61ea6c918b4a5e993cf44347c6bb53d86586ad539583d7a1360b3d

Observation 6f69844b-b07b-468b-9e95-fa2590d3990c · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MANBench: Is Your Multimodal Model Smarter than Human? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.770013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.770013Z digest=sha256:611a104f2e6052597214dd49339e5e8d35d234cd50b012a46ede3e6f8f344a07

Observation 85842938-03b8-4905-9646-413aee53cda0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.843690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.843690Z digest=sha256:618aea53d395e924494506397e9363f0980a857cabd7395d4a65187936fa3c98

Observation bc046f9a-de8e-4942-814f-7be1158cc679 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.920353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.920353Z digest=sha256:3a6adc385eff08ffa1d64a790ba80d21915eb972d1c89474058c9ce326d5e884

Observation 9a2fe13c-2714-4032-b3f4-6bdb30c1446e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

MANBench: Is Your Multimodal Model Smarter than Human? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.030597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.030597Z digest=sha256:d87278bfac6a4db112b7adfb97fc3e8d25fefebc0b7fd297484e49ae6c75f9a0

Observation 57dceeea-435a-43b1-87b3-602d38865fa7 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:36.917406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:04:36.122616Z digest=sha256:0ea359552cfd118ad0f48bca2ecdb1336efdf7ecf700407bdf6cec4ec5831c1f

Observation 94a18d1e-2c73-4c44-8f0f-5ff4448a9ece · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.266807Z digest=sha256:2bd91ab85caeffc7dea99bdc4292ee23fedca4d75e6a6d97956a50e66429157d

Pith citing papers

No inbound Pith citation observations are available.