Pith. sign in

Paper Citation Record · LEDGER

MANBench: Is Your Multimodal Model Smarter than Human?

As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2506.11080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11080 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:36.266807Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cbbc84f-f1cd-4c8e-94de-025e13f9b210 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

MANBench: Is Your Multimodal Model Smarter than Human? MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.019433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.019433Z digest=sha256:a90e9d27f3967d9f3d9f098c341c5060cef0241f7723e3c2dc7a44b141854acc

Observation 6fc41757-0bb0-4ca5-9a50-5ec21542d75b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.100287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.100287Z digest=sha256:4403c9a1f74284e27da7eafff5ade04935a3a494c5a6632c80bb8b56d4be930f

Observation c4c0966f-75f9-43c1-94ef-c43269d723e0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MANBench: Is Your Multimodal Model Smarter than Human? Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.203398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.203398Z digest=sha256:48e7c4b2047da4c69135a502afada8f0f01eaab6c23ea6568c96a61d4ee2e76b

Observation faf179a8-d2c9-4552-931e-acdd0bd8466a · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

MANBench: Is Your Multimodal Model Smarter than Human? GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.269754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.269754Z digest=sha256:5e5b250324055319a21e74ed9ca8f5d55d2c98347382a214ff1cf481fe36e24b

Observation e79d540c-5bd5-4e1e-b525-ad946d81fda6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:38.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.342134Z digest=sha256:511bf7aa433ebed40d88e068c38b1e03da5e38a6ea3ce14158f13533e7ac7894

Observation 46b97ec5-2d1c-424f-aaac-7a0c7fe1c17b · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.922036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.413796Z digest=sha256:8f9603ea2ab2ffe78d109647e0e5b138b34f1e24fc1ef87a798ea15e87c3407b

Observation 45ff32d9-7ea2-4adc-885d-acb2aed40042 · outbound

This paper cites The Llama 3 Herd of Models.

MANBench: Is Your Multimodal Model Smarter than Human? The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.473651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.473651Z digest=sha256:532d7ab40bdb0228887edb94e97992636efeb1ec7a9a82e5618e4968b0673641

Observation 6840f7bc-b91c-4ab6-8362-746e39cd4238 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MANBench: Is Your Multimodal Model Smarter than Human? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.576525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.576525Z digest=sha256:780a33d5763520f7922ce99fb5470a606b584fea6d0dbaec4084f9ccfe44704c

Observation a7295d2f-a5da-461c-bb3a-9925fd622e13 · outbound

This paper cites OpenAI o1 System Card.

MANBench: Is Your Multimodal Model Smarter than Human? OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.681836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.681836Z digest=sha256:f963943456193327df7fec3d13fdc8e8bf5a7f51d48e87798fc476031875bfbe

Observation 79938242-1aaa-48e2-8de3-877f771a804b · outbound

This paper cites DeepSeek-V3 Technical Report.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-V3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.753861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.753861Z digest=sha256:e64148530befbaaf42784f62275dc95211c7eed88b454eff4078783382a6e88f

Observation a1dfed1f-bfa6-46e5-a649-b08ab28235d8 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.852870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.852870Z digest=sha256:44dedf7715ce66c6d8ba9b887b657dae046be75017d12777db28712fd598a8e8

Observation f3cceeb6-bdc8-40de-8901-36b3e3619c04 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:33.952899Z digest=sha256:39af60689ad0ea1514823404d6cf9332481ea42f8e9d1956bf13efc90c2f7807

Observation ced02137-89da-4243-82e2-44d681229640 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MANBench: Is Your Multimodal Model Smarter than Human? MMBench: Is Your Multi-modal Model an All-around Player?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.025517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.025517Z digest=sha256:d780e5c4f62beae7f05f707ef78dc5fb8d874564f09f2f544c6721802638fd93

Observation c2780544-60bf-41f8-ae47-dc9314b71fe6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.099423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.099423Z digest=sha256:ebf7514569f018dc28d520f122c89b608d2d594a7a640625e4f4d0b2775c42bf

Observation d94c870a-ba9f-4acc-8d8c-bd44a87f2ac6 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.194730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.194730Z digest=sha256:2ce42a076c368454898dfbf1468303d441ea3d0537e4e2821d761f8525ac5f98

Observation de2927ee-074a-4a77-9fda-d5f316e2214b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.275314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.275314Z digest=sha256:1ef4797fc53e082ed370caa335f3161ae95da6ba1a7861d142f960b66a3cff47

Observation 3b4515d8-05fb-458c-aa7c-fd9cf9c1bd68 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.353892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.353892Z digest=sha256:df57a54724b89bccfc8619b553ee78c0e0f9d06c73773e5442fdfdd6a74fff7b

Observation b38614fb-0d8e-44cd-9e28-7398a53e3386 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.549855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.481908Z digest=sha256:68f67fb85f1873da4b17c42a997612931c4c7573f75f9f71cddad1fa7dc3b99f

Observation 6a764844-966c-4a55-adf4-7d89ad29ec5f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.555811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.555811Z digest=sha256:24946a923af9006cfe767992a2cde55a6ad8262f4a1165f4be01caa97b57de2a

Observation 84ea90b2-4610-4482-9b1e-57435bfd7e6f · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.627505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.627505Z digest=sha256:3756aabd19381998dbb591658bcdbe48d9b04b1aa9e83c75b278b9f6112ad2ed

Observation ce03bda5-2ee7-4d7d-b199-22debc836638 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge.

MANBench: Is Your Multimodal Model Smarter than Human? ImageNet Large Scale Visual Recognition Challenge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.736985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.736985Z digest=sha256:785626083e49f3c6724ac3cd476cc7d1dff50773649dd906981bd1eb906ed834

Observation aa645b74-9800-47bf-8a4a-9f9697541a3d · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.398739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.811098Z digest=sha256:d5de091776c6017c5bed6ff3b65f4e082ccec63547bc09d1f96160150253b675

Observation 10dcd35f-3ad8-4473-9d3b-2533e9f33ab9 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.245133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:34.906468Z digest=sha256:6aa4a30b269381f6f5ffa42d8b43f3dd4692cccfaf6fcd3b2e75967fab9b68f6

Observation 91cdb192-d7c1-4429-8c88-834a38467284 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MANBench: Is Your Multimodal Model Smarter than Human? Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:34.957896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:34.957896Z digest=sha256:6c2147de2dee4d5944bb47f81764d4d0162852ab32402fc4002d78c24d3f4ab1

Observation 178da340-62c3-46c5-b605-1794e37e26ea · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:37.077215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:35.053070Z digest=sha256:b86c359c85625206a913b3cf6f47e56ab586aaf0cadfc024da4d893eb18f5a2d

Observation 1ff12e9b-7eac-4590-aa1a-a924fec77b41 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.122944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.122944Z digest=sha256:39eebf44b7282c372189d7e6cc24e51a04d282cffb9450df40d5fe1398dd281b

Observation b143a180-e90c-4d7c-8a66-0230a099cedf · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.321699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.321699Z digest=sha256:4e1c14845dadeac4beaa7b2501c1237e004332d43e6f8db1e6a6b89534e089af

Observation cbe520ed-f158-447a-9e94-ed8fce53cbd6 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

MANBench: Is Your Multimodal Model Smarter than Human? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.401166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.401166Z digest=sha256:aca892482339240aa9a0768b270febb7bfbf50c5669a73d781fd3cf10a44735c

Observation 0a02c4e2-b065-46fe-8300-7b4fba8ddcc1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MANBench: Is Your Multimodal Model Smarter than Human? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.585470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.585470Z digest=sha256:b77f7f2f258068bbd1c85acc15864827b7860f430c2689df2901b0cd439f0378

Observation 9697a96f-5196-47ad-9dad-7c8174cf9069 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MANBench: Is Your Multimodal Model Smarter than Human? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.660401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.660401Z digest=sha256:c12e93c16152722960a07ebb23fbdc9e63a741bb70c8327bb912f1dbc8507429

Observation 6f69844b-b07b-468b-9e95-fa2590d3990c · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MANBench: Is Your Multimodal Model Smarter than Human? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.770013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.770013Z digest=sha256:d28e778460d7dc601726281e96cadc3b949ca7c6c173a05a06ca98fc94bae856

Observation 85842938-03b8-4905-9646-413aee53cda0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MANBench: Is Your Multimodal Model Smarter than Human? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.843690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.843690Z digest=sha256:8738455a4b9ec758890ace0ecbcc087b6a539b545fb3438b3e1befd8e62f9597

Observation bc046f9a-de8e-4942-814f-7be1158cc679 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:35.920353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:35.920353Z digest=sha256:cd802f6ae65ad8d881299491e182dfd89e0550d1cb9f8cdb109525060b1ce60d

Observation 9a2fe13c-2714-4032-b3f4-6bdb30c1446e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

MANBench: Is Your Multimodal Model Smarter than Human? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.030597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.030597Z digest=sha256:b8e530bd0dca0fc7dba5b5259ba1539d045e3971d993e1aac15d040d51d9227c

Observation 57dceeea-435a-43b1-87b3-602d38865fa7 · outbound

This paper cites an unresolved cited work.

MANBench: Is Your Multimodal Model Smarter than Human? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:36.917406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:04:36.122616Z digest=sha256:e7807a2f3abd6ca4e16e25c8e9e4f224eeee32b6b776dc9c0f5f88ba86b06306

Observation 94a18d1e-2c73-4c44-8f0f-5ff4448a9ece · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MANBench: Is Your Multimodal Model Smarter than Human? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.266807Z digest=sha256:aab90d38c3ac049e0ae4da2fa412989f9ad314161fce45eaa8c9e84e68ea0621

Pith citing papers

No inbound Pith citation observations are available.