Pith. sign in

Paper Citation Record · LEDGER

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 28 inbound Pith citation observations for arXiv:2502.09696.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09696 v3

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:55:05.990141Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:12:18.449858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.640745Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98444cc0-8a68-4581-9a43-5f5448bd715e · outbound

This paper cites Pixtral 12B.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.814584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.814584Z digest=sha256:af748db36d2369fb54c68a07bec2d6a82f40342a0a57b8a8dcad299fadf311a9

Observation c8d1fe0e-2467-4eee-b831-d50c64ba0559 · outbound

This paper cites Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.825741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.825741Z digest=sha256:5c93af0a082dc55148727fbf873fc3ebc9d6f4b88fad673a08de99f01a6da1a7

Observation 1536fbc9-d5a1-4d3d-bd9a-969cd4e1f9d7 · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.830828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.830828Z digest=sha256:545cad84f41b6ded48aa802f2f1610ff8dcdfb668b694e0b69a8f783c5022867

Observation b6a508fc-13af-4426-b05b-96e52b647042 · outbound

This paper cites On the Measure of Intelligence.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models On the Measure of Intelligence

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.836265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.836265Z digest=sha256:612e822dab3edf8a1b4bdc17313e29fec0d4c6ea4fa599c8cdd17fa17fa7567f

Observation 8f40461b-8772-4bd8-bff0-5f1e3331a42c · outbound

This paper cites ARC Prize 2024: Technical Report.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models ARC Prize 2024: Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.841893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.841893Z digest=sha256:a7c2ca7d5cfbf2ee2b588943cce6c3a232d0985fab241f87bdef3be1a9f38fe0

Observation 767a398d-e1af-4f8b-a53c-584161bc72c7 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models NVLM: Open Frontier-Class Multimodal LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.847019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.847019Z digest=sha256:804c531048fbcdf056215230146ec519e6afbca8e1793835f01273c14bdf0a3b

Observation 99fdb64f-d6ee-4be9-a54d-e96b954eae30 · outbound

This paper cites The Llama 3 Herd of Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.852088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.852088Z digest=sha256:e90ac4777c51e3f4cb3fe88ef23634b9cec2e8ae04281ffe6d76e558781fc96b

Observation 82ae0c76-dba4-46b9-8e4a-787f01ddfc51 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.857068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.857068Z digest=sha256:0b73b408eceb69967a636411df4f21e584d48deb050ae9f2f2fb9640e6151fc6

Observation 87ae7121-fab0-4754-8a54-8f688185c808 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.871183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.871183Z digest=sha256:ef6272a2c5ba7da7a7f36d5d6e0b65f1be7c7ed0855dbba2cb8dc31368b87acb

Observation 98842e2d-20fb-41ae-86cc-e0076730f568 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.875661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.875661Z digest=sha256:a4e601d136ee2856da66b503d408b4a5cc55ce15f7e2747ac304ea84bcec6b4b

Observation 2700c133-03ac-4b66-9c79-01c30d5e6daf · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.885072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.885072Z digest=sha256:d691f08a6c59ec62f7af7b91d7a7df61881c151a9a4cd8f9ed4a43e7f7945725

Observation 3d61ac6b-2eba-4b79-ac24-e0bce58fc303 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.890289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.890289Z digest=sha256:6f238f3081e0c8829b0229be076d5c5745d174d3ad219227dfe788225baf04a1

Observation 62238fca-fcb3-4dbc-b782-a8949f7f8a78 · outbound

This paper cites Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.895061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.895061Z digest=sha256:0c48b797c3e385c9e05a4b76f0b480880d2f5bc6e256a6d6c9516759f8f5860f

Observation d3f7a737-f74a-40f8-9644-b349ecab8e7f · outbound

This paper cites Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.905698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.905698Z digest=sha256:8d96ea6523987af4ec7dbc390c86e922336d0e8b3649808d4017338d1b58b747

Observation 9210a240-a627-4828-81f3-7bdd2adda71f · outbound

This paper cites Humanity's Last Exam.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Humanity's Last Exam

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.910610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.910610Z digest=sha256:90ce21ca691ad1789ca268d9d994d6e8eed5c32b9e49c9dc23b51e914e660cc6

Observation c9d65d81-e54e-430e-9d2d-426f45546a9b · outbound

This paper cites Does Spatial Cognition Emerge in Frontier Models?.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Does Spatial Cognition Emerge in Frontier Models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.915677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.915677Z digest=sha256:f23637b6555ed44b5dbe22c2b4204f841c116856b5645e367412a9996f6ddfc5

Observation 101bd3a9-40fa-4a20-b044-78c85ba72525 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.920807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.920807Z digest=sha256:5853a29e7c00a9d862911d55a43d589532d4c021a2608fa99903f01adac2d3ab

Observation e9e3135e-90b7-4461-ad0c-40efa29d97a5 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.926142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.926142Z digest=sha256:4a5750dc60c4224555a070bad5f451008416f479694bb0334c176393a223b8c0

Observation 72595cfc-cc34-4337-9f94-be24665ba15d · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.932609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.932609Z digest=sha256:ad084da867aa97459734493df4debd14653bc5a0a4e6e75e869faad96a023697

Observation 4f8f6fce-5fe8-4b96-b52e-dc6f83720153 · outbound

This paper cites GRAB: A Challeng- ing GRaph Analysis Benchmark for Large Multimodal Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GRAB: A Challeng- ing GRaph Analysis Benchmark for Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.938073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.938073Z digest=sha256:65a5289d7e3465dae9fd38367480b9dd0b8bcf2c9d9e4db577a4c0521fa9ca71

Observation 516db7e5-ea15-4bea-b862-da5dfe2b613d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.942731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.942731Z digest=sha256:0f2bc13a306d959ef1469efebb804dd6896af8dbdc0dd2de7f3f6589faa4ee4f

Observation 0effa166-3eca-4f2c-b395-f6d51264615f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.948087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.948087Z digest=sha256:fe77fdfcf6139987dbc999029f8dde0a4ab8353768a1930a33fa8bf3cd14c9c9

Observation 7e23e631-42cf-44a5-9e20-f6532ec450be · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.953189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.953189Z digest=sha256:a0f364901389aace2fc22182724e958c284f6c24607f935782820304dddc6198

Observation 07b145c7-c349-4126-9bcd-10a7ae592afe · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.958169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.958169Z digest=sha256:9c15d6d058c60fe2fc39ec0566c5b455e9f72f8e92a2c00013d8dc57802c6e04

Observation 3b19d9eb-6c8a-4aa5-88c9-bd1656193ae2 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.971251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.971251Z digest=sha256:c29093bcdeae31bb7c4e8b83750dac2aade6fbec34c2c9236a40dc0d9475175c

Observation 2d9b4615-be23-4171-8b8a-227386e1cbac · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.975587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.975587Z digest=sha256:25150b9532dc65fc1a12e56b625de8f8daefba0d37667db202e778c895f24e12

Observation b4757463-6499-4e27-a1ea-883ca2a50071 · outbound

This paper cites A Benchmark for Compositional Visual Reasoning.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models A Benchmark for Compositional Visual Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.979890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.979890Z digest=sha256:0c161cb7efb18cd66260db211190fbcd9ea8030834153d68dde03c9f79a9d202

Observation 781a1daf-41e5-4c0d-a0d0-9b7e712c6709 · outbound

This paper cites HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.985112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.985112Z digest=sha256:972ca640fb7ec426ac9f116d2c7f26669d758a95426c12a4005305db0e89be16

Observation 27902e1d-21ee-4f6e-ab39-e9afcad13154 · outbound

This paper cites Note, o1 pro was accessed through the ChatGPT interface preventing hyperparameter configuration.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Note, o1 pro was accessed through the ChatGPT interface preventing hyperparameter configuration

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:55:06.672110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:55:05.990141Z digest=sha256:cb6c737aecc267e9da80a3a3b40ede9b4d86eee7ed9490386525362e15377f37

Observation f9ad8685-9bee-4776-b457-bda53e46d68d · outbound

This paper cites Scaling Scaling Laws with Board Games.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Scaling Scaling Laws with Board Games

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.861532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.861532Z digest=sha256:f3726958076083ade3c3254ac9daa4e562d3b7feae2414703f74e92b197dc42f

Observation fd7ad324-428f-48a6-baaf-1f530e6f3ee4 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.967050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.967050Z digest=sha256:7aec4c5659f0f63049727a9b12647ca48486d6b31899ce47b271b7d6561a11f0

Observation b43499bb-492b-4552-8110-d242ea5d55db · outbound

This paper cites Hello gpt-4o | openai.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Hello gpt-4o | openai

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:55:06.712697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:55:05.899812Z digest=sha256:cfab68aa4e603a7ff578b8da6353a5dc414e3584e8370f236fbb87bcfd152ea3

Observation b1173d92-1231-48ed-a417-87fc67c79050 · outbound

This paper cites Transformers: State-of- the-Art Natural Language Processing.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Transformers: State-of- the-Art Natural Language Processing

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:55:06.688633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:55:05.962681Z digest=sha256:9adb78267245eaecd6f60e2692e9799da3a8484033fec07181eb05718bbf9493

Observation a2877bd8-df52-4cf0-ad51-32bb199e8e0e · outbound

This paper cites ReMI: A Dataset for Reasoning with Multiple Images.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models ReMI: A Dataset for Reasoning with Multiple Images

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.866702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.866702Z digest=sha256:69c0cfc6eba29ce3d1b23b4b10b54da90638b83cf9d10e58563915219cd96f1d

Observation 2cc69712-5543-4614-923e-709f5745dddf · outbound

This paper cites Mistral AI API (0.0.2).

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Mistral AI API (0.0.2)

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:55:06.732765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:55:05.820471Z digest=sha256:6da3624ddffffdba70fea57c6ae257183f66897928a559a65a7b5417e0973733

Observation 344fff01-4e55-47a8-a7fc-0618645bdd97 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.879967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.879967Z digest=sha256:aa423ab95b0f598d968f2990c2cbec6f60fdb0a4f89259ac2e623847619a2bd4

Pith citing papers

Observation d3d2f3e0-0ea2-4655-82f3-ded6840c9ba8 · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:f22f218b9c8f2524f35cc4f926d8844c512ca17bab014c151a4f3120655bd692

Observation 6328a92d-895e-4871-86f7-8f7f54284b80 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.449858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.449858Z digest=sha256:117a7ec796c79383490c95425d4b5dc5349ceb181974b228861b3f89f2e3d18b

Observation 34766d48-582d-4b35-a054-ef7c618cf600 · inbound

SCAN: Structured Capability Assessment and Navigation for LLMs cites this paper.

SCAN: Structured Capability Assessment and Navigation for LLMs ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T16:11:36.613334Z digest=sha256:be63032a047ca1f6cba5d4d1bfdcf7c403e39ffc8f4e9baf0fc12b433ec72c54

Observation e3f32080-9df0-49f2-8ed2-11638ecb1d4e · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:58d7895f0b61deb094cd3c2d11b9ba43ac16cef0ea70e0f731a817df74f893b0

Observation 85185ae9-fdde-4a62-abc9-5b7fa633c212 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:19.618321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:19.618321Z digest=sha256:d48e867d232c390ca3be363485d02288032c3ece5b981b4a24c45df03c5399df

Observation 72ce693f-81b5-4041-8af2-6027aa1b2613 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:8e9a31b98c3086dab8a1bdc4e8ab87cdccf1e600c7d0aba7700cb9e139c12fef

Observation f374187d-769d-4084-8ca9-cceac74b5a30 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.961173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.961173Z digest=sha256:06c848a1ffdbf532ceb276bb99a9697293ed911f4873afc3f0ec5b4e197dec6b

Observation 1e5d77b1-ae26-4692-935b-c591748238e0 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:4daca7846770b2ff38190cb5d100f5c0c91b833ccad60c28eeca5012cadbed1f

Observation 9bc0b383-70dc-41d8-b6f7-78728df2e511 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.784349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.784349Z digest=sha256:48ad3c4668ae2ab10b46de9fab7832d3983581e293743e0a45013b0d3cfe7d97

Observation 13a4d9b5-34ab-4d86-b96c-42772aa0706f · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.182788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.182788Z digest=sha256:f2dba4daedba85d717ec7e5718fb1bdea656bb76a903428515fa7fdeb1490b27

Observation f2747362-ce7e-4f6a-8c54-10ba720e97b6 · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.885010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.885010Z digest=sha256:01a99040de617c7915e4273e2b1c41dbc890b545bb06e3edbd77e63d76c81fd7

Observation f6b1e8b0-932f-4649-a42a-c00b089b2f09 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:204629bb573f7bb1d796faffe165cb94d3e20c001320f47b1b42127f217e1b8d

Observation fbd72c36-3362-4fad-a307-7bb6779e68a9 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:01863a466cf3a03751a8001cc2b3c872369e5b93eb364c927a0c0f2c256b78c4

Observation 7eed3b2b-8d9c-4ee8-9ac6-a42438f9b796 · inbound

Self-Distilled RLVR cites this paper.

Self-Distilled RLVR ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T19:43:46.623267Z digest=sha256:232bc5f9a2bb3f98b2cf72cbf280592d8bc89f66e25d376d46ed9e1a6859461d

Observation e2b0a151-a2cf-4c5c-9912-ebf4ada0572b · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:e063317cc65766b2924614af2497203ad52ee9d3fb6f7d95e246fec6bf32a51a

Observation 550446f8-6f7c-4807-91e4-2dbc659f2889 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:5cdfcbd5d5085fc6e238c841e0f6ca8b9a3dd648d40f32960f27626ccd3be7d9

Observation fb8cab9e-44fe-49df-b047-586e83a01698 · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:036f8c1b65d2422deb6ab04aa04da20b4df50a407ca62b88558b6400b49347d3

Observation 3d0e8800-e10b-491f-b257-ef0e97432a72 · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:11bf80411a3ea8aaadd5e5c939ac0313a9116bac7e6a53c7c09d32c67d91ae95

Observation 59bade78-b347-4887-87ec-819ddbbc3e3d · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:e233d55f63f8616079d4fe56c94ed4d4a288ca48a75208e7c1211f980a17e632

Observation 4bbfcb8f-eadb-4b75-a59b-a251d94b360b · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:2c7e4a2ba3aec6c1542877e39e30233d349616430dfd2819262477a4d71a2746

Observation 263ef561-e9fc-44e6-819e-76f2f5c527ef · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:29cfd0161422df405b6976e557f71fca265ab5fba33f2eb1bcdb2fe11604ed69

Observation 07df2b79-d2d4-4633-b939-2629537ab74e · inbound

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark cites this paper.

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 144

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T03:04:31.607329Z digest=sha256:8b14ea3f3e3633d10f9a43a802271fe2bd52d1aeda294bdb052d0a22683189fe

Observation a0a912cf-793d-4413-8cd4-8a37758a026f · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4c049c9abf49032dadab7eda9f2f140965584612a320bd10fd5d53111fa123e9

Observation a48ec12d-9f72-41da-94b8-f208e36eb6d8 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:c750c546173e116f0a50c98e14ed5c646a60c22fb8ad249f522eb16dc696e83d

Observation 2540585a-dd82-44aa-9c22-21aa383415c9 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:b07af207b8520165b1e05dfd943f1725fae2004bbfaa610206d3cfcd65266b25

Observation 0f0806fb-8a45-4f6e-80d4-a406a98866c3 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:743ec01d5654e85b3114c1b304af291e07d51a97ee6d33ac3ef5658d691b95e1

Observation 4f6d477b-d9a0-4c0e-9363-d00355a128cb · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:29.035002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:29.035002Z digest=sha256:bd637e7b4a11022f52a0e88edeefd8963caeea2d3104d93bb681118c26fe4350

Observation 4af62350-4877-4c2b-b949-36e3c2447264 · inbound

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition cites this paper.

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:23.411003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:23.411003Z digest=sha256:b673a404bc696858b3c8f423f5229121937869cf1e15f97e768b376f3b6558bd