Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:55:05.990141Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 28 inbound Pith citation observations for arXiv:2502.09696.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:55:05.990141Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:12:18.449858Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:49:41.640745Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98444cc0-8a68-4581-9a43-5f5448bd715e · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Pixtral 12B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d1fe0e-2467-4eee-b831-d50c64ba0559 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1536fbc9-d5a1-4d3d-bd9a-969cd4e1f9d7 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a508fc-13af-4426-b05b-96e52b647042 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models On the Measure of Intelligence
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f40461b-8772-4bd8-bff0-5f1e3331a42c · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models ARC Prize 2024: Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767a398d-e1af-4f8b-a53c-584161bc72c7 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models NVLM: Open Frontier-Class Multimodal LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fdb64f-d6ee-4be9-a54d-e96b954eae30 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ae0c76-dba4-46b9-8e4a-787f01ddfc51 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Measuring Massive Multitask Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ae7121-fab0-4754-8a54-8f688185c808 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98842e2d-20fb-41ae-86cc-e0076730f568 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models A Survey on Benchmarks of Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2700c133-03ac-4b66-9c79-01c30d5e6daf · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d61ac6b-2eba-4b79-ac24-e0bce58fc303 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62238fca-fcb3-4dbc-b782-a8949f7f8a78 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f7a737-f74a-40f8-9644-b349ecab8e7f · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9210a240-a627-4828-81f3-7bdd2adda71f · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Humanity's Last Exam
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d65d81-e54e-430e-9d2d-426f45546a9b · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Does Spatial Cognition Emerge in Frontier Models?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 101bd3a9-40fa-4a20-b044-78c85ba72525 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e3135e-90b7-4461-ad0c-40efa29d97a5 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72595cfc-cc34-4337-9f94-be24665ba15d · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f8f6fce-5fe8-4b96-b52e-dc6f83720153 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models GRAB: A Challeng- ing GRaph Analysis Benchmark for Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516db7e5-ea15-4bea-b862-da5dfe2b613d · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0effa166-3eca-4f2c-b395-f6d51264615f · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e23e631-42cf-44a5-9e20-f6532ec450be · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b145c7-c349-4126-9bcd-10a7ae592afe · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b19d9eb-6c8a-4aa5-88c9-bd1656193ae2 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d9b4615-be23-4171-8b8a-227386e1cbac · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4757463-6499-4e27-a1ea-883ca2a50071 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models A Benchmark for Compositional Visual Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781a1daf-41e5-4c0d-a0d0-9b7e712c6709 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27902e1d-21ee-4f6e-ab39-e9afcad13154 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Note, o1 pro was accessed through the ChatGPT interface preventing hyperparameter configuration
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f9ad8685-9bee-4776-b457-bda53e46d68d · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Scaling Scaling Laws with Board Games
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd7ad324-428f-48a6-baaf-1f530e6f3ee4 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43499bb-492b-4552-8110-d242ea5d55db · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Hello gpt-4o | openai
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b1173d92-1231-48ed-a417-87fc67c79050 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Transformers: State-of- the-Art Natural Language Processing
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a2877bd8-df52-4cf0-ad51-32bb199e8e0e · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models ReMI: A Dataset for Reasoning with Multiple Images
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc69712-5543-4614-923e-709f5745dddf · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Mistral AI API (0.0.2)
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 344fff01-4e55-47a8-a7fc-0618645bdd97 · outbound
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d2f3e0-0ea2-4655-82f3-ded6840c9ba8 · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6328a92d-895e-4871-86f7-8f7f54284b80 · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34766d48-582d-4b35-a054-ef7c618cf600 · inbound
SCAN: Structured Capability Assessment and Navigation for LLMs ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3f32080-9df0-49f2-8ed2-11638ecb1d4e · inbound
Seed1.5-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 85185ae9-fdde-4a62-abc9-5b7fa633c212 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 125
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ce693f-81b5-4041-8af2-6027aa1b2613 · inbound
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f374187d-769d-4084-8ca9-cceac74b5a30 · inbound
Kwai Keye-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5d77b1-ae26-4692-935b-c591748238e0 · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9bc0b383-70dc-41d8-b6f7-78728df2e511 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a4d9b5-34ab-4d86-b96c-42772aa0706f · inbound
Kwai Keye-VL 1.5 Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2747362-ce7e-4f6a-8c54-10ba720e97b6 · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b1e8b0-932f-4649-a42a-c00b089b2f09 · inbound
Kimi K2.5: Visual Agentic Intelligence ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fbd72c36-3362-4fad-a307-7bb6779e68a9 · inbound
Seed1.8 Model Card: Towards Generalized Real-World Agency ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7eed3b2b-8d9c-4ee8-9ac6-a42438f9b796 · inbound
Self-Distilled RLVR ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e2b0a151-a2cf-4c5c-9912-ebf4ada0572b · inbound
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 550446f8-6f7c-4807-91e4-2dbc659f2889 · inbound
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb8cab9e-44fe-49df-b047-586e83a01698 · inbound
Qwen3.5-Omni Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3d0e8800-e10b-491f-b257-ef0e97432a72 · inbound
Near-Future Policy Optimization ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 59bade78-b347-4887-87ec-819ddbbc3e3d · inbound
Co-Evolving Policy Distillation ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4bbfcb8f-eadb-4b75-a59b-a251d94b360b · inbound
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 263ef561-e9fc-44e6-819e-76f2f5c527ef · inbound
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 07df2b79-d2d4-4633-b939-2629537ab74e · inbound
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a0a912cf-793d-4413-8cd4-8a37758a026f · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a48ec12d-9f72-41da-94b8-f208e36eb6d8 · inbound
Kwai Keye-VL-2.0 Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2540585a-dd82-44aa-9c22-21aa383415c9 · inbound
MMGist: A Comprehensive Multimodal Benchmark for 2027 ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0f0806fb-8a45-4f6e-80d4-a406a98866c3 · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4f6d477b-d9a0-4c0e-9363-d00355a128cb · inbound
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af62350-4877-4c2b-b949-36e3c2447264 · inbound
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.