Pith. sign in

Paper Citation Record · LEDGER

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

As of 19 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 13 inbound Pith citation observations for arXiv:2410.04509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04509 v3

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T20:10:59.264484Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:54:23.129325Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-23T04:32:32.712337Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact58
  • verified fuzzy32
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 391c2e3d-8ab0-4bb8-a7e8-0f1680031792 · outbound

This paper cites Complexity in declarative process models: Metrics and multi-modal assessment of cognitive load.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Complexity in declarative process models: Metrics and multi-modal assessment of cognitive load

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.675921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2b7b566932a518e13bff4d2d952b2e7d1423e0bd9daeb3b0a0d684b3363d19ec

Observation 055423cd-2fb1-4a5f-a3ed-bdbab039e4d1 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.713023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ab8ade7bcea82e8285e198bb11c49281f21be1edb2baf358b5509fb045c05a3f

Observation 081df14e-80ae-4bb1-ab08-30c62b6666a9 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling laws for generative mixed-modal language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.680024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d066cc41657afd4ee0fe62b1668bc0c3e624cd394ff8ee1aae832eae85c36f71

Observation 13237bbd-396e-47ce-8b20-4238944258aa · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.615341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:79f107a5828aa7406978617d1bb9b579659c5d86ef0ad9c160f0a78ca468a93a

Observation ccb84f09-1457-4698-b7fa-8999f8a525ca · outbound

This paper cites Claude 3, 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Claude 3, 2024 a

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.683997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fc052f79e2713bf8df17cb73379e07913756550e340e1dae7ac81131a627da25

Observation 5e6cf1ad-5f67-473c-bd17-555a775c03ce · outbound

This paper cites Claude 3.5, 2024 b.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Claude 3.5, 2024 b

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.804632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b9866a34691476a457881cad804f63e30000e875cc89fef98d91e9e8742856df

Observation df8fda6c-3f85-440e-8684-a2f46e46d525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:25.026755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:06163b42573ab5bade6fca2939de8d42e15f59f5b7d3eb602d4ec0629faf99ef

Observation 8d8b027e-93cc-4830-a1cc-f388a5eda008 · outbound

This paper cites Turning large language models into cognitive models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Turning large language models into cognitive models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.022104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:1bccc1855a0c1a231ffc8eefb5c844d6aef43d58b5421ed90861575b20ec8aed

Observation 4ca4f385-91cd-4b24-95d3-a76458269871 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Theoremqa: A theorem-driven question answering dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.808080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e1642ab43f4831c1797856e206e0e863aff754ac7f78b54a10947edab95c14cd

Observation 3be8ded9-a037-42a7-97eb-24e30023ecc7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.993510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:88e0877d0264e116adb098ea17a984ea754bab360e182586754827e028c0a8e7

Observation ce162e8b-2af0-4b96-8ef1-f74d6e85f1d2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.981889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d3c507285764da4c4b13af5eca4bdb947f647910f6fd9adf30ee3dc3553f2138

Observation 9dc97bf6-0f72-4722-8dcd-ec1e2a400ba1 · outbound

This paper cites A survey on multimodal large language models for autonomous driving.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A survey on multimodal large language models for autonomous driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.793952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:81932ad355a31b1198243b991d9d42368d0a1edd35c49acd13ffb7c9b5162c60

Observation 16026aa9-6449-4604-a689-020543105388 · outbound

This paper cites Advancing mathematics by guiding human intuition with ai.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Advancing mathematics by guiding human intuition with ai

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.801048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:46e51b3c72f4a65d1cc446022382e0c25115b3edec6165728d00cd2f2f2aa7d2

Observation 21518bc5-3b88-484c-822b-15c1805cd1d6 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Visual representations in the human brain are aligned with large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.031867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:3b5cffc466e3327769d5b4d2b68f7293124e8dc8fcef0499e0c20ce0c8dfca3a

Observation 52922b22-c1d3-45ca-bf92-467dfa39452a · outbound

This paper cites Muffin or chihuahua? challenging multimodal large language models with multipanel vqa.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Muffin or chihuahua? challenging multimodal large language models with multipanel vqa

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.790354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:653be0999ede2d418a502259ec37cb3a4c7fe7049178b59270658fbedb560149

Observation c17fe3b3-ce1f-488e-ab6a-d6b5706b8563 · outbound

This paper cites Trends in Integration of Knowledge and Large Language Models: A Survey and Taxonomy of Methods, Benchmarks, and Applications.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Trends in Integration of Knowledge and Large Language Models: A Survey and Taxonomy of Methods, Benchmarks, and Applications

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.742490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:3fda5c6e456b3bae8c299004b5c50f8269b4def0ffe46028e12e7a6f2196c61a

Observation e53d8485-5e44-4821-8c36-a047c854cde0 · outbound

This paper cites IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.891338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:dc2efa84c4829dc2e831bce7014355c3e74c11ffedbc7a9833dd33b3b4cc2556

Observation 33cedaae-dab5-4308-868d-eabed4c4ea68 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.896437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:eec274393fe57d500765e7178d567ca07a6ab19f70c22bc3163e2f1ce318aef5

Observation 61b1e7d4-020b-45a2-a222-07c3f5485eaf · outbound

This paper cites UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.871280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8037bd1e05463f701a3598f9f87f4543b8f17f5a0c809d317f56a530affede61

Observation 868a8b7a-3584-45a6-9cac-c12faeedf4c6 · outbound

This paper cites PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.825358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:1db36e7d9fdcf54bc5b391993e71746409ccacbe96acfa2df7366eee43ea7c3d

Observation afc6ffdf-2f90-4c99-9de8-680fd9e1e13c · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.003350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:13588fa0f9133515c4d4fbf1cf24ff107b494dc2ddc9893e00a198dfe6d081db

Observation 61c2cfe1-22c8-449d-b678-ebbc4c2dbc27 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Measuring Mathematical Problem Solving With the MATH Dataset

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.728474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:24b419fd7cf22249d748b7b473187ce8e1c91648ee507595dcadc688a1fd9e4f

Observation 8900b923-fec1-4134-bf1e-df2e0167bfa0 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.850060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:bb4b5e21095964e27dc1885e6c9e90676658ad2827d0aa5a312873a6d650ed80

Observation ec757e1c-078f-417a-9525-dc588045b480 · outbound

This paper cites MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.854814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:f5988b1cb9730dc401870db14ab7f46e57be20f08f2afdb8ff1202de4abbdbb7

Observation ccc2ac53-320e-476b-a810-444d8c4da6bf · outbound

This paper cites Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.883048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b7108e09fdd48c0c59a61c2ead001242c16b87a3ed138b00ce04484b7176ade8

Observation 1817a52e-bc03-4163-845e-e23b32334315 · outbound

This paper cites New generation deep learning for video object detection: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection New generation deep learning for video object detection: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.797268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:679a5ba219351be6b4b00494747809b9c0f5d6e67798a92dd8432facc6d644e0

Observation 1e079590-4b40-45db-8bf9-78b8a447502e · outbound

This paper cites Learning instance-level representation for large-scale multi-modal pretraining in e-commerce.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Learning instance-level representation for large-scale multi-modal pretraining in e-commerce

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.823018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7239539e4ff5e251f300ca57bc0f3498265bda4298f172fa12dee51ebdec9232

Observation fd89f66e-3a2a-42cd-9720-86d0ad964e32 · outbound

This paper cites Large language models struggle to learn long-tail knowledge.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models struggle to learn long-tail knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.778712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ea1220a747d4dbc818d3de6b068697557eea5bcc68412e8c4dea58c0f0249633

Observation 10420b75-8e73-4209-bad3-e386591f6652 · outbound

This paper cites Scaling Laws for Neural Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling Laws for Neural Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.699095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:295bebd31deb727f53140f1dc6cfc9cf9fa20444d4d062c93c4f025db127b79a

Observation 73b4a142-6d39-40c3-9099-df106ddfdfb6 · outbound

This paper cites Cognitive load theory: An applied reintroduction for special and general educators.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Cognitive load theory: An applied reintroduction for special and general educators

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.710619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b96ea18fdbd710c21d3b25ad66a1b9782530ab8e4c9d6f4892f55108e2d575b2

Observation d3ceb4ea-2a12-403f-92e2-70ad657c6bfc · outbound

This paper cites Large language models are zero-shot reasoners.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models are zero-shot reasoners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.770497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:182e5947171df5a742c267201b224ef91bb59d9c72f3560dec55ab274cd887d4

Observation bc925469-d196-4f80-a92e-8f719e36071c · outbound

This paper cites Solving quantitative reasoning problems with language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Solving quantitative reasoning problems with language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.766527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fc0d896798d7286f80771c6e512c92c3e09cdeee42cd5d205fcbdf342159c2a5

Observation ca615388-f73a-4a0b-9df3-4df1a512c422 · outbound

This paper cites Bringing Generative AI to Adaptive Learning in Education.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Bringing Generative AI to Adaptive Learning in Education

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.068959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2aa6c7f8a6fea09e2353bd5a7393aac4aa090d76a36d1abda8fbc5dbefaf6d84

Observation 14d40029-233f-43c0-99e7-30d1188c3bdf · outbound

This paper cites Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.840001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:aa61dba500beacb617c59a53e678663f9f2ae75abb153252b785a9e2246bd7a8

Observation e4a372f3-2979-437e-8640-101eb7831d61 · outbound

This paper cites CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.901744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2128f1c45955e4fd471f8e26da768f6718fa199fcffdbb162e9e8152ebe259a7

Observation 981eb948-3025-48ee-83e2-0d6677026d28 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.758496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d2f333e64bb6e8a953a091244bb7f7a4dc1ba45b83e0957dc5a38e8f4ec13ecc

Observation a4fb8dfe-28f1-432b-9c51-6b8ead269c3b · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.804839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8caa5f0b88cb3ded7aa4bddec27d9025b35b02a6ac8767b5aa4c21e934c959ff

Observation 213f5ab0-6c1f-454d-809c-372f739c275f · outbound

This paper cites Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.074552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c058ab80ee267ac740e2129d4912c91bd01f3e2406410e4e357fc9f75fe4af46

Observation bb2f1cd8-6e24-43f5-8ac0-7e500414bbac · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.784013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:475eaa05e6e743dcc86a797593aadad9565e713efcef0958536ecfe4bc95b713

Observation 4add98e9-d240-4094-9622-d47fb5d22cda · outbound

This paper cites A Survey of Deep Learning for Mathematical Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Survey of Deep Learning for Mathematical Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.737121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7952815126bb273c9f6abcb8328f333cb8bfc8d7f4fb6ff3d37cb10ed589027f

Observation 740ff5f3-23ef-4f57-877e-1a072ddbc86d · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.722988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:9953e25a8cdb9d1197607b5f0b8630e9051127cd03b82178d9255dbff7e72ef0

Observation 5ae4255b-12a5-483b-a923-2ccbaea78250 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Chameleon: Plug-and-play compositional reasoning with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.762166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7d223642126d7ce6c28ac7c98f94138dff54a8b845eb19a70a83c32b3ba84fb5

Observation a4d2500f-7082-4485-a0e7-69bbf93779de · outbound

This paper cites Large Language Models: A Survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models: A Survey

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.844854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4c16ecc858acd47a6ffdec6272900d3312fdd94fe3fd54c656c6eb0e1381b5aa

Observation 1d180948-b687-4322-b825-7e8fb872c1e7 · outbound

This paper cites Scaling data-constrained language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling data-constrained language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.750446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fef7a4f303e27fb41a72cbdc2d54f412294c54418a20544b2b1a534e41118293

Observation 91aa87ec-9310-4dcc-a70c-71666846f229 · outbound

This paper cites GPT-4 Technical Report.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GPT-4 Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.757361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:279df73800dda0d987f487607623e09e17debc82d165b35737ab8c7db6c6b994

Observation bc700717-d009-44ea-9e7b-272d59800955 · outbound

This paper cites GPT-4V(ision) system card, 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GPT-4V(ision) system card, 2024 a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.743969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:51093bf027f873b3c839151b88fdab2d9dee2b04f6af884417e0daf0ce33c650

Observation bfbc8825-449e-4ddc-9c43-9c4e997b51e8 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2024 b.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gpt-4o mini: advancing cost-efficient intelligence, 2024 b

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.738654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ba2f7c23c6f92b276a670122063d5115907b144e35c4b3eaa1260974d1273cc7

Observation 5d29731a-20f6-4213-8362-89fdc2eb83b0 · outbound

This paper cites Cognitive load theory and instructional design: Recent developments.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Cognitive load theory and instructional design: Recent developments

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.786625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6e706694647928af7a90401151acf79f1ad72d9193d43f41588c4d7af3d87f89

Observation a1e38f9c-de1a-42cd-9fb2-0d325458eb1a · outbound

This paper cites Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.747536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:42623138b511d151c54841edd8f5ee38dff21ceddbea6cb519832d218318d441

Observation 0f08144f-7f53-4bce-967b-e99f9d347ae5 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.789526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:a3b7528e67c86cb86e4f96a479a31a9509f4bcc0226af4da7bbb7f4a589cb323

Observation b92f1d45-4b97-46a8-b689-3c3579cf5d2b · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.631903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:abf570f3ddb3d0c80272a9a43bfea2e69500439eb591bf0c2a67a2a46916e70f

Observation 6109d41c-3620-4bfd-bf38-7ec138d14b65 · outbound

This paper cites Elementary math learning through piaget's cognitive development stages.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Elementary math learning through piaget's cognitive development stages

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.730301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:0c092882398843944450f81b54ca9b7fd5f15b2d9626f3815eaa41c151bad6bb

Observation c5183aab-f5a4-4d10-ac62-c2ca24b4d04e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.597040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:92a3ce6e350ab3e87b907ae03513c012fdf5fa13bea38c7baea483c70ec33aa7

Observation 1cdd9210-316a-4bb2-a8c4-de08309256d0 · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Detecting Pretraining Data from Large Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.831510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6db9912a4754707c3c2b051b04d9799f686ff04832859114e836334ac25209a4

Observation dd81da68-afc6-487a-b30d-16383dfef233 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.675026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:191a75062cba21c9f9365c8549ea57966d77afebbe659c55980c59e6a07075e4

Observation 89ecd939-1a04-4291-8044-b49b18e9cc38 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.058416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:65f5dc0610d8fc28523faf01ca8cd792f0fcf21164cf3a1afe3d539f1a1a5f04

Observation 81b3a9e7-e85a-4bea-b552-7cf98bfc58c5 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scieval: A multi-level large language model evaluation benchmark for scientific research

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.722923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:bf3c27c8ff47e7d24cc30728c31f42fae4f49fc8bd8d88e8e9e759de86f89985

Observation 6696d416-a815-4cb5-bea7-38f0a3071df7 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.796726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e736e62505e60b62c201ccd30e30b63bf8b04b0926405ed723beaba98adb2e9f

Observation d0ae4849-dc13-4b84-8d16-27965b2d6f13 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Memorization without overfitting: Analyzing the training dynamics of large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.734523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e0ed6d97cc5d5f2bcd7d2c67f48d98d4dd360c9839f9a1268c9b8f6841555088

Observation d7ddc412-30ba-4a2d-acbc-506790c88a42 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.769190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4d4f3af59b4a08722eb4f498794f7452838bb943eef80a97f58538076db4285e

Observation 6d1ba5fa-f238-4a5e-aa60-3ad5e5606653 · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models for Education: A Survey and Outlook

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.812679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b88b350a6978fb89799b683305dbd38be45cdf408c8ccdd08d048a7b0d3db3c6

Observation dedad807-91d7-4d41-ad1f-0185068474a2 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.656773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:22b3be23375fcc76f955d72453e473ae3a5ac12a78174818ed49ad39b9b30e81

Observation 2841365e-1eb1-47e2-818b-53502139cdcc · outbound

This paper cites Large-scale multi-modal pre-trained models: A comprehensive survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large-scale multi-modal pre-trained models: A comprehensive survey

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.706479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d7fbbf2b833f0fbfa6acd9074b2968b387737cac37cacfd44b2f29d9b8df93c2

Observation e1bbde25-40b2-4a34-b348-a6b4424acfbe · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.665670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:39d9c52928f93d036d9c4032a1a9226a9fd634acc1953270d5ccee898dec2615

Observation 0b2da07f-e806-4b7e-a552-f13f425ffe47 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.651777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8d9e961255de3eb608c149681dc92b7a8f0fd7daba7d62dd1cab95365680fe75

Observation 2e8d19e2-7d15-4fb1-98f4-a5faad70e5bf · outbound

This paper cites Are deep neural networks adequate behavioral models of human visual perception? Annual Review of Vision Science, 9 0 (1): 0 501--524.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Are deep neural networks adequate behavioral models of human visual perception? Annual Review of Vision Science, 9 0 (1): 0 501--524

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.718670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d32102f33471d4abce7f0ebb2fdc88e9c329bba0c452a918cf4d3ed362db4ca3

Observation b7bf2295-098e-429e-bd4b-71e3631f542d · outbound

This paper cites A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.643282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:68379ec9897e2154ded1c779f1ea0a4ca8929b48e9eeefd926821f74427af7ac

Observation ad01e102-61e5-4ae0-be2e-9cc1545c302f · outbound

This paper cites MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.637635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b0dc583e50ecb9d6639024e73a6f2a5d7056b603cf7c145c081f33fdb034ac1f

Observation 752b09f9-3757-4f31-99a5-b587d6533ca1 · outbound

This paper cites SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.609009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c0612e94fe16ac536563a526bd52c844076221776ccb43ed2a9c75e5cd778b21

Observation 6ed7585e-c8b0-45b2-b324-b47400a2c4e1 · outbound

This paper cites Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.590916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:dc3ffdbebf30020d2ef3bd1e08e017e3b581a462d1aa659b1f05ba9d39b17209

Observation eeb102f7-ccbe-4efe-9ff6-0a01b33774c4 · outbound

This paper cites Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.049022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:0d479d415cc58428830b96f4b460416ed518d38db21df5b526e1b6889ec65ecd

Observation c1a9acb3-51a9-4a54-8acc-a199d0c57ce1 · outbound

This paper cites GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.693768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:1f507083567e3f3ee6d3a3cea429697bbeb2264fe6e3582febd76f0907ea26e1

Observation ae8143ff-b337-457f-ac72-5a7d537990fe · outbound

This paper cites Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.689996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:9017d004284695de139a0aeeef27a09a87a8b6889c30f406c6a9b5c49e5a2661

Observation 0372a539-9a1d-41d3-9673-a0093cb903f5 · outbound

This paper cites Exploring diverse in-context configurations for image captioning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Exploring diverse in-context configurations for image captioning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.694605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:694ddbfe5bf4701226b41a6efef9e32d965be9af147f50cd70e9691da43fa5f1

Observation 17c4ad5f-dc70-4c37-ad5e-3797dcbb1393 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.906671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:07ae31233f2a6190bec8533b2f5b61f7facdadd6c311019c2bd47033d00ae852

Observation 2bab7c3f-75c0-4163-9282-e0ca44687f23 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Yi: Open Foundation Models by 01.AI

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.686872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:be790e5371901ff62248d5615ee38fec2a052364fbb7ec53e82ae44e1379e7a4

Observation 7dd62501-1f2d-42ad-929b-c34c289c1fb3 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language model as attributed training data generator: A tale of diversity and bias

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.699491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4b04aad4fdaf400f3fded9d6bf919ae43fa5fcac98ef4ce74fe1d4d98774bb04

Observation fdf4f964-e45d-4c93-ad69-829831adaa9a · outbound

This paper cites MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.818351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:40162bc6af8b3e2fb37d450ad9f97c3678138f053c3468349864557e15a07727

Observation 1fd7739d-d3c1-447c-9fb2-ef9a29d4363c · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.708257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7198d943c1028983fdd3a7fcab22c004b2043bfb9b17cda4241b51a60a704f19

Observation 60de8a06-b8b0-4a34-a59e-75de3c30507a · outbound

This paper cites A Survey of Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Survey of Large Language Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.620935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5b757313b0afb11e2debbf239d97c36f01ca6a4aed12094eb700b0e10a6abadd

Observation 26e69ea4-d823-4e60-88c0-5da7588b24d4 · outbound

This paper cites Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.681459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2b583aa42ea83ff682b1e87863c9f5ceaf07b593569a6eff9706e27d3c452fbc

Observation d7ccc97d-03ac-4148-a9a0-67fb8b57f831 · outbound

This paper cites UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.703873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:11bb7ef4ff8601e2f33d3e1caa0b7d9bb3f95f7598cf5c779b000918d8bc5fdb

Observation 6a4c28c4-ad13-4c28-8dfe-f0bb8de28cef · outbound

This paper cites Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.864286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6d122f0c6630a71d1fef5aa22dac3602aa3a007279eded639c93d16d6220f6f0

Observation 774c5a7e-bd97-4aff-a820-5a3e8933d4cc · outbound

This paper cites Large Language Model for Participatory Urban Planning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Model for Participatory Urban Planning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.777188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:20fe059cc1dd1171287cae115a9e66f0b56dc5ebe22b4e90e1481c1a6b426578

Observation 6be2948c-6684-4b74-9e22-fbe581ffa754 · outbound

This paper cites Large language models for information retrieval: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models for information retrieval: A survey

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.037969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:665d64363819ab5f4f5562e7f463dbeea9389ceae80f35b667960b977a2a0eac

Observation 882d7d89-01da-45cb-8219-31bb3569df79 · outbound

This paper cites Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.043471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:077505758ac6581906dc61cdaac69034ba069f0dd07c08058125369068c3bf73

Observation eb14ca03-8ff0-4d6f-988c-96131b7302d4 · outbound

This paper cites Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.817974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:360ece1e987e2f3325a63635727c3ede36b02992138452218f562583d4c0635e

Observation 0ce46bd0-3a45-4866-b6a3-476d24ef4787 · outbound

This paper cites Object detection in 20 years: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Object detection in 20 years: A survey

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.829449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c0dfdc386505ab15d8b010dd6c7eda3dbf57ceb69783e925b3a1ed5bbe4aac33

Observation 6892dd8d-ae52-453f-bd0a-45fa91b59f44 · outbound

This paper cites write newline.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection write newline

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.774675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:11f9dc6bdfab6b8fab4a25199cc22fbdb4bacadf39052375473cf4b5b315df11

Observation 2229d2cb-fbcf-4124-8669-856d25a312ec · outbound

This paper cites @esa (Ref.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection @esa (Ref

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.754450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6e20f9d38c56eb043b4871f06f8e99edf715af191d22b01e2c632153ade9286e

Observation b4822c65-1ae0-4888-937b-1572a7a59bd0 · outbound

This paper cites an unresolved cited work.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:13:25.783042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ba27f2e6d977c67dc9b18af1c9d8d3c886e709a17ec843f3fcbc4813b028730a

Observation 0393e8f1-3007-41c8-a9e9-5de496b1dab9 · outbound

This paper cites an unresolved cited work.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:13:25.814025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:536a4a1b1eb0b3c02e0b0a7700e83906ae530f1fc3fd57a6f0908e08660b13fd

Pith citing papers

Observation b878d1b5-5129-4ce4-8ac1-7417cb9acef9 · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.129325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.129325Z digest=sha256:1a93a124b2c12607a0188525f9e95c21f7ca26b92959e16d0cfd2d3b7892623e

Observation caee159e-ef80-446f-ae9d-b3b0e6de9ef3 · inbound

Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions cites this paper.

Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:50.093213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:50.093213Z digest=sha256:3430b3ff0084856c3c3764c8b939075da3f5214730e57849d43e3bc063f5ad2c

Observation 9c605137-5d0b-461f-a0c0-26dfbf061253 · inbound

SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization cites this paper.

SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:33.961647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:33.961647Z digest=sha256:86a505f55f3589068d9870cd916cfd019a25d92a4b5cd55e201a311d6f78b9b4

Observation d9f32bcb-afaf-402e-a709-a0de4b010f27 · inbound

PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models cites this paper.

PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T21:58:03.596038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:58:03.596038Z digest=sha256:c0b05880a4735e5bf65344093485321989ea5934dd0ead170f779ede8432494c

Observation 0eabe05b-56f9-4448-8a95-946e31910d21 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 224

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:32.716530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:9822f333e3f5eac01e6f2c8db02f44833e5e9d29eb9ea6a30ac0e300ab99fc2d

Observation 54fc93e1-c672-4bc6-bd01-f06f3c2cd15f · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.439693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.439693Z digest=sha256:b71dfbd36e3ccc9275478f346105ea1d0190ccc8ee392f1195a7a5ab464830bf

Observation 485367d5-0e3b-41e8-af3e-5c3530cf70f2 · inbound

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring cites this paper.

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:37.844699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:37.844699Z digest=sha256:c85b5fb06fcee263d60083c02da8721a8b5fe906a0b6f34fc6154ba46c20c4c7

Observation f56b0be1-a536-4fe1-a031-b6364e5d2db8 · inbound

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis cites this paper.

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:20.619522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:20.619522Z digest=sha256:21a15c16bf962a1efbaa37f0ae44233cf55b12bc0464601dbfc78c388db728e1

Observation b15cb36b-4ccc-4a3c-bd5c-af9a360c38a5 · inbound

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities cites this paper.

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:23.892898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:43:23.892898Z digest=sha256:fb0eb702069646734e10d3c6912ae5bb050c9bb47dac0c69652d104666a487f3

Observation 8ffce526-ce49-4519-9858-9056ea838602 · inbound

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models cites this paper.

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:39:59.415150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:39:59.415150Z digest=sha256:8e55c6200d1852ec30a242f0b2681b3676e9d4aad9ebbeda952a7916d358a440

Observation 592386d8-cadf-4471-b632-729753b65e9d · inbound

Evaluating and Improving Large Language Models for Competitive Program Generation cites this paper.

Evaluating and Improving Large Language Models for Competitive Program Generation ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:30.637102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:30.637102Z digest=sha256:6b498abaaa574ba95406649cc20cf64878937442eec560de8219b8437ee0685e

Observation 6a1fae2a-88ec-4283-b890-89b596714d48 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:28.516071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:28.516071Z digest=sha256:d8f56ccef0eb8294e57843f056e238f5879fe8859192ce4b010f0248c8038f42

Observation 5d475765-52f1-4439-a241-eff51990eeac · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:53.846243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:53.846243Z digest=sha256:ade5d494ee8d658621ef76c97bfb0725c814bd456ea9d2f4beedd623920642f9