Pith. sign in

Paper Citation Record · LEDGER

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.12023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12023 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:10.054216Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T20:13:24.898902Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e4a372f3-2979-437e-8640-101eb7831d61 · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.901744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2cfc91bf01b33f1f4bf0ce868e12858003ee770cc082574c1c8625d5e25a56a1

Observation a025b83d-6424-46d8-9c2d-5465ccc4248b · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 297

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:10.054216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:10.054216Z digest=sha256:05e66cde8be31af7c32100dfff7830dcc55949be2ed2477a448592587d65ddf9

Observation 3a2d1bc4-7646-444a-aad8-f7820aa139d3 · inbound

Large Language Models as Computable Approximations to Solomonoff Induction cites this paper.

Large Language Models as Computable Approximations to Solomonoff Induction CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:54.680465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:16:54.680465Z digest=sha256:fd08bfc4f633270c2fdb2014cb8ebee3da30cca52311d8c7e208dc393982f74b

Observation 3e9f344b-1c98-4941-af2c-1bf31b044908 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:08.364386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:08.364386Z digest=sha256:c517ed8a1e68503ba21032fba6d9a8d4cef61bc5f3cb9ec3375f1565f57bd039

Observation dd46b12c-2414-4c4d-a52c-c89d43e00e35 · inbound

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines cites this paper.

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:56.096946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T00:38:43.897231Z digest=sha256:6deaa79d9c091b133ca9a9a71ba40bff1d700d90fdd1af0d88dfc0c52b86dd33