Pith. sign in

Paper Citation Record · LEDGER

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.12209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.12209 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:12:38.725841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T20:13:24.799209Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4fb8dfe-28f1-432b-9c51-6b8ead269c3b · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.804839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:52f5e70ca09a053754ad30bd5d360c80c6b26de9d9b0467d2ab4846207726d28

Observation d02a36cf-61f9-4022-be44-d6e2618d4d31 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.405426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:0068fbaee42f26946aea9835e10d2b0044a048d96de66c2fcca894bbc8665372

Observation a4ac3de0-180e-4084-8b91-8ad34a3a22a8 · inbound

Computational Experiments in Number Theory cites this paper.

Computational Experiments in Number Theory MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:00.891133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:28:21.017594Z digest=sha256:82f6878abc7b57ab96babad4773b7a785dabc09e110e1e1b5684eb2b808de466

Observation 2593c843-c04f-47c9-9431-139f71010cd0 · inbound

The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants cites this paper.

The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:38.725841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:38.725841Z digest=sha256:1f260ed612a703942b96e13d02bc8f563b4531a6367bd5ea745af97803659a79

Observation 74b9d6e3-901c-4cb0-827e-7c7c600334f7 · inbound

Evaluation of LLMs for mathematical problem solving cites this paper.

Evaluation of LLMs for mathematical problem solving MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:06.681146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:06.681146Z digest=sha256:5719d14f196c746a308a2956efee5ee1ca0363b5703181097469c5de7489c82e

Observation f9ff443f-a8d3-498e-b19a-149234f32a6a · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:18.377180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:18.377180Z digest=sha256:90cdd9dff6e90a2ea554f665733486a1251f5b9fbf09a86f92c3822a18b3df6c

Observation 193dd8cc-9d80-4f51-8211-c7db35f3a691 · inbound

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models cites this paper.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.838439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.838439Z digest=sha256:f89bde15c3f3b87eb91f12150518665f9b682e0c478be011263820c0ff6cd927

Observation 2549b930-ffc8-4b3d-934c-a762e4745d65 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.067924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.067924Z digest=sha256:2b1ebe4419e232245101b391f6618da9d5804eff65a127b3341ba82bcf5fa1dd

Observation b3be25f3-592f-4891-8055-e6dcb953c31e · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.235673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.235673Z digest=sha256:2330a2bf3aa7774cf11e9708bb41996034103a6207722464771abe4edfa97e82

Observation e8e54732-adc5-433c-a87a-d9e101bb78f9 · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:23.844792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:23.844792Z digest=sha256:7c2dc9f1574b83e09e356f853dacd1aebdbbdf11547a66d4ec1bee922ce9f6f2

Observation 2701aed0-f1ed-42b5-a67d-fe4931bc0d23 · inbound

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy cites this paper.

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:12.995111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:53:13.427364Z digest=sha256:5b82113403998e38d161fc0a7885c73d591e639c56d8738642bf55dd413b7adc

Observation e6a6050f-5bc0-4367-83fa-007eea97f816 · inbound

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy cites this paper.

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:53:53.592710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T00:50:11.410735Z digest=sha256:305bb12be1f655cf985675bbc42ad6b74c030b1d1a976429a7b0146ce337dc05

Observation d1a36c05-5cba-460a-9733-094f2a14c761 · inbound

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models cites this paper.

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:12.831388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:53:18.670312Z digest=sha256:cb94bcb68869f44eba42bad487b781ea5beae0ac4801eb12070b82715b6f2ae6

Observation 05f4609b-a3dc-4fc3-aae1-217c8bbbd5bc · inbound

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models cites this paper.

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:33:52.825418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T00:29:19.064290Z digest=sha256:aee95b3686a40e800024951ed8c12525aacbe953bd8c99f6739d3608bf58f652