Pith. sign in

Paper Citation Record · LEDGER

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

As of 19 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 5 inbound Pith citation observations for arXiv:2501.13766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13766 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:40:39.950474Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:34:14.534071Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:25.670647Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 112c66d7-04a1-4d1b-a21a-e7d108d90a83 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.544581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.544581Z digest=sha256:aa0c90700f9d05cb64ed34961e788bfaa1240c0f5ed583c3fe45ddde81c992e5

Observation 53e3c48f-fd35-40f4-b26c-8fc824980633 · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.550749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.550749Z digest=sha256:fe07eca803989431b8016ecde9a047ffa8c423c51e9c7b25378f20c366160c15

Observation c91449ec-0e50-4a9c-b3b5-43239d8bd634 · outbound

This paper cites Llama 3 model card.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.556201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.556201Z digest=sha256:e14d0ce1b32f0825b40717332671bfd332d6b4c4420f912aee6750deeb93de5b

Observation 04d744b4-b66d-4209-832e-ef129bbd370c · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.561212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.561212Z digest=sha256:7987107269fdd26dfcd35c86da0a56b9fb7087e067a4d7aa0721214cad613381

Observation cedec0c5-2d90-4b56-97d7-d3461e4a3b75 · outbound

This paper cites Claude 3 family.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Claude 3 family

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.307721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.566715Z digest=sha256:7bc34a7f5062a7b278113b52710279ba08a8a985a197dac3790985c810c1bfdf

Observation 83950808-6d63-42db-b4f7-8fd3b7983326 · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Llemma: An Open Language Model For Mathematics

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.572302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.572302Z digest=sha256:c984b499603cb4c364ca50921f9a66e52b67f0d1d5ffd570343b3089318a70be

Observation b537a3de-8faf-455e-836c-f489bb6c4a66 · outbound

This paper cites Numinamath 7b cot.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Numinamath 7b cot

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.577730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.577730Z digest=sha256:764a86176c0e7694e8431c7a569b813c75261af1507e157eced448a2fb1107fa

Observation 8f23a55c-7aeb-46a0-91cc-5793641ba7ad · outbound

This paper cites Natural language input for a computer problem solving system.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Natural language input for a computer problem solving system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.281044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.584189Z digest=sha256:f1d42cccce83dd505816d7b6b7a94e811b3d7491e6a31803fe3dbbd4a5954d79

Observation 0c29df4c-4659-4dda-bfc5-48162fea5bca · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.589484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.589484Z digest=sha256:ec0f508b24c07bb0acfaa8d5f3f1e45fc473e02c422ca1cf02254acd03468c74

Observation 2c239a0b-3c6a-4e05-9ba8-de94667fed2e · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.594235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.594235Z digest=sha256:515e6eb594edb3be877efb860a73c3b21cce93a3b446a04037402ccf37a47b85

Observation b1c5f3d2-255d-4d38-9d90-b5c2442894f2 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.599408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.599408Z digest=sha256:bdf1eb2d224eea026b3234a7c286468af83af531d3db469f3075e37dfb10e9cd

Observation 9ffa50c7-3651-4513-b136-2d80c55f9f43 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Theoremqa: A theorem-driven question answering dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.253873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.604841Z digest=sha256:2f73f4401f06ec583eb1fd904669a716e08c6eab04f7bb616d27c31f9bd0b4d7

Observation 3d0553c2-64d4-40b2-981f-03f6f2cacf01 · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Premise Order Matters in Reasoning with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.609708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.609708Z digest=sha256:3025181f952ecdc590b02d9844abaea65dcecf41b56d198752d05e77799e0d70

Observation 4b5d623a-009c-45bf-98fc-8027e215b48c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.614608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.614608Z digest=sha256:4aa568fe82ecd42d2f7a8783c7180b26fa8e884ab0e4f50510c5aeee33524474

Observation 00d1b62a-186b-4e91-9f97-523cf32dc4fb · outbound

This paper cites Evaluating language models for mathematics through interactions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Evaluating language models for mathematics through interactions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.236541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.619641Z digest=sha256:35b1507c4cd234ecc553a8af5d59bcede4dfadcb7fd5c46721e562f081a4e6ea

Observation b56fc7bb-e869-42d5-8e8c-bac6c4a40447 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.624353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.624353Z digest=sha256:1dc28df9c31f4f7a88be6c42ad7554fdaf8a5c9594fdea6b36d7875527dd71b0

Observation 2ffca2f8-04ee-4faf-80dd-38c1b1d26ceb · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.629739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.629739Z digest=sha256:4a667c2aadbac4c5d274165ab131cb81950e7aba72c166993c9cf288e9dab663

Observation 02230e92-85ab-4b48-af74-053c66d8d60d · outbound

This paper cites Investigating Data Contamination in Modern Benchmarks for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.634533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.634533Z digest=sha256:a9dce2b28f94df3fec49c061194524c7ea8a0945c45d5f4f28b34919d84f0288

Observation c3126efb-dc84-4429-b7ff-800f2f3c8d47 · outbound

This paper cites Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.639354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.639354Z digest=sha256:6e95e08525faf9ec0611aafa04e1242dc840772ebc38ee911afccaa24552cfa1

Observation 077e38fe-1be2-4295-a652-7ca1a967eb0a · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.644278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.644278Z digest=sha256:c3b5fe129250731cb4793401c4c814653f18ec629c9c0663aa6761ca013d94bd

Observation eabd3cea-40ad-42db-a1e3-f89c85a61189 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.649051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.649051Z digest=sha256:05eaa3fd21c548bc6950f625a3b9f4511d10236fff56df4987eb2efe4a6fbcd5

Observation 88227e78-1063-4d19-8980-340995b03504 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.653756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.653756Z digest=sha256:076e10e2ffedf7498315fbaf0e0bc0adfe908c99206b527bfccbe08364eed510

Observation 5880b332-0ed5-4872-b79c-90421489126f · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.658649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.658649Z digest=sha256:30bd3fc92c7f687a195f4853551293640868bd1f6a8143994fc9e035788daf50

Observation f6f556c5-0b3b-4e02-8c07-d2fca9794773 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.663637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.663637Z digest=sha256:1c31cfe195bf9db1d4fa6bff5116c92407f3003a3d59ebda98ae7a47b38e35cd

Observation c65a5993-c6da-4384-8b0e-fe79d8886ba3 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.668864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.668864Z digest=sha256:af2497d4e144230dbacbcfd531d9b87151e2ae0e762c230e7d764856a4407bcc

Observation b3d0c280-304c-49a7-a989-ab3e5fc7b561 · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.673875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.673875Z digest=sha256:3a7ea5c973a7760807cb05334d12f7b12727d3fd505ed19c9ed7cd1f9b89b9cc

Observation f8781234-6a3c-4196-bc36-f758951902b5 · outbound

This paper cites Mistral 7B.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.678599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.678599Z digest=sha256:9f711b4a15052e0a7e46043d2a905311be225608947876037aca1c9129d14a26

Observation 9db68948-0a87-4acf-9eee-84d5d90e3d3f · outbound

This paper cites Investigating Data Contamination for Pre-training Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Investigating Data Contamination for Pre-training Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.683641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.683641Z digest=sha256:ef9815e16eca1b979f9fa41d04a8adbb28f34dd993e271dda465c4403eeece02

Observation 9078c86e-1a07-4823-8cd8-d942510a60f6 · outbound

This paper cites Large language models are zero-shot reasoners.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large language models are zero-shot reasoners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.688524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.688524Z digest=sha256:741201c76af9df00a7f859e9d0bef4089891698502947e242c3d6a5f7b8de92d

Observation 0dfad2ef-30b3-40ce-98e1-586cb125aa77 · outbound

This paper cites Mawps: A math word problem repository.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mawps: A math word problem repository

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.195779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.693132Z digest=sha256:e090f30aa692c02b8694a6178636fa531327c27eed550eb3fe487236a14cf9c5

Observation 74f166fb-caf5-4f6e-86cc-e6d481253606 · outbound

This paper cites Solving quantitative reasoning problems with language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Solving quantitative reasoning problems with language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.697680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.697680Z digest=sha256:e448104aa6bc909c5190e00f6cfa34fbc49ed8484951aa8fec347a45e15c623d

Observation c921ab3e-8305-40ce-ac3d-f7a0b99bd94f · outbound

This paper cites Common 7B Language Models Already Possess Strong Math Capabilities.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Common 7B Language Models Already Possess Strong Math Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.702855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.702855Z digest=sha256:72df69a0dbe4da6623c1b0abc0aa13c85dc09db268becab41731eb0303b2bebd

Observation 3aa96eeb-b20d-4d50-bdeb-84b6ff0231e9 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.707901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.707901Z digest=sha256:cc5973721fb4f3953de2edee0872b754ec679624a34d72de50337f42418bc4f4

Observation 433d176f-0b6d-451d-8053-e732d9c6b62f · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.712585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.712585Z digest=sha256:f0f018225331c19b5ffac781dfa31fc4805dd36b32cc09c21061cbcf5e60fe63

Observation 37d24616-1827-498e-a9f2-0b7911c0c595 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.717433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.717433Z digest=sha256:d75b977ef3315055d5050b5d0651189b6971b768b1ea5249b030d0f043370b9f

Observation ccbf91a8-5b0e-41c4-9b04-f03f8f0a8a8c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.722180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.722180Z digest=sha256:6378f7a754269844e613cc2120e9f57824e51c5b0a8c5a84b9df90f4fdb6942c

Observation 14c2d6a3-7291-4e88-ade2-344cbdc376f4 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.727254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.727254Z digest=sha256:a4a59326cbaf16ac3ec5d7eeda206b2267cccbfccf0102c41a31eb697889b851

Observation 599c6f31-cc86-45d9-93bc-32d5d4cbf542 · outbound

This paper cites Fairness-guided few-shot prompting for large language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Fairness-guided few-shot prompting for large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.169782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.732105Z digest=sha256:39338e9781e31e9e9c97f36336da1157472bf572fa4b2e7666abba3f00737b30

Observation a38806a0-1aaa-4941-8bc8-f89d5f7abca2 · outbound

This paper cites Mathstral.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mathstral

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.736586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.736586Z digest=sha256:3acd9885304a7b908d6a3a227b27108ab1c3cd712d0351c15ef58ea9968518e1

Observation 2c317ae4-a445-4d8e-8806-c85647dfd5a3 · outbound

This paper cites Mistral-7b-instruct-v0.3.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral-7b-instruct-v0.3

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.142404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.741261Z digest=sha256:4da70db3162043471392e929e51cd1bb45ac42fdb5ce9d7856cf2c5047f491ea

Observation 3b1b4e39-5e71-4bf9-94e3-4a461cc34d4d · outbound

This paper cites Mistral large 2.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral large 2

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.124881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.747037Z digest=sha256:8ced41f0a13e102926821e05ebaa9b00fd0d24f5d0cd2d97ac87ae127f7fa5e1

Observation aa9bbae8-0374-4506-9469-32bbc1d62eff · outbound

This paper cites The future of ai: Trends and predictions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models The future of ai: Trends and predictions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.107384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.751486Z digest=sha256:7284dec8d1707c73b90e0e4026efd1c2a112a0282dacbb2de3f847412efc7839

Observation 459c3e02-77f1-413d-9cdb-203c0485619b · outbound

This paper cites mistralai/mistral-small-instruct-2409.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models mistralai/mistral-small-instruct-2409

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.090764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.756854Z digest=sha256:370761e1e45a94b84650f4c6eeb3c704872fb3e4aaccfde4e529e86c3c6b6f88

Observation d97bf66b-9ffc-40ba-b4a9-771b8196c7ff · outbound

This paper cites GPT-4 Technical Report.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.761299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.761299Z digest=sha256:92f562630a35fde38a4e0755f981f7d51e4b8ec3ed8e21d7c861ce718a28f04b

Observation 457d8c92-56a7-459c-9e29-df681c5c9e25 · outbound

This paper cites Hello gpt-4o.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Hello gpt-4o

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.765864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.765864Z digest=sha256:ffd4854a0883fbc6ac94d1ba85ae06c71bf32d02b985d9cd8e8bb6386014c308

Observation 21421130-e9a3-4aa4-8fe7-5d00c6be18b3 · outbound

This paper cites Learning to reason with llms.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Learning to reason with llms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.770409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.770409Z digest=sha256:d2fb11ba8739b03a6c6c9c5599a784737f2ff92f4755bacf11c385bce7cae68f

Observation f11b28a2-617b-4ab1-af5e-66097d738444 · outbound

This paper cites Training language models to follow instructions with human feedback.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.775015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.775015Z digest=sha256:1e0ac6ba9950ee582d313833e819049b7b5ebb0f1a3ef30d5111aa8083b8eb84

Observation 024796c3-3452-4db2-ad50-c22106483284 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Are NLP Models really able to Solve Simple Math Word Problems?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.780365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.780365Z digest=sha256:812a8ae75858063e6d5b17a7096f6813ceb9db0a178bc30ae186875ad76cd49f

Observation 68217314-83f4-4253-a44b-77855e4a987a · outbound

This paper cites VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.786212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.786212Z digest=sha256:505a2d73e7ad39febb5ba7733b294066d91ca6e107f8997f964e41fa3e9fa1b4

Observation a0d1a685-4a18-4770-9069-0b01d45904ae · outbound

This paper cites Impact of Pretraining Term Frequencies on Few-Shot Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.790994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.790994Z digest=sha256:e5fbc3e27cfe614774787b553b6aa0d2157fd5a2b96faeefc9b752f28de5cdbf

Observation cce3f2c0-f7d3-4efc-bf83-9852a15f3771 · outbound

This paper cites To the cutoff.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models To the cutoff

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.042710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.795631Z digest=sha256:99654ceae46856d35201262b6252e31e250350621a3f5ac13351f6f0c621e7f7

Observation 1a048196-af5d-4ebc-a63b-cc9e3787ae95 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.800241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.800241Z digest=sha256:a1f3401fd7b7a32cdb204c44b3f19bf521582b9e3f947d44d5216d701b25d78d

Observation cf56f8a0-24e0-4295-98c7-21a16c8832ae · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large language models can be easily distracted by irrelevant context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.805110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.805110Z digest=sha256:b6965adf1a5f882f960efdca1a0d8221d32fdebb456260a2bee2e39d939e49ab

Observation a7e5437e-8496-4c30-8a34-ceca47b43642 · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.809872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.809872Z digest=sha256:035b199a31bd482b1526db4ae26cc5ca3f22ba35db625f40bf51e20f65dc995e

Observation ad55903f-86bb-4535-becd-65104a268541 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.814540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.814540Z digest=sha256:9abe47f7940b559d2d1f6ac8e298151dc6144c74ec7cec16dfa5406a6c7fa670

Observation 934e076b-8366-41db-af36-00a957359448 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.820487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.820487Z digest=sha256:5e7b717637bec7bd7900b5c888cceb7136034aa39e1b6860f12b363b91c645f3

Observation f2f0a032-c36b-4daa-8248-92a6c1db513f · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.825147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.825147Z digest=sha256:fe543c501c405d5414bfb520f83b20126f619f8b33befd48bc89afcee0958e8f

Observation 2a1c0d03-b390-4700-83c0-5a4ed1f5084e · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.830578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.830578Z digest=sha256:f1ff407051222962e2bbc0b57b1cd94033ff243ef6665ad85b0d2c1e1725a57e

Observation 89e4f12c-0e1e-45fe-a848-9dfbc2b60afe · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.835782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.835782Z digest=sha256:6e5d9e8654bf4e56e118cc6be6ea00d35954616db5004de86014ff8a88d9e7de

Observation 1a947438-dea3-4669-ac75-a59931a82729 · outbound

This paper cites Deep neural solver for math word problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Deep neural solver for math word problems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.015819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.840532Z digest=sha256:a1c1f2d7434123e4ad88196851e374647910e87c93810baa7ad435748f8cc043

Observation c01740dd-8ba5-46d9-bd61-ac4061053c35 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.845175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.845175Z digest=sha256:ee9dd874a284c99b95363bdc9c48c32b929f88b078dfe1de9444fb642f9d0304

Observation 2cb89d01-d296-4229-9d77-6f07b88bb2ab · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.850504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.850504Z digest=sha256:6e513ae159a6cdfaf92d6e78bd0527bd162410921dbc05ca01d3e387157f2012

Observation ec7cec21-656e-4d3c-bdc1-f3db0a5496fd · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Crowdsourcing Multiple Choice Science Questions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.855073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.855073Z digest=sha256:2d70bd1e035fca27a9a88f12a6a128e79699e151bbdf74b58d39ad3e2d7474c6

Observation b7273789-96d0-423a-b457-a92f31e57fcf · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.860944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.860944Z digest=sha256:897ecc5287bb7e61b936f17475bf2ccb69b123e977cc0a1e343ed3918468cd5f

Observation 9cd00303-8f6a-4d48-b38c-ded731bf435a · outbound

This paper cites Can We Verify Step by Step for Incorrect Answer Detection?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Can We Verify Step by Step for Incorrect Answer Detection?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.866556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.866556Z digest=sha256:52696a87af149fa6aecf9f23858a5147bc555eadbc171ee9e388ee8fb81fa552

Observation 23f0feda-a3e9-42b3-95f5-ecf2217b8a83 · outbound

This paper cites Can LLMs Solve longer Math Word Problems Better?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Can LLMs Solve longer Math Word Problems Better?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.872020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.872020Z digest=sha256:6cb1b7ce87dadb4a2fdc25182520da74dc54279fefaf7d2fe3c65a3fee0dd3e0

Observation 2220d358-a814-4ef2-94e9-06802039777d · outbound

This paper cites S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.876942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.876942Z digest=sha256:7b2e404321859cb15a0043534a6810ee07ae3353776cf488affc26aa159183e6

Observation 026e68af-1ce4-442b-8fda-b46bd6829f18 · outbound

This paper cites Qwen2 Technical Report.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Qwen2 Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.882278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.882278Z digest=sha256:0bb23f06eb27c421268c5566864c0e295c9efaaf220e1d0ca7f92361070543bc

Observation 47b04974-d71d-41ad-a51f-c9a1a33f51dd · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.887051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.887051Z digest=sha256:bb461f9b72cd8e0e008174954fd0694709452fe599f7862fd4839605e997cc59

Observation 4e1cef28-7bac-4abb-a63d-fa80d8a41e69 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.891943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.891943Z digest=sha256:0d475fe667490a5497f9a9a6654c7a22e6ea0699781643105d1a18d33fdc4286

Observation b89a4fc6-2a52-4f64-830b-90fc643e0569 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.897050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.897050Z digest=sha256:ea35b7b96124687dac90b0292c341b04d3e97ff656185f365390fed3149b973e

Observation a245528d-1133-4eff-8ad8-e3739da4b8f6 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:40.987548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.901836Z digest=sha256:4358e589960157e85013aa107336c077746341b29e919973cbdb4610f3046ccb

Observation 927131eb-c1e4-4242-935b-097e1159a221 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.907302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.907302Z digest=sha256:fae2fc299830c2b35ac17f4196543ee91707b50f98e0c3ffccef5628cc35520c

Observation 29d7a177-961d-447e-b8a0-95c4f6d97ac3 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.912031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.912031Z digest=sha256:d8e1d9e768d4fa2c3e415c5b8f631e3bbbc19394bff7e7ae2164d7d35d122110

Observation 05349184-bc8f-46eb-ab58-95304d74812e · outbound

This paper cites Cumulative Reasoning with Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Cumulative Reasoning with Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.916871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.916871Z digest=sha256:183bb1952d2f3f9288ca9965a443245b1b8e2e6e00ba5ef7f53bd111c9a3ec2d

Observation 6c3ec3ec-f72d-4321-bbf7-c7c786e381cb · outbound

This paper cites Progressive-Hint Prompting Improves Reasoning in Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Progressive-Hint Prompting Improves Reasoning in Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.921948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.921948Z digest=sha256:9d75e2461e8bb3659d38cdc523349c182d869889ac6a3883c50473bb050ce9fc

Observation 23df265d-f769-483c-a362-c3dfa64b58b4 · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.926795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.926795Z digest=sha256:dc83814c21ff6f7086af7e848eeb0fcc827ee74e5d2d8780b08df53d0e8f93e8

Observation ef799b5d-3e13-4643-8f81-5b823d7740c9 · outbound

This paper cites write newline.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models write newline

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.931716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.931716Z digest=sha256:c2fc44121628649c8fab66a6ee5ed9f58f3aa6be9fd6db0559a66c142b9bb7b8

Observation 00c936af-350d-401d-9171-b91ecdd63caf · outbound

This paper cites @esa (Ref.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.937620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.937620Z digest=sha256:4dbf81cd5a8b438300db7daf0d83c3aeb8b38a0c3044d5edcccffba04f6024bc

Observation fb76f3c2-b2cc-4f7d-afa2-9552682e12ae · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.944207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.944207Z digest=sha256:3765aed311483ebea99736247e085201623b9383ba71497a57556ca410769873

Observation b46613b4-3af1-405e-95c1-bd144fb6d148 · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.950474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.950474Z digest=sha256:1bc63933fce99a5ca6b6f5d8295526430e4f47d9d189ae85d99eb013bfed3c05

Pith citing papers

Observation db9ebfaa-48ac-4194-918c-b6c7ac841939 · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.891540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.891540Z digest=sha256:4c152ba95c1312eea21e88421271419e483d584b7d757a7c884e609a76fcf00d

Observation 2d14513f-748d-4736-b5d6-42ef87186ab8 · inbound

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification cites this paper.

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:59.476120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:44:59.476120Z digest=sha256:b8decf933e618ef3905af2184dee56d6c042664f0bd2643197cb91e74af7e089

Observation 0a6fec3f-c75c-4d2a-aa25-1949f32e86ed · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.330394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.330394Z digest=sha256:1224dac1f6c0a2cfcfe6b3c06bed6dc5f4ed2ea1a459164379a9fc28807f21d3

Observation e7eed9a8-b5c5-4b7a-8a56-6dc6792e5537 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:25.672136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:4db2d53d053cd13bc2b1d933514c0ce19244a1624a48518af18440ed4e5cd269

Observation da24fb6a-f16c-4309-85ae-be896ba80983 · inbound

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers cites this paper.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:14.534071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:34:14.534071Z digest=sha256:b9d40fa8755a6e6e75a04990e2ecf0403bd7f63eb943bdb01f400bb528633a8b