Pith. sign in

Paper Citation Record · LEDGER

GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2402.19255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19255 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:22:15.777907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:39:01.102254Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ac0af52a-9b08-4286-9cec-ebf9de1c221a · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.742230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6ac4f725535f2c356ee51d1b58d204021882ae1efb9ab83cbd937bb0fef86224

Observation d82f4801-5635-4d86-9de5-ec53712e3d72 · inbound

LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models cites this paper.

LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:15.777907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:15.777907Z digest=sha256:fb96781d36ddfde86f72f30dd9f1dc22befe69798e646ff1a933988cfdf127e3

Observation cf3710cd-db74-44e2-8166-541cffa73122 · inbound

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation cites this paper.

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:58.814245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:43:58.814245Z digest=sha256:c4b2d7542203074cb85e5b751e442f34e837be3e7b8afb3d020a82cbe9c2090b

Observation 3ee6d899-f5c3-4d6b-bd9b-fb6ed2167fa8 · inbound

CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective cites this paper.

CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:38.902389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:38.902389Z digest=sha256:8603e3a02cd950a079022ae866f427c28160323079525eda4a16d6669204d15c

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · inbound

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation cites this paper.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:218c0aea0a4e47e8f95ee2e1ffce93fd5c278fb97225e13ff7288702e677185a

Observation 473895c4-83b0-4638-8419-11988695da1a · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.343286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:d79da6bbad283316820da488a8e4fc10bfa24db5aac80b78d455a5eb1255f188

Observation a3c25241-11df-4826-b648-320d31f26066 · inbound

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding cites this paper.

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:50:51.172050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:50:51.172050Z digest=sha256:115e9b51f9d1e23b5554bbba056fc49ff978b14ddd5c9dd7f9f49ac6478d99ec

Observation b1e3dcf4-64f9-4e84-bc91-83c1113b8d8d · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.340074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:4b818ef790e58e8970adfeb579e1a9740cbd4865074a2c3da342a934446b1326

Observation ec94d46f-565c-4263-ad3a-7b9b41f0a9cb · inbound

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards cites this paper.

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:46.556756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:46.556756Z digest=sha256:0f918ccc269141a7bdbea6e38022828f3d54276516a2eac72792cbb57f95de64

Observation ee31aa89-88a6-40bb-832a-e7dc0b7595cc · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.279904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:c4b6f8f5f2a794937fe0c636fc7ad62e0cb9e066197c82d8eef2f171b400b6ee

Observation c26b4548-2955-4d1d-8e15-10a847931c94 · inbound

RTTC: Reward-Guided Collaborative Test-Time Compute cites this paper.

RTTC: Reward-Guided Collaborative Test-Time Compute GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:44.161130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:10:44.161130Z digest=sha256:3c76e17c7b23fa48edada4b185323fd24eebc3fe824c2540436defe3809b98ac

Observation bdb0e41b-fc1e-4ce9-8803-6e85e63f5455 · inbound

FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline cites this paper.

FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:20:56.968306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:20:56.968306Z digest=sha256:215c991e455ea411012555136a9c0bbf58b37f5524caddeaf1a26ac74bea571c

Observation c0a41c17-c5c8-4748-bccb-6ece7fdedc66 · inbound

Model soups need only one ingredient cites this paper.

Model soups need only one ingredient GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T02:51:55.289859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:51:55.289859Z digest=sha256:92efba14bd03ae32d4a7de93f963678968831f1dcff9655be7510f6190d9c21c

Observation 4a1f310b-2525-410b-8401-16dd2bcd567a · inbound

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models cites this paper.

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:00:04.688097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:57:48.890310Z digest=sha256:daeb279ad93e7a9f90effeaa26368a0a6b6803d8c26a7c3376d27e1a15900fac

Observation 5d5d65c7-5b94-41c1-ba2f-150ec4a2f0da · inbound

Mellum2 Technical Report cites this paper.

Mellum2 Technical Report GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.503173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:58:35.397914Z digest=sha256:38b723dd3a343b9c8ebcf3ffc6153abe2269abf7d981fd3c2b0aba5ac169d0dc

Observation 047da0e2-690b-4182-a75e-6af01102d05c · inbound

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks cites this paper.

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.204133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:27:30.923556Z digest=sha256:25035eafae7625027a8aa20d6592bfa19567f6659a14cf0548b97ba3fe2d65f7

Observation ff603d16-9f8c-4b7c-9906-6b7827ec79bc · inbound

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories cites this paper.

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.340135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T00:36:14.316735Z digest=sha256:e0ad407465c024695c1302ce1bf813d0325b0b77969617ea8cc7f0beb2422e4f

Observation 1f60ae21-2436-4792-8642-6fd1f1c402a0 · inbound

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories cites this paper.

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:39:01.103776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T22:36:12.063779Z digest=sha256:ea09040046eec7c4779d131e2fd769d395b37af9e2ce2f817286edf7235d2d0a

Observation 4868ac3e-7355-4b84-8942-39d14afd2863 · inbound

Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models cites this paper.

Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T18:56:31.772618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:56:31.772618Z digest=sha256:aa189e28e9af03250d7195bea164d382506720deaa0bed35c36e96ae8075dd27

Observation 7dee0196-0d50-4c69-8992-8c15a07ff5f6 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.502614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.502614Z digest=sha256:bd5c9eaea2f4e3cda3ad86e9ad47d371ac5139b2c97c8c74e4290a2252ef2542

Observation bd5024cc-f644-4e05-92c0-1e7f1bbd6a80 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:22.594941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:22.594941Z digest=sha256:d7eaa60e4ccf841e0bfd8f8fb3c6fcb604e82977ca45b42443e323cd1092f4db

Observation a5b350ca-398d-4f5d-9ab4-02a053f0ee48 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:33.784184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:33.784184Z digest=sha256:7bbd95a7ba1511db807dea1f8844a86bb29bafb0496ac3da5072dfc030a60490

Observation d6fdfc55-92f6-4009-adf5-75351a2daed4 · inbound

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs cites this paper.

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T20:19:22.327652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:19:22.327652Z digest=sha256:0a67365acb9e87725daa32a1191abcdc4cff060cf49c5115170c3d2d0ec7561b