Pith. sign in

Paper Citation Record · LEDGER

GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2402.19255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19255 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:26:07.602912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:39:01.102254Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 273b4196-04ac-4f45-b529-a955dc87d4e8 · inbound

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations cites this paper.

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:07.602912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:07.602912Z digest=sha256:80cac6ac0fef7abeec7caf829dad9420f948dc3e9e6b28d1fbdf52b4b37e34e2

Observation ac0af52a-9b08-4286-9cec-ebf9de1c221a · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.742230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:dc158b32682c7fb561855f5a8cb7815affe38d1fd9706bf2c8afb5ce1fddd837

Observation d82f4801-5635-4d86-9de5-ec53712e3d72 · inbound

LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models cites this paper.

LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:15.777907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:15.777907Z digest=sha256:fb96781d36ddfde86f72f30dd9f1dc22befe69798e646ff1a933988cfdf127e3

Observation cf3710cd-db74-44e2-8166-541cffa73122 · inbound

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation cites this paper.

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:58.814245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:43:58.814245Z digest=sha256:c4b2d7542203074cb85e5b751e442f34e837be3e7b8afb3d020a82cbe9c2090b

Observation 3ee6d899-f5c3-4d6b-bd9b-fb6ed2167fa8 · inbound

CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective cites this paper.

CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:38.902389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:38.902389Z digest=sha256:8603e3a02cd950a079022ae866f427c28160323079525eda4a16d6669204d15c

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · inbound

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation cites this paper.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:218c0aea0a4e47e8f95ee2e1ffce93fd5c278fb97225e13ff7288702e677185a

Observation 473895c4-83b0-4638-8419-11988695da1a · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.343286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:64212344c3c464ddee3bc4167b39823462f812d6c822b78a4e987178eadafcbe

Observation a3c25241-11df-4826-b648-320d31f26066 · inbound

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding cites this paper.

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:50:51.172050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:50:51.172050Z digest=sha256:34fb4156e304a26826e71e38cd5d46717d29a3f0aef8d9dda659100c8bf0808d

Observation b1e3dcf4-64f9-4e84-bc91-83c1113b8d8d · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.340074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:2118b26196e7f25f7b0676371ab5a861adb1118271ab1cf66520a4d21b3237fc

Observation ec94d46f-565c-4263-ad3a-7b9b41f0a9cb · inbound

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards cites this paper.

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:46.556756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:46.556756Z digest=sha256:0f918ccc269141a7bdbea6e38022828f3d54276516a2eac72792cbb57f95de64

Observation ee31aa89-88a6-40bb-832a-e7dc0b7595cc · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.279904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:8564b2f93b17cc710680009b32da2ec566b0839e0def7b231c1125857cfa8d2c

Observation c26b4548-2955-4d1d-8e15-10a847931c94 · inbound

RTTC: Reward-Guided Collaborative Test-Time Compute cites this paper.

RTTC: Reward-Guided Collaborative Test-Time Compute GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:44.161130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:10:44.161130Z digest=sha256:9a8ab4e0e63bbeeea80954179690241be8c4c6ffd39bec958fcca4a012ab9a8d

Observation bdb0e41b-fc1e-4ce9-8803-6e85e63f5455 · inbound

FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline cites this paper.

FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:20:56.968306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:20:56.968306Z digest=sha256:215c991e455ea411012555136a9c0bbf58b37f5524caddeaf1a26ac74bea571c

Observation c0a41c17-c5c8-4748-bccb-6ece7fdedc66 · inbound

Model soups need only one ingredient cites this paper.

Model soups need only one ingredient GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T02:51:55.289859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:51:55.289859Z digest=sha256:92efba14bd03ae32d4a7de93f963678968831f1dcff9655be7510f6190d9c21c

Observation 4a1f310b-2525-410b-8401-16dd2bcd567a · inbound

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models cites this paper.

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:00:04.688097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:57:48.890310Z digest=sha256:2c4eb6453a15759ec32a13bdf58ebdd626e99413a4f76a239d213c4ac7b587d7

Observation 5d5d65c7-5b94-41c1-ba2f-150ec4a2f0da · inbound

Mellum2 Technical Report cites this paper.

Mellum2 Technical Report GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.503173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:58:35.397914Z digest=sha256:e8191be16421a1a727008a41771323842f6aaea39421ce0db3c8ab99b365ef0a

Observation 047da0e2-690b-4182-a75e-6af01102d05c · inbound

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks cites this paper.

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.204133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T09:27:30.923556Z digest=sha256:a6563a2e3037343a440154578b0370134cd543018d6f98876e37280601e2ded6

Observation ff603d16-9f8c-4b7c-9906-6b7827ec79bc · inbound

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories cites this paper.

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.340135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T00:36:14.316735Z digest=sha256:9ac5b0a8bbbafb41cf8dffbf8aaf45017dd4d61f1e99a415ef544582f614757f

Observation 1f60ae21-2436-4792-8642-6fd1f1c402a0 · inbound

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories cites this paper.

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:39:01.103776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T22:36:12.063779Z digest=sha256:3aa1a46b92ac16b8ac1e0c5971cd307c7ea298bee0fc38d8b7e0fa6660540584

Observation 4868ac3e-7355-4b84-8942-39d14afd2863 · inbound

Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models cites this paper.

Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T18:56:31.772618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:56:31.772618Z digest=sha256:aa189e28e9af03250d7195bea164d382506720deaa0bed35c36e96ae8075dd27

Observation 7dee0196-0d50-4c69-8992-8c15a07ff5f6 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.502614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.502614Z digest=sha256:bd5c9eaea2f4e3cda3ad86e9ad47d371ac5139b2c97c8c74e4290a2252ef2542

Observation bd5024cc-f644-4e05-92c0-1e7f1bbd6a80 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:22.594941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:22.594941Z digest=sha256:d7eaa60e4ccf841e0bfd8f8fb3c6fcb604e82977ca45b42443e323cd1092f4db

Observation a5b350ca-398d-4f5d-9ab4-02a053f0ee48 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:33.784184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:33.784184Z digest=sha256:7bbd95a7ba1511db807dea1f8844a86bb29bafb0496ac3da5072dfc030a60490

Observation d6fdfc55-92f6-4009-adf5-75351a2daed4 · inbound

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs cites this paper.

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T20:19:22.327652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:19:22.327652Z digest=sha256:0a67365acb9e87725daa32a1191abcdc4cff060cf49c5115170c3d2d0ec7561b