Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2402.19255.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:22:15.777907Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T22:39:01.102254Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ac0af52a-9b08-4286-9cec-ebf9de1c221a · inbound
Why Do Multi-Agent LLM Systems Fail? GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d82f4801-5635-4d86-9de5-ec53712e3d72 · inbound
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf3710cd-db74-44e2-8166-541cffa73122 · inbound
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee6d899-f5c3-4d6b-bd9b-fb6ed2167fa8 · inbound
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 2003
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · inbound
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473895c4-83b0-4638-8419-11988695da1a · inbound
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3c25241-11df-4826-b648-320d31f26066 · inbound
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e3dcf4-64f9-4e84-bc91-83c1113b8d8d · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec94d46f-565c-4263-ad3a-7b9b41f0a9cb · inbound
TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee31aa89-88a6-40bb-832a-e7dc0b7595cc · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c26b4548-2955-4d1d-8e15-10a847931c94 · inbound
RTTC: Reward-Guided Collaborative Test-Time Compute GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb0e41b-fc1e-4ce9-8803-6e85e63f5455 · inbound
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a41c17-c5c8-4748-bccb-6ece7fdedc66 · inbound
Model soups need only one ingredient GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1f310b-2525-410b-8401-16dd2bcd567a · inbound
Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d5d65c7-5b94-41c1-ba2f-150ec4a2f0da · inbound
Mellum2 Technical Report GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 047da0e2-690b-4182-a75e-6af01102d05c · inbound
Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff603d16-9f8c-4b7c-9906-6b7827ec79bc · inbound
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f60ae21-2436-4792-8642-6fd1f1c402a0 · inbound
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4868ac3e-7355-4b84-8942-39d14afd2863 · inbound
Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dee0196-0d50-4c69-8992-8c15a07ff5f6 · inbound
GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5024cc-f644-4e05-92c0-1e7f1bbd6a80 · inbound
Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b350ca-398d-4f5d-9ab4-02a053f0ee48 · inbound
Implicit Reasoning Steering via Concept Chaining GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fdfc55-92f6-4009-adf5-75351a2daed4 · inbound
Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.