Pith. sign in

Paper Citation Record · LEDGER

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2504.17665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17665 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:28.303524Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62a960ec-5e32-450a-96d9-08c7fe067a1a · outbound

This paper cites online" 'onlinestring :=.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.159490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.159490Z digest=sha256:888d34e20f304e45946e72c50cbba463be8f156a8b647a6058bc940b2ca0d051

Observation 506503f3-8331-40a1-9a64-ba58128b6830 · outbound

This paper cites write newline.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.256611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.256611Z digest=sha256:f9b4b639272ef156ca115f4c2c645ff8f24db6c409ec1d5310fa048c3710084c

Observation e79519a5-0baa-460e-a67d-5feeff3aba02 · outbound

This paper cites GPT-4 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.277342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.277342Z digest=sha256:130795ad2abe7f08f75d0e1695f331a7be0c5b6dc92e032128a2562df18a1b72

Observation ff58524e-b33a-4208-8c31-6305b3699408 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.285575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.285575Z digest=sha256:d8faae58c6592295ca6f87e57a821c92924e36f3f94ed7bc238d1ebb18acf947

Observation dada4289-5d98-4b00-82b1-e2e86c9b2d3f · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.305525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.305525Z digest=sha256:6696e5eddd18b00f6454e0057dd59d1a0ef0010fe7be4bc9e60e81d0854b0317

Observation 2f794079-b4e6-4394-8ddf-6b4cdcceb854 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.911559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.408830Z digest=sha256:ca4898a387fbfe2eb40b9f3eb4ea19f8ebd164cde9aeec903e76a38ccce3fe51

Observation 2d75e65c-7674-4401-aa0b-92df0c40e0f7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.435610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.435610Z digest=sha256:2f5a29a5e3294f197e6827e80332e473f0f04be8b60100f52d04ed82fe6656d0

Observation 32c92f81-8e5f-4c35-b37f-9c1392d60637 · outbound

This paper cites MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.447288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.447288Z digest=sha256:5755db3aa2c9446a2a89b79ae9641b14ac75dc52be6ba7261c771ada4b23e38e

Observation 4f422dc0-0654-44d6-bdbc-bda1a0d83758 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.534750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.534750Z digest=sha256:0ba57fccc99d3c3412d47064503ebeef2d08a2babe1471353360a270c48217b0

Observation 8749a97e-aa6b-4800-bce8-8b7d742c5859 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.580437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.580437Z digest=sha256:fdd05b3f9e6d1f76a65adb639d7607791cd379777c2363874cf96aee27a4a3e4

Observation 7b4b1caf-91ca-4e0c-bb3d-20effa818813 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.842314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.678437Z digest=sha256:2e15362d72dc82947bd2d8cce5dcf066bbe9f373b6a2a22409b673c1c075fa92

Observation 1d544743-a921-459d-993d-eb76b5a5ada0 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.705902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.705902Z digest=sha256:13df7f459ed044881794b7640abce63ccf489252f529f41009ba8049f45e32d5

Observation da02a7ba-5c01-4ac7-af52-ff898ee082b6 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.721732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.721732Z digest=sha256:22aa01cc56bf7edf1d3908ee9504e6944dea87586aba7c41651cb71a835cc480

Observation 7b7701e4-daa6-4806-b3cb-94f779d066c7 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.731733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.731733Z digest=sha256:7ef667f7576c8d881197884b704dc190b036a4c129edd2ab362bd4fa563fabb1

Observation 83f28c52-37e4-42fd-abdb-b02d042ec52e · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.817668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.817668Z digest=sha256:50729747a0da26c91cbb1a98891aff083b434fb4d910717b0bf1d2e07d8a2ab3

Observation a1e98042-fb50-45d1-aea0-72ac99a2a6e1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.856738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.856738Z digest=sha256:114f543579cb0dfab1a842ec6d6816658d53ce741cad7a03351ec0a03e8e43e8

Observation 7fbaa48e-c3f8-409b-b597-87b03c0d2784 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MathPrompter: Mathematical Reasoning using Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.861159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.861159Z digest=sha256:6114981721a8fe6ebdc8e7b3ac1eaae3ac5a5ff02acf80b9f77e07ddda9a96a2

Observation 8d49385a-1479-481a-a344-7af00853427a · outbound

This paper cites How Interpretable are Reasoning Explanations from Prompting Large Language Models?.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics How Interpretable are Reasoning Explanations from Prompting Large Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.876660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.876660Z digest=sha256:4834af339fa40d0905ef5e77d50da052e3638538cac3dd195590f313410e155f

Observation ac83fd1c-703b-4bb0-af44-4e9f7c456e1b · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.884720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.884720Z digest=sha256:9600cba80ed3aec378fa0a9c0f1c34e9fa5f3e5385df1b0c75139dd68dbc8d29

Observation 1b945a6d-a71f-40d9-a7c0-fc80f0ffb998 · outbound

This paper cites Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.976338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.976338Z digest=sha256:976afc70c1a5c7a229e318938096baa64a092fc93e5ec9ef98f9fb5728d2aba3

Observation fdf7c8da-a15a-4dfa-8579-c6ceee6a6bf8 · outbound

This paper cites Let's Verify Step by Step.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.008119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.008119Z digest=sha256:9c5a3a9f95b84e483d1bf307b0f5e41f9fae6146e260f822744614392f0c5134

Observation a5df8063-c254-4073-ad11-dbe39036e38c · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.012469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.012469Z digest=sha256:e3f50994b146904776793a0a47832c20fce3f82eccf41025dea3696573ae4aaa

Observation db3513a4-b966-44bd-af55-159ec022d21b · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics StarCoder 2 and The Stack v2: The Next Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.017095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.017095Z digest=sha256:c1dc3906d959dbcbfac8c5ec8da54847841437a754ebb80496e6fdb9cf8e0a09

Observation 4cf74b68-30c9-40b6-8623-6fd74e7b2993 · outbound

This paper cites Faithful Chain-of-Thought Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Faithful Chain-of-Thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.021493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.021493Z digest=sha256:6dceb4931692f47112062802dd47e0321c7bfe1c2b807d12dac2c54167a088e9

Observation 9ef4e6f1-ae80-4668-b1f4-36bf68506952 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.025866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.025866Z digest=sha256:274b95daf919b42759c54da034bec43f32bddf706d5e65bfa6d78a41bb2dc180

Observation 8453e4c9-49b1-47bb-8ecc-f7b35f7bbb24 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.069599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.069599Z digest=sha256:607bc845a50586f87d2a0ee770306e23b0a06718da900e7768b340f4747d8fdc

Observation 1e4e2223-d431-4ae9-9af0-a9e31ef49758 · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.135870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.135870Z digest=sha256:420b8c2c4a32c61abcb10ed77fee282daea41dfddcd9e9b38649749396217e75

Observation 10a79d37-2910-49f3-aff6-d11fd8c59d4c · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.140228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.140228Z digest=sha256:039d40d534f3b0dab406be7fc1f14e1f0ca4d55ec8a62cc81a6ae05681ea1cdb

Observation 63a95e6a-6204-4928-98c8-5864781550c9 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Code Llama: Open Foundation Models for Code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.144373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.144373Z digest=sha256:f05620a2deccadab5d128e7965347b9ea26f5da11368efa511ed56eaee6759b1

Observation 280b389d-db6e-44e7-bb5d-e48cbf25c1b3 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.148351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.148351Z digest=sha256:6f4eb1ad66377c7777eca4d8765afa66e4a038193cc5c629235bc177981ee357

Observation 870b4498-931c-41cd-a6e7-abbc8839b71d · outbound

This paper cites https://lyz-code.github.io/autoimport/ autoimport.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics https://lyz-code.github.io/autoimport/ autoimport

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:38:28.666950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:28.151866Z digest=sha256:6a44ac269f30c606887acf09765b4bc185d5576cebcbf7b02bfaf8043097d273

Observation 444747c2-aba4-48de-b594-2650b1891a9a · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.157210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.157210Z digest=sha256:9a2547dfca592ef0534b0e36e49df6c5ff851262ca883bcd1939fb6e3c3950d5

Observation 0572d174-3895-4ae9-b9ed-00e10be25196 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.161671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.161671Z digest=sha256:2672224b34d760cf5a3eb3f47088d195e4e7638c85385e152f910e53641dbdd9

Observation 37de6bae-2ec7-4da3-805b-3b2633deb28f · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.165082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.165082Z digest=sha256:9fd0fdd5e1b911d3fa233d58ea726a3efcfe405103717203637319d327769181

Observation 386764f5-7076-40d2-9519-fa12012ec13d · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.255976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.255976Z digest=sha256:2a03f15d9fd75def4438baf72490bb0b4c157fc63eb7fa1fbe4641c2a9db4a21

Observation a7e1db0b-a0d7-4091-b74b-704b71752da3 · outbound

This paper cites CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.303524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.303524Z digest=sha256:f8e664a2b0843a76240e5d337d292e4736564d5ce8e4d8ca5d32453193ade840

Pith citing papers

No inbound Pith citation observations are available.