Pith. sign in

Paper Citation Record · LEDGER

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning

As of 13 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2412.17397.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17397 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:32:05.187906Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:36:23.845366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:36:24.154635Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6139c537-2c34-45e4-8e74-2e0cbe47a633 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.007464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.007464Z digest=sha256:65a76fb28c5e6cb88cfff31c0c76ac2dbb9fba591e322f975d573b01a8eed30e

Observation ef7f8bf7-0f0d-4e82-9a03-493d51278cba · outbound

This paper cites write newline.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.012085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.012085Z digest=sha256:ab573c9645c8daf5468d8af9bbffa97925fba09ab6702272754626c6686359cb

Observation 6c59b594-4057-4559-a966-d31d539f04ed · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.016894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.016894Z digest=sha256:65503f68d78bec411d80445a545595106f0d190cf0f226aa922ce11331c7853b

Observation 716b68ea-4806-4bab-b082-434262358aff · outbound

This paper cites RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.021572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.021572Z digest=sha256:cedb64a180c099f426588a1a651476d4464b6ed9ca7f8fcf6f8f400870fba627

Observation 9ae8fbef-08ce-4121-a41a-c5cc708d1804 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.025822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.025822Z digest=sha256:99168cf9458c2c26010fb98142936d8fe49cb492aee728f0463dd94adb2e862e

Observation 66c77d0e-e133-40d8-a43f-febd697f0b96 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.030152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.030152Z digest=sha256:c5f3c6404eb010ba967c3eac4781b43adc551b9e04a5f41f506107e8aaebe508

Observation 41634b16-a98d-4af0-8a17-582fc290f46a · outbound

This paper cites The Llama 3 Herd of Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.033716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.033716Z digest=sha256:bf29bfd1330fada30df2aec0aeb34fa6da1b67371c6813cd55d1647165d1ba39

Observation e6724a79-e296-4dd7-a083-36322016908b · outbound

This paper cites Stop Regressing: Training Value Functions via Classification for Scalable Deep RL.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.037290Z digest=sha256:bcf1c192925d6386dcd5efc515aa182d95983faf6cb5da611c773dc90ceadd38

Observation 81f014e3-d931-424a-a8a8-12e468bd6d2a · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.866211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.041677Z digest=sha256:64b622dbc2ab536cb85cfd83f2f0c655bb26530d1c1e2348c4e2bf0620764bda

Observation 3c0ff638-9bb8-4f13-b778-5e620ca44c38 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reasoning with Language Model is Planning with World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.045671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.045671Z digest=sha256:8ce0d0165d80d0e8290849bdbd54057cb78a981965d306f9bdad056eb12f43c0

Observation dbf429ce-cf85-401b-9596-23268c5f20e6 · outbound

This paper cites GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.050096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.050096Z digest=sha256:24f6e4ae179a3efc626624fde9fcf6e641d47de4f3f862365582acade88b8291

Observation bcb1c06b-7139-42f8-81a7-08aa29e10fe8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.054517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.054517Z digest=sha256:44d41bcac24a4a91d0817903ca68900b1fcaa55400001e16eea7c5f6930ddb1f

Observation 1842be97-7bc8-460b-a9ba-ba85042970e7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.854172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.058941Z digest=sha256:94cdfcd2fc6ee5fd68f4d3711794963f7caca6f1f4ccb3fdb964f43955426030

Observation 65d61a2c-4590-4299-a8fe-b9d85fd0e639 · outbound

This paper cites Mistral 7B.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.062776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.062776Z digest=sha256:aa3347c6fd06f529d9b7e5d7d1926c804d740b15fecd0b391af7fb47ad69db6d

Observation e7b1a721-e5af-48b7-b9d8-048a594e227d · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.842298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.066919Z digest=sha256:f57304aa953977f81801af830fbbd363df472f3aff2c3d7a05a1c61df51f25fd

Observation 74cf42fa-1b7a-4a0f-bf3a-6bb4590f3474 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.071278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.071278Z digest=sha256:0dba4b251ca38ce15d41ed978e93f6e4e30a93a35cdc62f01f35f4f03426426e

Observation 68af0466-d675-43b6-a933-312d41ea6f2e · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.075477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.075477Z digest=sha256:a996d8961054adbec173466b6074c083ab72d0da04c358560aa2f2cfe8cba839

Observation c2b6c74d-03e9-4bf5-8291-ca00818a1fc7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.822646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.079473Z digest=sha256:26ac51cfb0bfd0b5f142defb5fe963af0df7295189c088870a737f585f6aa0d1

Observation 87c9b8d4-ace4-4bca-b25c-80f786c7afcd · outbound

This paper cites Let's Verify Step by Step.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Let's Verify Step by Step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.083498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.083498Z digest=sha256:ee7615182dc6880a705d316b40817c87c5f42809684de94d30458b121d2ce6de

Observation 6e460728-aeb0-42c6-b628-d7877ce4171a · outbound

This paper cites Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.087359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.087359Z digest=sha256:9bd1cad31b178e3863d27cfc42167ffb570f1f2a9563cef6d879d0f892881c00

Observation 1c3b8cd5-0cd1-4746-8667-0f53affc4735 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.811023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.091033Z digest=sha256:0084ebb41affcb4cf8b5b8e39b4fe569b0d6dd5c91d31a828f2371bdfa8369b3

Observation 10f51c8c-cbf8-4e23-b0af-e854b271ee2a · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.094613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.094613Z digest=sha256:7f0356f2f97ab647b1c4df5d14219200750e256e745dbbe759c96cb4dcbb7617

Observation 087d3004-cef3-48f6-96e5-55305776a496 · outbound

This paper cites REFINER: Reasoning Feedback on Intermediate Representations.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning REFINER: Reasoning Feedback on Intermediate Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.098268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.098268Z digest=sha256:98c5f87945df207bec2be5c551974a3b4c0a3f2a79821de6de8ca63e45a5e532

Observation c47a144d-ef43-4a36-b39b-70530ebf90d4 · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.102208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.102208Z digest=sha256:351c1b311a348733959a48679c748651dccbfe7496991b4e0917278410e675f6

Observation d6fe4b0d-cfa8-4917-ba07-97a6fb580142 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning D.; Ermon, S.; and Finn, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.106426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.106426Z digest=sha256:c5d00281309a975c2a9e122b731e5b86a4b34a012b81c4fc9bf086e69e5a5987

Observation 590eabcc-daa2-4879-8a24-6f44dd336a21 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.785053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.109858Z digest=sha256:7cd6bf34cf4c88ec5829d3562556f9755f9aed0d304c450710f818c47706de1a

Observation 8f2afc7e-e1c1-42a0-adc0-1045b8ccbb5b · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.113680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.113680Z digest=sha256:8f2891f99abe103895c2a40b364597bde2141353e555563c8ec622876c66a951

Observation fc4032c3-1d91-45f8-9527-a34ed4c73115 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.117598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.117598Z digest=sha256:f38f9b3597e520709a3d4ddb971992a5c81aff16dfe511787dab28718daacacd

Observation e2dd6853-cff5-4612-8f08-88757ad4fca9 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.121465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.121465Z digest=sha256:dd455b17fa35147ec1a262a6bfe9edfb45ec8bd77e18a341d92153bf5db50311

Observation fe3f1d92-ba3d-40d6-be20-3812083b914e · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.125316Z digest=sha256:121f411b5672c4b91d047fcbd13beb1176f6f7f45f9905af3b9587baefa31982

Observation d2f7f02e-3b18-485d-b1cd-f459a8631c03 · outbound

This paper cites Offline RL for Natural Language Generation with Implicit Language Q Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.128971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.128971Z digest=sha256:93617e49c62ae7df6479959a3d71d44f2b0a34cf70bf93d34169a98b4c679bc0

Observation 8a40f0ab-ea8f-4ca9-8c65-e908bf7b7388 · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.132531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.132531Z digest=sha256:beec033d2dab5eccc56cb5e73003ae56446835755106c9cf165e672dd3d102c6

Observation c4fd0710-afef-4a15-b087-da2858a9caff · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.140190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.140190Z digest=sha256:58639c63ff47212299a7f917aa3282810b8be363f87a0d2379de48666fe220fe

Observation 2aa358fe-ae80-4232-ac82-68d297a01d65 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.144130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.144130Z digest=sha256:6381f8bcf793bcac974cc999ddfa793398d17b7676197041c65fc6b93288a9be

Observation 174d3ed4-ae9a-4391-a886-334acd0be3ba · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Generating Sequences by Learning to Self-Correct

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.148451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.148451Z digest=sha256:a5e8600e45203d27a5c6fa0b9a201f85100e136820a697a5d350888ab6f1843e

Observation f1af8582-daa7-4077-a3b8-63c3cd2b011f · outbound

This paper cites A.; Ostendorf, M.; and Hajishirzi, H.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning A.; Ostendorf, M.; and Hajishirzi, H

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:05.774600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.153337Z digest=sha256:4c4f9c7366770cbbeb8d347cbcd85d6a594a20550b3da1a53dd1abb3b7ce098c

Observation 8af40692-2384-4830-b52c-7e7b4d258317 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.157765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.157765Z digest=sha256:776784113a0c9e8e738f784728c17afaed303803c7dd304bb6f89fa90a11a6aa

Observation 68b738bf-6638-44f4-8d49-efedbaff11b0 · outbound

This paper cites Self-Evaluation Guided Beam Search for Reasoning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Self-Evaluation Guided Beam Search for Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.161940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.161940Z digest=sha256:c85ed24e80826212e076c9d1e7342ddfce4f2ab57ad768bfa57d3dbf4e3c490b

Observation 87c2715b-ed7d-4e77-8b7e-84faff3d4faa · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.166120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.166120Z digest=sha256:6e2f163bf1f90a68cba475460d0001f93cfe0c921bacb6bcb4af7bd81d1a2093

Observation c7adbc43-b337-4f3d-98d4-cdbd1ffc4ac7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.170329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.170329Z digest=sha256:3f239daa9f4c7dbcf1165350eb69d8efce2948ffb17db913f20603f8e6bdb46a

Observation 5b01707f-56f0-460b-bddd-0eb8ccca49de · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.174422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.174422Z digest=sha256:3d74ba94be1cff7714ba368b9fa5e2439d7ea12219d6c90adfbab5c6ca378767

Observation 1be35d6f-3f50-4936-88ef-3fd03c599e9a · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.179010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.179010Z digest=sha256:657ec28f845b6088ad4f2ca3b3d4c4be8e0050e6014bbc09632b134119aa5be8

Observation 9c14088f-63e6-409b-82e9-e14c9c121d84 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.183607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.183607Z digest=sha256:c363fcf428ac91b29ae36243aa2aeb239c8e7eea06f8de7bcf302e72a3a0aea8

Observation eda91110-2467-45dd-8607-670cd82758d6 · outbound

This paper cites Solving Math Word Problems via Cooperative Reasoning induced Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving Math Word Problems via Cooperative Reasoning induced Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.187906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.187906Z digest=sha256:7a4735c321903a56dbab8c458bde12334b3a37f24fd8af07a781acf9d3582398

Pith citing papers

Observation 20d65286-dcfe-4a98-bb6c-fa62cf15d1cf · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.157579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:7fc9a2622d52586339f26f657d0e4edc8b6c636402a7f10f67081f97e92b18f3