Pith. sign in

Paper Citation Record · LEDGER

Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2410.06508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06508 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:25:52.090103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.951632Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3e000929-d484-452e-854b-999eaab8a447 · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.090103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.090103Z digest=sha256:e98a63e550eada290aefd3d42af416dc7744ddc39f12b5133f84dd29524e5e08

Observation 490f703c-c4cf-40e5-a257-a31153a9bfa5 · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.580233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.580233Z digest=sha256:3d97f29bbe6983a818c5db0dbd40899407e98deab68b9eeeb8cdfd450e247baf

Observation 9907e356-e0ea-45e9-9e39-a81b4b91975f · inbound

MARCO: Meta-Reflection with Cross-Referencing for Code Reasoning cites this paper.

MARCO: Meta-Reflection with Cross-Referencing for Code Reasoning Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:34.648272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:34.648272Z digest=sha256:b33bb97c97da9f31d1217cfe5a511f3cc4baaef303c0db0b94fe3bc94f50923b

Observation a085871c-a7ba-4de6-806d-5b693720461d · inbound

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning cites this paper.

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:18.952147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:45:18.952147Z digest=sha256:f977e7a4c6cb2bd423fb02a9792cdc8dc2ddee3f32911991c431d64b4b709847

Observation e83cabac-7a02-4576-ba82-f0d1e5c6e3cb · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:15.655281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:15.655281Z digest=sha256:df7ca36975fa82473b0131967d54220b689d9a88f91a353c884cf3eeda996944

Observation beefff04-48de-4cf9-b381-03c75111e970 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.770333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.770333Z digest=sha256:70dedc51574856e471cfbc59223221baf686b604e581a712dada76198d4618de

Observation a2352f16-e9ff-4379-9aaf-4d06f4f24489 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.953706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:7e3b523e711939ccba628048c1b6dabf19c1cbbd900c8df8f2d6bd7518b9134e