Pith. sign in

Paper Citation Record · LEDGER

Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2309.17179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17179 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:42.924167Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aea7076a-c96b-45b0-84c0-e62d2b2c97e4 · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:34:15.902550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:6a4c929095d4428741e9340988d6b8ee935d5c96e9eb46cb70a04ba88efbb678

Observation a9446a2e-c207-429b-a153-efdfed86db07 · inbound

Improve Mathematical Reasoning in Language Models by Automated Process Supervision cites this paper.

Improve Mathematical Reasoning in Language Models by Automated Process Supervision Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:45.903535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:53:45.878221Z digest=sha256:14a13f334be566c7145382bc3ab4897820a9760d0c70576552d9499cc21c1407

Observation cf78593d-89e2-49b2-a804-19c8160fc570 · inbound

Natural Language Reinforcement Learning cites this paper.

Natural Language Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.000653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.000653Z digest=sha256:2db85e54f7957ae4d7ddf1301b969ec286cfe9079d8bf317812fc74ab1c3d8ff

Observation 251325a8-35f2-4446-92ba-b6cccf010c91 · inbound

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions cites this paper.

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:55.094867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:17:55.094867Z digest=sha256:b5b8d7d84d76b4633112d4bd9d06e4e99c2eeb97127c0addbe66d91fc8f5156b

Observation 7645cd5e-ad11-44b8-8dc5-db8d4b2aa4a2 · inbound

PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making cites this paper.

PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:43:44.467524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:43:44.467524Z digest=sha256:75c358bf8f3ab740a9fc61059f1689f468990d59d9731f45d925951db2490b28

Observation 45e7be58-dd37-41a3-aea0-62891708d141 · inbound

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration cites this paper.

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:49.604476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:41:49.604476Z digest=sha256:bcca54801a6cf61a760f0ccea3cf403923dcf252c2499b8c7a8b92644266e1f8

Observation 59da8fea-a7fd-4404-b724-dbec5d540773 · inbound

o1-Coder: an o1 Replication for Coding cites this paper.

o1-Coder: an o1 Replication for Coding Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:11:28.970300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:11:28.970300Z digest=sha256:67b6c45b765e37540e589c4b186449460b5daf36efacd6e29463032d2d2cac6c

Observation ec385732-d4b8-4987-9308-d193a39b7662 · inbound

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models cites this paper.

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:14.895690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:08:14.895690Z digest=sha256:b01a74fa87fb6dec2ac053303a68d416cee1528682d73b25ae1c085b6267460c

Observation 7c41dc8e-5dc9-4f61-ad6b-dba38235787a · inbound

LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments cites this paper.

LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:55:52.756862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:55:52.756862Z digest=sha256:1015a40f06eb79a0731f088e876cf4d0cc8b7c5fbac5c8d66cecf766cfb47e6f

Observation a0b38c5f-97f0-465f-bf7b-bfe49f0da580 · inbound

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning cites this paper.

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:08:56.359209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:08:56.359209Z digest=sha256:45277ee0a9f3553abc9fc8de3a25df4cab0b3192ef85e392adcda863275ec55c

Observation e41f556e-da6e-452a-af56-9974dfee858a · inbound

Offline Reinforcement Learning for LLM Multi-Step Reasoning cites this paper.

Offline Reinforcement Learning for LLM Multi-Step Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:19.067660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:50:19.067660Z digest=sha256:6e8f277d841e3dd8cde0ae43fd58d99416c98ac04853181ae81dbe0f8f15afdc

Observation fffcf75f-e6d4-474b-8407-31f54595ab11 · inbound

Efficiently Scaling LLM Reasoning with Certaindex cites this paper.

Efficiently Scaling LLM Reasoning with Certaindex Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:44.269878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:44.269878Z digest=sha256:12aada81659a886a4088dd718ae65b596edf704be6e772f187f7246d3f302d0b

Observation d7acb99e-dda2-4a48-8d3c-85826cb4ee4a · inbound

BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning cites this paper.

BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:50.718863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:50.718863Z digest=sha256:1b238c514942c74e0003078635ecf22a79347aef23daa5ba335c6da214e3b03e

Observation 11b85831-dbcd-4a8f-acf3-7e651b70d0c8 · inbound

Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design cites this paper.

Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:29.165127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:29.165127Z digest=sha256:4465a8b482ea71859c6e26da18ab98afdc7b876437bd5b20b4dffc52f5eed4d4

Observation e221ac50-fc44-411f-b8c5-01529f990319 · inbound

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning cites this paper.

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:15:46.007079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:15:46.007079Z digest=sha256:aca4ce12f3e32ac18ac5ceb641d02e9e9a6f5f8c255d9b4a7b2926cb3cfaa3e8

Observation 18bb0357-4e3d-4883-8230-f7d2ff9d439a · inbound

Large Language Models to Diffusion Finetuning cites this paper.

Large Language Models to Diffusion Finetuning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 1969

Resolution
unresolved
no resolver link, observed 2026-08-10T14:01:55.493801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:01:55.493801Z digest=sha256:a27cdbea32fe0aee771ae8b9156bb25eb30a966365aff277cb6a0781c7ca29c7

Observation 469e47ad-84b8-460e-a008-a747f6fe8abb · inbound

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation cites this paper.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.626962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.626962Z digest=sha256:a47e342ff34e528c2a6b76fe1ed7c1da646096e10ba1f4d9c1ce41b76a9e85ce

Observation cfbc782a-e2b0-4875-8f83-4673f0a6e7f6 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.242429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.242429Z digest=sha256:3fe26e894fd13b44e4975e7d2e97354d0dddd955a17f75f443ef2cfac4e047ed

Observation bdd185cc-189c-4585-a351-a8aeff4a8531 · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.332278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.332278Z digest=sha256:17568168466d1c31e296f8e5c6b80294745fc915904482a4d5de6d2a53240345

Observation 4ef57d7c-c201-481b-b7ab-80775d952152 · inbound

Policy Guided Tree Search for Enhanced LLM Reasoning cites this paper.

Policy Guided Tree Search for Enhanced LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:31.753947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:20:31.753947Z digest=sha256:d3ba078b530edfb5bb7a3af72f7e0191fb0b46025af05693be55c590e452d330

Observation e92021b1-0c6e-4369-8774-ec43f7b0ccf9 · inbound

Bag of Tricks for Inference-time Computation of LLM Reasoning cites this paper.

Bag of Tricks for Inference-time Computation of LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:24.496163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:24.496163Z digest=sha256:76ec055442128c4d0238071cbd19c72f59b4d7fb28b9ec9e5d8c3fd6ea7d9bd1

Observation fb68f2ca-0621-4020-bd19-b463ba31ccd1 · inbound

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement cites this paper.

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:31:43.547304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T12:31:43.494099Z digest=sha256:bf7d3a31018d68cd9a359be8d5f41dcfc8be23735b299d271328c7be5c8394f6

Observation fb995c4c-c156-4409-a4c3-748bfb0c287a · inbound

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model cites this paper.

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:45:04.025581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T19:44:16.630377Z digest=sha256:293a8554d513e097f83a624cc39f2a8e2a745d05002ad6622cd58cf24b33fc62

Observation 86b7f0f0-f364-4cc2-b6a8-e887299adba3 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:42.924167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:42.924167Z digest=sha256:c7e579a61f9603c0f8fc125789f58e714b49cad6bb96263043f43651621bbdd1

Observation 4e2c4003-3996-4161-8e02-189bb2881a20 · inbound

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning cites this paper.

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:34.880310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:34.880310Z digest=sha256:a54250557e8c37319ed0ef2846a6d4e19fad71f9b1a0ff91d6b6c95f1e5c4b3e

Observation b176da2c-6234-4d68-9946-9070a536d876 · inbound

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies cites this paper.

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:20.370996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:20.370996Z digest=sha256:197e42034099ecf0d8626fc6f39e83aaee5c299fc647aa095684d3a82644bacc

Observation 98b5ea52-b3df-45db-a343-775bc98530d2 · inbound

Safety in Large Reasoning Models: A Survey cites this paper.

Safety in Large Reasoning Models: A Survey Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.087814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.087814Z digest=sha256:dc87f7d6d597bf8be8e1e042af933086db6b45adb6fbc782534ce7b28079fea3

Observation 5b3048f0-41ac-45e5-baa7-7ad496a900d7 · inbound

COSMOS: Predictable and Cost-Effective Adaptation of LLMs cites this paper.

COSMOS: Predictable and Cost-Effective Adaptation of LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:17:37.794174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:17:37.794174Z digest=sha256:c7fa43d908a092a43fe76fd13dcf53a74489669a2c3e987b1c5e7356070f3bf2

Observation 1da66c8c-df16-4f25-9390-592bbbca1659 · inbound

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning cites this paper.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.863816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.863816Z digest=sha256:a00ee08a3ff29bc8fb4c7e8d0aef44fb0df62ee87e8092fbf4319f3f57634bab

Observation eee365db-c203-4bf0-a66f-19a157c902cb · inbound

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO cites this paper.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.146231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.146231Z digest=sha256:e4d5cf2591d7dfef3c18f332b1834288fcc11d0240600a9b934323be2f84a987

Observation 23cc425d-47c5-4bf8-842b-656c2bc06366 · inbound

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision cites this paper.

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:23.309802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:23.309802Z digest=sha256:bd516feee45c2c8860cb119be8477f356294e81de6e8d65cf922a86c2de18e04

Observation 03ff09e9-1f0e-4415-ad95-afb2a0f01966 · inbound

Generalizing Large Language Model Usability Across Resource-Constrained cites this paper.

Generalizing Large Language Model Usability Across Resource-Constrained Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:53.403306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:53.403306Z digest=sha256:2bc68ea908deb119266b441552fba80b38449c5dc664cbeabea746011fd95bee

Observation beb2dcd8-ce09-4a45-ac89-9df398fc6e18 · inbound

MMATH: A Multilingual Benchmark for Mathematical Reasoning cites this paper.

MMATH: A Multilingual Benchmark for Mathematical Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.912353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:23:07.912353Z digest=sha256:3103811987bf47a5c98490636f528a4a8e153417293cf4aa753936effe19baf4

Observation db427559-1d5c-4206-92ed-026d48eec8af · inbound

Can Past Experience Accelerate LLM Reasoning? cites this paper.

Can Past Experience Accelerate LLM Reasoning? Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.952185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.952185Z digest=sha256:32e4c11af816d9e71f212087a76072d538e3296aa2596cc4365e6e925653328e

Observation 8769437d-e1a6-40e5-a546-f22ebb8d114b · inbound

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning cites this paper.

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:21.438917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:21.438917Z digest=sha256:e147b301065b35109969ca6f204e830a7e02afad21cef73101bd3db062e2358a

Observation 9ad7cede-ca98-4c61-8f14-ecbbdd2bc5ab · inbound

Structured Pruning for Diverse Best-of-N Reasoning Optimization cites this paper.

Structured Pruning for Diverse Best-of-N Reasoning Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:23.561140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:23.561140Z digest=sha256:db1733dd2b78006d9d8a3bc2f0e6b342e07cf8ae086ca7cc277d449fef38f728

Observation 650cc220-635d-40a7-87eb-790fc13501dd · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.180416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.180416Z digest=sha256:9098236cd44247a9e80351ce21b956f545ee0bd41ffa79a16f90b07d44843131

Observation 4234ff1d-a7b8-47fb-b9ec-4092ced0f56f · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.616592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.616592Z digest=sha256:fac34389807310850d4eb91345f98d6b8fdec8a0de0e74e97234efc9fe8208aa

Observation e67df85b-d344-4f51-8833-ae66fee42643 · inbound

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty cites this paper.

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:38:09.465060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:38:09.465060Z digest=sha256:63d710758126818cdf99a40be208d1792b745a9f4c648356a070b388eca361e9

Observation c54a79fa-7e5e-42a8-85a3-f7797b28d389 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.750335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.750335Z digest=sha256:89751ea5a10e9b2b37a463b03fbbdb4d7211f95e4909f44706989c30cabffd4e

Observation 2c47a8aa-1244-4b5e-bdc4-7b2deac56277 · inbound

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality cites this paper.

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:37:08.989689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:33:08.719028Z digest=sha256:4e4d7e5d4d94811cdbe0b1a7356ab44851748aeddba0f5f5a317b4b04ba531b0

Observation 6b7d3b68-72f8-48e2-94ec-d7f9494fb526 · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:14.714930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:14.714930Z digest=sha256:70369abd76f28b98351f6da1b2b9e75ef6a0804e18f59d51a6491c646cf146f6

Observation 5eeb45b6-3c6d-4900-93a7-03e4147598f1 · inbound

Reasoning in machine vision by learning fast and slow thinking cites this paper.

Reasoning in machine vision by learning fast and slow thinking Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:18.579086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:18.579086Z digest=sha256:8c35bc7b92bdb613d268441cb7502b4c6c133dc79075b8f3f8f15fa08b7283ee

Observation 2a1f5f4c-bd4b-4007-af55-873c9248df71 · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.040552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.040552Z digest=sha256:8739c6d279af975dcd24961276a77fab5c8bcbaa58f4b4c8b1ae24d5fb9becc5

Observation e1d90dd1-c705-42d0-baef-f00aca515744 · inbound

Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS cites this paper.

Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:12.161024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:28:12.161024Z digest=sha256:95bd24e334b709a89e58a3fcbbb75baa388b5c08b023bd01a271336b9be9e36d

Observation 60d6e2b2-8357-4380-9b77-2c4b9d9fc6b8 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.209380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.209380Z digest=sha256:6ccf9b84b32ef9817612bbd3db8ae581f80b5ff61091a9dbbed8dd994f280111

Observation 2a39a040-6db1-4621-9b31-ed1e1b145f3a · inbound

It's Not That Simple. An Analysis of Simple Test-Time Scaling cites this paper.

It's Not That Simple. An Analysis of Simple Test-Time Scaling Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:04.281051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:04.281051Z digest=sha256:9e440fce731a8c1fed884916276cf48a8b389b5547b0fee4e49047875d54683e

Observation 6bacaa48-b2da-445a-9bf2-4aa70ba0a491 · inbound

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra cites this paper.

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:42.237425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:42.237425Z digest=sha256:aac01b7f916e1d3cb1f47f0179cb7b3543687ee07af8622b058ff1cdd2258c41

Observation 84845823-462d-4b71-adb7-056d11f47575 · inbound

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs cites this paper.

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:46.343818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:09:46.343818Z digest=sha256:3b4f0c514fbac515a6c69e1e51a866f0cbb500035c41847023a43e94b8e61db2

Observation 14285457-87e1-4d05-a671-243df4b6e228 · inbound

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time cites this paper.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:16.716582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:16.716582Z digest=sha256:c3204c8b065a638b53440303c9515bbf352e3a74a3cd5856f98728ef09661474

Observation 398b7170-af66-43e5-b8c6-f39c12d228aa · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.654971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.654971Z digest=sha256:5f29d8ef5087e89953c38b7590f304f00dd0a767cc531248daa893961a8a499f

Observation 3f246477-bd08-4ba3-8922-04f6a7bc2519 · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:26.719841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:26.719841Z digest=sha256:4e92ec9bdac99698a56991d519402ca0f6667f0235c90d7f44b33f9dfbb9160e

Observation 028a9c3e-e00f-4664-bf40-d95bf6581261 · inbound

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs cites this paper.

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:21:07.901256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T17:19:58.334098Z digest=sha256:5fd8cf769450a22e1224319857316126998bbe5a421fbb060892304f768967ad

Observation 788d1018-d4fc-4a17-8e50-9b4dfe58040a · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:22:55.276086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:ff4c66f68105decf84005ad3a0f4dd0da6461bf72e4ab59b6820da6f72db3668

Observation 0754fcb6-bf24-483b-9a0b-0deb0bafddd8 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.052180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:65f77f5b13f0816f018ed6ef9fb21ae09b7fd4174e11510fd979d9607638dd1a

Observation a5728366-d70a-4ffb-9660-d96556437e88 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.542090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.542090Z digest=sha256:750b67739071611ea8db8bf386f991ac2c78ce1972cbe2ede5f6befbfb60d1a2

Observation 2a2df38d-1ab1-4573-b6f5-5d192859296b · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.722018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:2624e079375784a9d3f373a66c49acb8b517981170df8bafbcf4c7d1767e9fe9

Observation 761dd4d6-29a3-4eb1-8972-b07c89f37a19 · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.418936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:d2afd4c0abf978101b96053071b763759f3e4ff7a8a7b3bd5dba0dd30b3babcf

Observation 66be9b4c-024e-473f-ae01-4a23529ff515 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:05.589442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:9f83906adfa95a0432f3e1e49ca19540f107465a511bb4ff88203fc614f5c911

Observation 93f3611e-068a-4ac9-b5a0-fb844a575fd3 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:08.919771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:7278d70f2dc3674a5df9a20b7a586dfcc097e6458132799faa2105f1c8308d9a

Observation 64f389a3-ece9-4f54-8a47-654524236766 · inbound

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees cites this paper.

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:42.293790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T09:34:33.271537Z digest=sha256:d4f9e3b19f02bb62cad9babb71a8dd794bd079d8772a92d2ff7a43a0718b136b

Observation 5210ee3d-d234-4f0a-a335-3a0f50e898ff · inbound

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning cites this paper.

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.956367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:16:00.499779Z digest=sha256:68d0e610ceef528514f417316256c22c0857593e17ac576b6506e361646cd532

Observation ff54ab4a-a1f1-47c5-8e7f-8ebbb47b3abf · inbound

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation cites this paper.

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:25:58.264821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T02:28:20.366674Z digest=sha256:5fba40f00ce5e598aac5271d8791f92e08244e46a4e1fd66cf51087f74542cd3

Observation ada3a1b9-1026-41b1-9cad-17bf54fbcee6 · inbound

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning cites this paper.

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:42.171923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:03:01.612516Z digest=sha256:9ad38c49e1ec7b4e5b57bd365b8c503ad0096491772f0279aa48188f613f5937

Observation cb9cf998-12a0-426d-826e-0ab8d015ecb7 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.023684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:01726c25b47ceb5f36ee285d8dc323141c2a14d223162e0b5fa3907ed34845eb

Observation 4f34326e-528b-4d96-9ad3-1fc96a76fa5d · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.262296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:29f0891814b318aa16921ef5364d515a72d3fc280dbcbe0173cb02099b1ff63a

Observation 4077066b-ce29-43c3-ad50-fc741a014a0d · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:30.042468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:1c7f79a593974c1111964477da87078009dee8ae71b4c0c6a973f3740e783f82

Observation 9c889228-de92-44ce-852a-8375a79b6547 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.956800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:b06e76bcfbf33a595e9a6f4bf5780509ad59c3236fa150e2990d33bd450b315a

Observation 1cc95fe5-48ec-40d7-879e-988120a46c37 · inbound

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination cites this paper.

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.900391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T19:58:32.016341Z digest=sha256:abd34f548d847c94bf8cbcbeaf6d249543770aa4bb86ada397cf218ca2350158

Observation 3c60c9d5-bffe-4fc4-937f-e71ab1969215 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.750084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:36724b6c9f7a7d13bc44b37a51dc75a1d1e6a7b67ef3a98be2a8c1dd248a5ef9

Observation e84f6b6a-14d9-492c-acf4-7991eabb59ea · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.320578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:72f2e5793aa127e0a64510b0e5a88b8d5c86e767f07b98d4d7d4183adee1f7e0

Observation 956695dc-6775-41b4-930a-08ea6cbae14b · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:26.571855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:26.571855Z digest=sha256:8102d58097490e4a723340cffe4a0546318a947b42bfab9ad2f64970430c9008

Observation 26e4a92d-54d2-45ef-982a-7632f69be8e3 · inbound

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies cites this paper.

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:44:59.784865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T18:38:48.512777Z digest=sha256:7727688c9d65dc2da24808a16c0c94305d9f8fe4d96240b9996655f9faed01f9

Observation c83a8d58-2ff5-4cae-ac64-129a2da31a2c · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.805075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:ef00a39380cf5a6cd6ca1f5732e01d5af5104783f570f71603aa15e439bc2d36

Observation 9fe9a5a7-d48d-46cb-a485-730dbb33a8bf · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 165

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:45.694539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:45.694539Z digest=sha256:ce2fa6358d16b62bde7520490e10da23b1581dae008914d1f3f8ddb04a58f8e3