Pith. sign in

Paper Citation Record · LEDGER

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 2 inbound Pith citation observations for arXiv:2507.14295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14295 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:14:09.490536Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.854490Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:22:47.413783Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17eae788-e7b5-4906-b24d-02c7ea37ba91 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.177595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.177595Z digest=sha256:baffeababdec8ee1c4dae8749f6c8a68928a4c177d4271ad9a68f15e8d4799de

Observation ebcb6810-5c35-407c-b9ff-cab89baa01bc · outbound

This paper cites GPT-4 Technical Report.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.291984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.291984Z digest=sha256:fee4530714d195f8a02221871eb59866bf7d59a2e4dde5853d887cdc9824874a

Observation 3df69451-78be-4620-ae58-f933f98096bf · outbound

This paper cites Qwen2.5 Technical Report.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.416492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.416492Z digest=sha256:e6d8a58faa3f77f2e0e86e0af3d6abda900fcbd4bb34ff369054b435afd73a30

Observation 4160e1db-202c-4c32-ac9c-371e39e8193d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:02.648526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:02.648526Z digest=sha256:9dc5262f35fdd8d48600eb77fabfbd346325bd18bb4c741a11fb14c15090fee0

Observation ce9e97a8-c8be-469b-b310-bf62adf99a20 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:03.913893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:03.913893Z digest=sha256:fdf8288d5d606533e40b5c90f219ce1fdc647cba0e26a771983a91e07d143720

Observation 6de5f99d-b149-450f-a171-fb3dcc223bb0 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.049831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.049831Z digest=sha256:87370b12555420f4702a22eb107d7ac312a2cf1d6b11e72dfcc0a7c1e5728537

Observation 69b166d7-93b9-4d6c-8956-1bad9a1ff0c0 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.197657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.197657Z digest=sha256:db8544728791666b5d8565709da3e107a19dac2656b0e3bc7191c630cded09ba

Observation 7a687a27-e2c6-419d-8403-c693e0e42247 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.352421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.352421Z digest=sha256:6b1c015d6d3b98d5ec35df968842a11a3151c89a05088df065ab34c49eab5cd0

Observation 47cd9019-307f-493d-8c78-e3e39d7fd96c · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.482017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.482017Z digest=sha256:6f9037063214ed1261fda69860c46cfe66c085af9f215b5f33311e45fc33d89c

Observation 1cdbd948-6853-4fe3-8dc6-4761fff620e4 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.685622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.685622Z digest=sha256:4558d255f73e577eae2a82b533647c1a99279401f32ce0ed78dd3c56dda3187a

Observation 871f8b7a-38b5-4a3b-ba7f-27d4ff4df2bb · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.795371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.795371Z digest=sha256:2867488ed0a079bdc5e99e1cca58a7ebff9790019125e05d29989bab63e2f9ec

Observation 7e0cb956-51fa-47cc-be48-1410d9573aba · outbound

This paper cites MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.918820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.918820Z digest=sha256:250e752de7e0a4a0cd764e2377fd49792c51af7bf9ce011e296bec090561dc39

Observation ce9125b9-71b3-4eb9-9426-ede0b7403805 · outbound

This paper cites Simworld: A world simulator for scaling photorealistic multi-agent interactions, 2025.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Simworld: A world simulator for scaling photorealistic multi-agent interactions, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:10.660770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:05.133546Z digest=sha256:179a9a5394bc763b62329290e0c0673bf1e251b024004725bc9125ca0c86d77e

Observation 3ddd5e56-30c2-468b-955b-142c29dd4d9b · outbound

This paper cites Gonzalez, and Ion Stoica.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Gonzalez, and Ion Stoica

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.285192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.285192Z digest=sha256:7dc93dd920a77c670a29e0714209eb4adcc79b18bf708d83e474498a671e1331

Observation ebae5172-5e07-4063-b061-45322dd9e06b · outbound

This paper cites Training language models to follow instructions with human feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.398814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.398814Z digest=sha256:81c65ade064c62fce3c6780eeaeb1b85102e12430d1f9a0f57112948a5a89ea2

Observation 5ae237ca-edbb-438a-b45f-14e8dbdf840c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.544503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.544503Z digest=sha256:8baf7e4b92173d4a42944855b4b314e7071f79f6f01a1557e6d71fe530a81c4f

Observation 2cac3e1e-7103-4240-bed3-e2854c459812 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.665511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.665511Z digest=sha256:d53869697ed78c4299f73d35bfb7cc68628bf579107767af612ddc7751729aa5

Observation 5ae4401e-7a2f-4c3c-9d3a-980c350a22ea · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.766444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.766444Z digest=sha256:537a40d214bdc8b239811428c7aa93103fca5097711e442580563e873f23b9ea

Observation 8574136e-8b3e-4922-bc44-b2c01fd2e3a9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:05.889070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:05.889070Z digest=sha256:012787bb874e4fbacef6398c8c029acc252d291350d5398a132a1d96a7691a9a

Observation 298c3469-3378-4fea-bdf6-944685f4ca02 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.014532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.014532Z digest=sha256:9eaa36d682704d408479f5e9bfb4dbd7c78aa186ed328119baaa82e0cdbf4f2f

Observation a1fbccc2-743f-43f9-b951-47742968c637 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.124607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.124607Z digest=sha256:0929cee25551222f1dccdcb8734148cc0ec17dcacc01d364e5d1c5d5a764e173

Observation 181581e7-9293-4cf1-95a6-cbf1d8f68053 · outbound

This paper cites HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.230944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.230944Z digest=sha256:def12f0da7ac1925971a6c3d36b8add60be071e5c3da26811e2ce0631fc45c01

Observation 73dc1681-ec6e-4b8f-9a9b-84e118096ee2 · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning TheoremQA: A Theorem-driven Question Answering dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.336764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.336764Z digest=sha256:af994882b76ded12b32875381c56d790bb455f65ea04d9cf006142073ba3f2cf

Observation daabcfd0-3f11-408b-af82-a64257058c1c · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.442279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.442279Z digest=sha256:277547f94b6401cc0cb0ef75d52a0bbaa675fa9ba7e8729ef07cf97d4594bed7

Observation e41177a8-a2d0-47bd-937b-617b59d41f6e · outbound

This paper cites Reasoning over Public and Private Data in Retrieval-Based Systems.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Reasoning over Public and Private Data in Retrieval-Based Systems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.552532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.552532Z digest=sha256:1e5c5d0a37108bcf8924d80bbb6e18bf1ad7f075d15fbaa3fa4f07ea183e396f

Observation 49182346-dc44-4c8e-b8e1-319a2d0820e8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Measuring Massive Multitask Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.625049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.625049Z digest=sha256:9f24730ff3b315dbfbddc0246ccd3cf11cba449d51b2adb319512ecbd06a4cb0

Observation afae9fdf-48b2-4f3c-856e-656045a2e324 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.701604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.701604Z digest=sha256:66c70900ebf08dafcf02ced8e0ff0056f78afad26b533562bb1bcc171214e14a

Observation 951c7202-42bb-417f-9270-96f43c8cebe4 · outbound

This paper cites Graph of Thoughts: Solving Elaborate Problems with Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Graph of Thoughts: Solving Elaborate Problems with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.824346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.824346Z digest=sha256:12a532cde4abe1b91f351df75c0ede947f4970097928dc6f50348b255f3377d3

Observation bc84f671-1e7f-4305-8939-c34d514f0e3a · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.901577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.901577Z digest=sha256:e36e283d6c8ebc5b4c2860f342bf48d0c10cb709f1c8f73be84b0e9cf3d57fbd

Observation 5aeb5db2-1b6c-4c40-84b4-26c61b1f9262 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:06.999304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:06.999304Z digest=sha256:dd51c7b704863438a1ed4b9282c5e0814656ebd734ad0f3332e4d1a2b3ba3099

Observation 7d33ccf3-a93e-4e24-98cf-a71390969d14 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.142246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.142246Z digest=sha256:8717cc611e37aeda5573284cfc615d072d6dbf5ea1f2af97ab525bb2d303a843

Observation ccc5a4b1-fbe3-434b-bd90-85dc181f763e · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.252826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.252826Z digest=sha256:3c7c7b6e370171ddd4f29d65bd0fba8c8e2f6ddbaf6cc361330e06bcab8dde5d

Observation 3c18fef3-5c6c-42f0-bad3-34cd8047f5e7 · outbound

This paper cites Large Language Models Prompting With Episodic Memory.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Large Language Models Prompting With Episodic Memory

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:14:10.255823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:07.377051Z digest=sha256:64d7dc4f743e343b3c9c29afb27c41ba976e49fd74eb82053be1100541197e96

Observation 3cb608d4-b8b8-4cbe-9b45-b488bf86d5c5 · outbound

This paper cites Larimar: Large Language Models with Episodic Memory Control.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Larimar: Large Language Models with Episodic Memory Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.482480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.482480Z digest=sha256:dac465d59c1510e2458a37faebf468de43ea8f63752c6067e01e8a364a7bcdc6

Observation 4344a43d-7ea8-4d60-8166-252265d65b68 · outbound

This paper cites Deep reinforcement learning from human preferences.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Deep reinforcement learning from human preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.581685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.581685Z digest=sha256:8a74eab808770fdab61158e0509a23c6a51cade33f98f72a7ea0b23a4c434609

Observation 2e6002a7-da98-419a-a470-c755754598ea · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.594492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.594492Z digest=sha256:cb3017629ea492162b4b810fe9181d90daadc4f53fd7b4be6ce6515204a0e017

Observation 48114e8b-2682-435f-b466-1d18e8140ded · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning On scalable oversight with weak LLMs judging strong LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.651045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.651045Z digest=sha256:06a13999ad2088ccb460e78aaed95ae0378d63a37c2d32125ca611bc1d040c69

Observation 194dec1d-eefe-415e-b1a5-da239a12b62b · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.746870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.746870Z digest=sha256:f13068ea1d71ecea00902c6c8c41ddfb5555852109a09d85253c5c2a6f41fa40

Observation 4ac7c2ab-2fef-4fce-a820-1cde53e3f2d8 · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Parameter Efficient Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.820412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.820412Z digest=sha256:43288a97b910af031e3485ed3305babca5210d8022ad37c0dc15690d1b8aaf1d

Observation 10ca8d94-9bbf-4789-a5a4-e5ad3429da5e · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.896851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.896851Z digest=sha256:47362533a2a44991a93d3be0f1ff6515fc4f622d86411c9622f708029167f531

Observation 4fc4a1fe-2bfc-4f79-8670-f9e59b038f75 · outbound

This paper cites UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:14:09.915290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:14:08.019297Z digest=sha256:4677b0f4676328c8c6ee397a97f5c825ca8cac6126ac9f94b38590945fb93285

Observation a65de3cd-6f1a-45ac-a9de-7ab66f3fa671 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.106325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.106325Z digest=sha256:602a04360f24171fa6d8b612dd5c042aacb5ff14559b11ed28c720998c3c8bda

Observation 767d4daa-4474-44d3-9b51-91a5bf50c929 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.189265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.189265Z digest=sha256:19de1685b216a74ed842c9a17dd8e70d6a47a26fb1f17dab3413ff0847cf99e8

Observation 88d476c0-e02e-47c5-a815-cd4a96ff86cf · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.303551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.303551Z digest=sha256:e391416e15eb48bb2a1960596a73361b334358ae690e2fb971c9e7feea199f70

Observation 6b3fe19b-6338-4233-8b83-17bb197f16bb · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.374932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.374932Z digest=sha256:808d57c27aec97eb40c1808f84fb60263fb479fe1a00131fd76c38910eb55925

Observation 7b333d75-1361-4f5f-93e1-2b7aa7042db9 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.466387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.466387Z digest=sha256:fdbb9a1352e756fadf5f6264b5332bd05299f715171b57339a122d5291299fe2

Observation 86811f4b-5b09-410d-a948-d896f4aa6c56 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.565734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.565734Z digest=sha256:e89f875acb558b1751c60d25aee2606f6ca974cae6951399d7029305c8760995

Observation f9866ada-72fb-404f-a504-70cb42a61217 · outbound

This paper cites PAL: Program-aided Language Models.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning PAL: Program-aided Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.667441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.667441Z digest=sha256:68b7cd917054b21b7e167ee48b6801ce602ce3f7d1ffabdaa16b609751446f76

Observation 5e7c3e03-f7f3-46b8-bced-357d7694b112 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.740558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.740558Z digest=sha256:94e7206e24199eb52c411084130eb9c22e5a97d6687db573078de35719fb8f01

Observation 15b7b6c7-8e84-4df7-a904-398435224e19 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.865042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.865042Z digest=sha256:0643221e0da8ad9800b979e62047073c765e16cb38adab5230f1116c2159f4c6

Observation 4413c3a9-3758-45c7-8799-5a54cd77fcc4 · outbound

This paper cites Let's Verify Step by Step.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Let's Verify Step by Step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:08.959676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:08.959676Z digest=sha256:b883c8a2c63e91a6c9d0cedd8c05b50365dff59a3b8d02d70cf5a65a89313657

Observation 144f1987-5e3c-4b69-8ac1-a239733e5a99 · outbound

This paper cites AutoPSV: Automated Process-Supervised Verifier.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning AutoPSV: Automated Process-Supervised Verifier

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.056162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.056162Z digest=sha256:0a3dfd55cfd0ac94154f3e4c8640b07c87772c7afe0eca900716bad76f0892c8

Observation d16cba06-c5ee-465b-89a1-82f9c7880e9c · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.175411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.175411Z digest=sha256:a42642571d9d674b341b7a8329aa68e5af2721dfc5ef1d5c20a235aa7b78c941

Observation 972a76c5-c47e-4911-9642-3e5d1b12bd12 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.245253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.245253Z digest=sha256:53c3675f366f6cd5caac9134e85294b4925b54039961d597bcd5aacd7beb2ed2

Observation acf55dc0-563b-4636-8500-950e53af27a0 · outbound

This paper cites Ursa: Understanding and verifying chain-of-thought reasoning in large language models, 2025.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Ursa: Understanding and verifying chain-of-thought reasoning in large language models, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.305497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.305497Z digest=sha256:71e7f3bcde8d1f94978391a6650a75b185ef6052f8f52c2a5b4cbf79a6da3a2f

Observation 35ee2956-b30f-49e9-a549-c2ea74da1a31 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.393008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.393008Z digest=sha256:2be8f9af55a6989911e2d15cffdd1615229f5801f7dc0d82b16662d9038c266a

Observation 1dfd7eaa-d8bf-485e-9537-a01cb4c3251c · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:09.490536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:09.490536Z digest=sha256:abf5254b63c71f5ec358deb2483b9867745d91afaf20c32ac9c5c296f669a2f7

Pith citing papers

Observation f75ca2ed-f3e2-4825-8c23-754567efa4a3 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.854490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.854490Z digest=sha256:66b304b9d499a353f006ecd6e6b2f21955081de8410dcdff188263ed507c53ee

Observation 223af9c4-7897-4da4-9fb1-83288a4d7796 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:47.415781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:3d8ca8a15d09bd7b99f8a68374abae27a5907181f92f95b4a5f951086fee5125