Pith. sign in

Paper Citation Record · LEDGER

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 12 inbound Pith citation observations for arXiv:2506.08745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08745 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:46.121588Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:00.494940Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:51.882252Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6ffe7a6-8cd9-4dea-9d5d-dd5cad7aa535 · outbound

This paper cites Kakade, Jason D.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Kakade, Jason D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.922731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:42.214409Z digest=sha256:3d780a135830e215eb33996d6993268d11277e02e7d57c74969ad50960865027

Observation 33bb0019-910d-4abe-974c-3fdef402e725 · outbound

This paper cites Open r1: Evaluating llms on uncontaminated math competitions, 2025.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Open r1: Evaluating llms on uncontaminated math competitions, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.913745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:42.272800Z digest=sha256:926fe07262e56fd27f203a886b6419a0fe6219ebd65be1292f3dd12ece189b67

Observation 349a80c7-031e-4bf9-b36e-80b6ca054ddd · outbound

This paper cites Program Synthesis with Large Language Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.339097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.339097Z digest=sha256:0f48ccf1722b8c3954fbe5d43e96182178343fcfbd6ce618f4ec0563b801e380

Observation a028dc47-6b3c-461a-b440-d01b6c095bd8 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.395768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.395768Z digest=sha256:8fbbaa6d4f739e29da5393126621abeb16906887196d01ff212ed154d2e62390

Observation 126ac31b-a149-4f20-a242-22e69feca703 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.459366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.459366Z digest=sha256:7438ba1af4f4c19efb902824a713c37527fdd33b36bed547db3b4f3d35844315

Observation f7a66c95-c231-4ac5-800a-9fb4f645469d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.527494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.527494Z digest=sha256:3febcea3d89269eac481ae56bb75f2a5203f99678190933a6533e8b054be9a42

Observation a96f279b-e9e2-4748-8db7-cbbe8ca149b3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.531213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.531213Z digest=sha256:e085f2c50d64ae6f94f95b8c7ec48d6e598e2d6254f7f020af037d915f2ab147

Observation 6ace5b57-f3d8-47b4-854e-7baf906a026f · outbound

This paper cites The Llama 3 Herd of Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.533735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.533735Z digest=sha256:3fe0c1bb2b79901cbbcfac7cdc1ef2fb73a17575214360f3f07eb1a38180ac43

Observation de5575b2-157c-4405-a3db-3e040724983f · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.537444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.537444Z digest=sha256:39a15355859a24bedc26574f1368417491fc9d7fd55a853c06debfa9d173b070

Observation f6f8b083-dc55-41ed-aaa8-cce6433666c1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.613608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.613608Z digest=sha256:5368603ae84100946869685f31149ffd073dcccfd88c3091d41a98868812ee5b

Observation 740110bb-d645-412c-a1e5-969cf9887c27 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.903681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:42.727320Z digest=sha256:e1b40331e665ee12e4ee8125f7f666dd9f340d411349cad25b9270ef5dceebb9

Observation 47db5fb1-f488-460d-a6a4-2b79b25aeae2 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Measuring mathematical problem solving with the MATH dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.895231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:42.897467Z digest=sha256:63fa644731947e51139423f95c75436b7e633bfacad91c90cd892029b170f568

Observation 8f58d985-37a9-48c1-be99-122ef30061de · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.068905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.068905Z digest=sha256:6da54b98fbc503a68d0dfd6c4203fd7b95117ba3017c087bb9268dc98bbd18d1

Observation e1a9ab83-840e-422d-820b-e131a5396b53 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.246700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.246700Z digest=sha256:0a0b436c6dd2c8ede43e1198f242a517a249ecc4ef5ae430533bafcede5337f3

Observation f59237a1-bee6-407f-a99b-ea33077367e3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.415963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.415963Z digest=sha256:66a739559aa58e7a418df8a5e84ca483ce1b8cdd11e9e3c251ad44058a3e37c9

Observation 68bf5204-e622-4186-a76a-b84f075529da · outbound

This paper cites OpenAI o1 System Card.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.589672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.589672Z digest=sha256:510babe7d1e6f530e53f3303f2647a04871f7ae8fb6d7a9bcaf7abb45ac20498

Observation 8d0e0c34-3cd6-4d97-ab62-faadd798ef38 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, 2019.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.886118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:43.711859Z digest=sha256:9f8acdce2ee6ada33108cf476449324200c2faea58fa0218b16fbe3070617ebf

Observation ba7da1c7-439b-41fe-a38b-b24b3222e832 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Efficient memory management for large language model serving with pagedattention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.874897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:43.850545Z digest=sha256:fd12e611ed499015be6115a901cb92ad16b88aa8c50a38385be3860e0baac0a8

Observation 1d58092d-d4fc-4dee-97c8-cca1226af5c6 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.878319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.878319Z digest=sha256:a9afaceb825f4e429fd58ef00a7b3dee11ec3d35250de01756c7f069ce03b782

Observation 0d523d22-44c7-4159-8605-808ce22af111 · outbound

This paper cites Numinamath, 2024.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Numinamath, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:43.926312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:43.926312Z digest=sha256:ad0234e86310dec0b0808e25925057e19261386d8db55f4c04457f3478866c17

Observation c71c6a20-c3f8-4038-9297-01d8d6730a0b · outbound

This paper cites Let’s verify step by step.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Let’s verify step by step

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.860484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.039561Z digest=sha256:364fc1e6ca11487c9d2c7c10ffc0589a8e01d40033ea207e97398e74d4154c58

Observation e1c935f6-3a28-48d5-aa9a-b50a8efa91e9 · outbound

This paper cites A Survey of Direct Preference Optimization.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A Survey of Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.129023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.129023Z digest=sha256:9cfb940a55d89089ff9ec32041e8474b42ccdd34f12109cecb944b672574a898

Observation 59a01924-d375-4698-afa9-7e01cd44e84c · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.200132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.200132Z digest=sha256:d74e6b88a75bd40d88aee4eb7ab2d4eda9a9f8fc83c23e4de9fc30d5c28af438

Observation ac2164ab-115b-4381-b36e-a55838303da1 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.269991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.269991Z digest=sha256:69947b6d8aedd4d5f9f643d45c32ab0ec2ae324bd558043d30851777c5c74be4

Observation 85670879-67ff-43f4-9d9c-f41cc542d951 · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning On the global convergence rates of softmax policy gradient methods

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.851299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.340424Z digest=sha256:4c6cfe917eb9da9fa805136ec2d702635c3f8b9e08b143f68677b4f326a47323

Observation a7d0a55c-b43f-4d58-98ad-299082650a33 · outbound

This paper cites an unresolved cited work.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:09:46.761702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.397534Z digest=sha256:b3683862f82ad7d47f179fb98bd41d32d3d6b90f2f77f64146313e7cb74f854d

Observation dfad0528-64af-4b13-bad9-d394c6b8df70 · outbound

This paper cites Iterative reasoning preference optimization.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Iterative reasoning preference optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.751714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.452991Z digest=sha256:6b01ddb1614a52482d7781f0edc208be0e58351d34fbc6bcf95cce2ed7b22041

Observation 0d617e2b-74d7-46c1-8a74-dda527e8d774 · outbound

This paper cites Are NLP models really able to solve simple math word problems? InNorth American Chapter of the Association for Computational Linguistics, 2021.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Are NLP models really able to solve simple math word problems? InNorth American Chapter of the Association for Computational Linguistics, 2021

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.741637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.519249Z digest=sha256:9aee51646565274f6acb09c09de6a088b806e430ced890c970b6d4978971b1ea

Observation d553303b-f8d8-4160-b70d-91155988e8dc · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Manning, Stefano Ermon, and Chelsea Finn

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.731168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.589116Z digest=sha256:7b6fd82c50dd9aa45732f6089cdd7dc8579deef50be26b239cd4cfa4b4e335ac

Observation 0fab6a5d-2b04-45ea-97a3-48a1e71374dd · outbound

This paper cites Semantic cosine similarity.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Semantic cosine similarity

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.722065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.644626Z digest=sha256:16e92beb3f7c0071eb4b0bf15224a021f9fdbf072892fdf27dd4cd018d34cb85

Observation ddc33db8-3652-4bdf-8cbd-be304684d6a0 · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477, 2025.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.700943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.700943Z digest=sha256:57da4380ec99e43ffea6eedb29227ce9a85083634967d4de452206734f5a91ee

Observation 7b4faffc-9300-4686-b4b0-3249f98dd90d · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.769436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.769436Z digest=sha256:dcb20e970c602a7cfb3e2a8c0adadb2f6382c8d4b5e9e2aa7e40fa9e8a384341

Observation 6d3ca9b9-f63c-42e7-bb4b-d15eb909844e · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Gpqa: A graduate-level google-proof q&a benchmark

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.709087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.827505Z digest=sha256:3eb88160f25b3b8ee4bc65d4163c0104cf1646eb5168ff2bd4f6f5fa020978bd

Observation af519356-a0cd-44dd-98d6-0ae87e455b75 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.871042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.871042Z digest=sha256:a4e6586e2f541e7aec4e9b791e20f0e8b37a1e2aa956750b5838a7d107aa3e71

Observation bca96c70-9f78-4100-bbdf-d6bd0ef41128 · outbound

This paper cites A mathematical theory of communication.The Bell system technical journal, 1948.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A mathematical theory of communication.The Bell system technical journal, 1948

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.698740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:44.951161Z digest=sha256:1858f5e4ee985a1d8ef8b54b1a3fdf4199b322bf442030fc018e4d75881ac7aa

Observation 6ce762ba-95d2-4297-88e7-3e04074a7a03 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.000333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.000333Z digest=sha256:d55c3408c369d6e625d985fa49c89ca662e307425e6eb99f3c0deb2a0b71b932

Observation 89f5da95-4b79-4e1e-ba65-f83acba760b6 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.062113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.062113Z digest=sha256:299a45266fe8a06a5dd1f4e1013901bfe1c7713cddec94f8b5a9c5f204d29d69

Observation 255ae497-a8d3-4e96-987b-23795bdf76c8 · outbound

This paper cites MIT press Cambridge, 1998.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning MIT press Cambridge, 1998

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.112502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.112502Z digest=sha256:3be9680150e316c9a8b268f4430c324bb5db9d47220df6b662e78953197d032f

Observation e5a48ae3-e6ce-46a4-9b33-513fc885edc4 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.684681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.174942Z digest=sha256:4d691f8989bc22ea1b14a894f18ecfc2d05c24a43368e179daf67d21e46f787f

Observation ed5c02dc-80b9-4fe9-9eb3-558a69b6038d · outbound

This paper cites Le, Ed H.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Le, Ed H

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.675301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.241174Z digest=sha256:3591c898dd75b76ddf786671d7813a009246c91e18b7b6514b92fbfd527b2edb

Observation baf741f9-462f-4292-aa7b-819939d4589a · outbound

This paper cites Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.364256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.364256Z digest=sha256:fe8dae295c798a5964b41d499708ee3b1537ce0aeab2811af573c010769ecccb

Observation 706440a4-c0a9-47c8-96af-1477ac5189b2 · outbound

This paper cites Wong, Zhuosheng Zhang, and Rui Wang.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Wong, Zhuosheng Zhang, and Rui Wang

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.665635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.411810Z digest=sha256:bad07d7150d1ad83b6a56e793aa1d3db489029e978ac4d6ff220412639dbddfe

Observation c10d9300-976e-4ae8-8d9f-0ba9d28f4d26 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.655371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.462434Z digest=sha256:9e000c9d9f59b46f3bf48fe5d2dff94b5f5e3fe46f960013f5d58ad42e3beac6

Observation 84ddd584-f7dc-4d90-beb6-3439d62ea7a8 · outbound

This paper cites CREAM: Consistency Regularized Self-Rewarding Language Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning CREAM: Consistency Regularized Self-Rewarding Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.503921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.503921Z digest=sha256:541a29f93a24355212b6b1ad3b56a7b60ce5a273efe31d4c661494e946a9e41d

Observation 89ff06b2-70d6-4386-b387-4bacf9fbf833 · outbound

This paper cites Chi, Quoc V.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Chi, Quoc V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.645771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.544523Z digest=sha256:8590ece8ee721c512d1df9a2c6b065f70278e47c0b7ed3631c63566f6a93f42f

Observation 1fef196a-87d2-4433-a6a7-37af59368ef7 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.585513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.585513Z digest=sha256:df7511ebe5a721ccf925619c21f4239fb18194b75ca07df7dadb07db20613efa

Observation b844e0e9-3376-412c-89aa-c3ac0f50d2bd · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.619823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.619823Z digest=sha256:e53f546c444eb3379d75f813f8cf1ebdf6496a8b707adda09325aba49b689678

Observation 3d8ec851-bcbe-4610-825e-f5be039ea645 · outbound

This paper cites Self-rewarding correction for mathematical reasoning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.657899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.657899Z digest=sha256:5e56093759b30eb5ec53a1dc31f4b009a5d6b610c14e758cd3672b9857289343

Observation ea577258-48cf-4c2f-b79d-f26916e0f3e0 · outbound

This paper cites Qwen2.5 Technical Report.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.688442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.688442Z digest=sha256:bec8137a491be785d03b5f62856183545a004644b7f294da28ae9c4172e01a39

Observation 56f7a275-2495-411c-a8cd-15faa612de94 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.745522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.745522Z digest=sha256:d1c1d227c17dc0986081a0cd28482839f65a820c15843fe2c041759e15f87b5c

Observation 62e43cf3-1427-4cee-a808-e6800a55c9ed · outbound

This paper cites Self-rewarding language models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.635670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.777807Z digest=sha256:2cf4bfc54feaf46a7d61940cb6ab87764c7d54041a987ce0e8e7fdab029691dd

Observation 85a54cc2-bdcb-4eea-9f1e-d18d4b028f3a · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.798657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.798657Z digest=sha256:ea4a9f2dab8589404f71d9a7f3a2fb20d3a77eebdacb5f63dd4ecc5fdc459d41

Observation 6e11b935-4602-4260-818a-7337de2eb813 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.826337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.826337Z digest=sha256:e9b12b44c09b3f190a3e0d77e05681a09f9b84a460b94e692280f9bea4fbfc5d

Observation 5d2ffa49-c387-402e-bb6f-05bb86a0dde2 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.849450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.849450Z digest=sha256:f3dc1040a34901d2b4d8ca2f13b659af482cd4d7971a64b1711938bfeb248cdf

Observation f99f6ab8-b206-410d-9992-fa579703938d · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.877654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.877654Z digest=sha256:385347a3d255cc77af077ad4c8c338fb77f644f903923b688195d425af18ac0d

Observation e6ca236c-8010-4364-b52b-f308d2ab6a58 · outbound

This paper cites Sample efficient reinforcement learning with reinforce.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Sample efficient reinforcement learning with reinforce

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.626201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:45.898683Z digest=sha256:6408b280fc9e3c091668ddcc20fe5b7371fe1f87725196f8270a08ec814bf4dc

Observation 871a9734-5061-49f7-98a0-123e3038f3d5 · outbound

This paper cites Reasoning with Reinforced Functional Token Tuning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Reasoning with Reinforced Functional Token Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.933741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.933741Z digest=sha256:6fe81a4da8cfcdd8573e0dedce63e20b2c1293655e91e24734e6248143d49f83

Observation d96c8c71-2b16-4466-af1d-22da45d8d444 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.969598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.969598Z digest=sha256:38d47e7a387799a62588a824a7cbc661b267450d76c7f4b01db49fdce0d428d5

Observation 30fd01d9-4a5a-4067-8c7f-6140ea62cb82 · outbound

This paper cites Process-based Self-Rewarding Language Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Process-based Self-Rewarding Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.000207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.000207Z digest=sha256:bce1dc194618b4ff6ae54ce2f51d69c79c378e8dfa808f699e5451423b7b3b4f

Observation 97b1f075-2d3f-4a02-bc76-f2266f53f3a7 · outbound

This paper cites Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.031769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.031769Z digest=sha256:0b477fd69b44d68a9fd26e04fedebf79a334494a99ff7f2446677d7951bb8b78

Observation 0f1fcc14-1af2-4ead-8c0c-fe95686c89db · outbound

This paper cites Calibrated self-rewarding vision language models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Calibrated self-rewarding vision language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.614123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:46.063365Z digest=sha256:d52a509c1daa68811c9294c97a22fca557f2454445a95e15b6aa398d98e53813

Observation b4962ff6-9a42-4688-8290-a53355ebb185 · outbound

This paper cites Landscape of thoughts: Visualizing the reasoning process of large language models.arXiv preprint arXiv:2503.22165, 2025.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Landscape of thoughts: Visualizing the reasoning process of large language models.arXiv preprint arXiv:2503.22165, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.096209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.096209Z digest=sha256:2c833899939c5a1885c453d8634f450bd1e2b8e311f0338a1d4fbd3b9f248b21

Observation 2ba2c7f2-b2cc-4807-8c01-8a8bf8181b62 · outbound

This paper cites Texygen: A benchmarking platform for text generation models.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Texygen: A benchmarking platform for text generation models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:46.601801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:09:46.103070Z digest=sha256:88bd7545d60292ec44099947b78249a287567818f21856e272b3ee67b7cc7597

Observation e93e34ce-8ff9-4559-8d0c-652a94ae0a90 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.113846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.113846Z digest=sha256:311d2c4e2813dbf8a841d4670f63f7e01fc7b0a660eac69b9e33793e1e5cb5a3

Observation b2d13d38-dc1c-46a8-8c78-dbf3c121cc03 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning TTRL: Test-Time Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.121588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.121588Z digest=sha256:5d970f6f26d69d10e78a2dbb48d75d2670348ab680e2ef88bcf5559c4908e04c

Pith citing papers

Observation 29319d30-5e71-4196-a3bf-7d2e642e1d67 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 180

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:14.968757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:e873f98f7da7dcc6be151ca9143eae2285ba644ec873deefff864f40b4758572

Observation 9cadc56b-68ff-4aa3-b874-4412ffc301f0 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.494940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.494940Z digest=sha256:e215aad88fdd4ab54da46107ac5659812ccd46c02f23f21bfb2fa58719682718

Observation 5ecaf77b-8d39-4576-b78f-eca6fc09c686 · inbound

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training cites this paper.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:20:57.469347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:56529573ef729783c372d4409c7cf821b4ced2be2bd82ec218eb12d4605d464d

Observation 24083e58-3a7d-4795-ad66-57b4b709d172 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.961406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:990c33e783852279bc4efcbd6fbbd069a26957bf9d9f47dd1c9e35f71ef95b9e

Observation 10de83ea-bbcd-421c-9cd4-af79fe64ff8e · inbound

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces cites this paper.

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:03.819469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:04:02.263300Z digest=sha256:ecea2ea933eca2115d1aef8651d0a7b0d599f67ebe1a5c9e9414721c3934151b

Observation 2b93e0d7-a109-4f6d-b3a5-c4f04b6f73ab · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.335231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:ee488c9d822ca9ee3384a2ff5a5d94871c29feec0fff3f220009237a71713616

Observation 87aaad09-9d25-4276-aa0a-da8b8d618830 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.369084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:bef9af6f759ed77872ee5e1276d9e0bc0e2752f9a38f6bfea586e461f70ef693

Observation cd01867c-1484-4581-9ce0-6109352f8bd9 · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.234876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:82f4fd598d1321187499f9edd48e7830e90c352e06fc1e4c2a45d6749f9cddf8

Observation 3c3e53ac-6859-4942-a940-366d111b5910 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:40.790106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:2a9c8f6de55e8fa8f44bf1ea0b573db5c3a585a0bb1284d78cabfc23c41a5e7b

Observation ed2c3dae-401c-4cf9-9c43-135827c3186a · inbound

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment cites this paper.

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.659582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T07:15:00.398111Z digest=sha256:6dfe6612253d895731c3fc39900cba2e5aa234504c58f9af5b84b56fd7116cc1

Observation ddbafa54-e33e-4dc9-9f48-ee9e5c8f1759 · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.883679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:083216d2514ee0064148ffb0f8acfd35fdacd1cbc96d115ad7aec6edd07f3331

Observation 0c0481c5-6659-4e70-a0b2-28879ea86c06 · inbound

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation cites this paper.

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T09:29:14.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T09:29:14.105502Z digest=sha256:ab5adeb89a75dbc1d3a531fe64689c13239fa52bbdb7ffc53745d021c0e4edaa