Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2509.06024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06024 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:42:06.520147Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22cd9d77-5bc3-43bf-9b6e-012dd366a55d · outbound

This paper cites L1: Controlling how long a reasoning model thinks with reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL L1: Controlling how long a reasoning model thinks with reinforcement learning, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.155669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.155669Z digest=sha256:9a82276757a22a5053ad5a060d02a08d6181f736c11440a65602b77302bb41a7

Observation 2c9f5930-0553-42f5-9e99-0b969903eda9 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.213407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.213407Z digest=sha256:1b81e2540b71c0bc2fe657c0932d04cea5b5919eca87494c7abc2dda3fac3069

Observation c27a294e-71ce-4d0c-af54-188db47e8299 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Evaluating Large Language Models Trained on Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.231836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.231836Z digest=sha256:7f506c4fee3e2c20aa8f08c2ba5b029a861e856793e40609829203fe4134f40e

Observation 69fdc157-bbdc-465c-b495-18e1571d5675 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.339231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.339231Z digest=sha256:2c86f654fd7aa7fe210513a5dbf3a0cd457e938410aa855d402294ad94e53406

Observation 7e793c40-6723-4edc-a07b-443d5e24281c · outbound

This paper cites Empowering LLMs with Logical Reasoning: A Comprehensive Survey.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Empowering LLMs with Logical Reasoning: A Comprehensive Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.386445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.386445Z digest=sha256:599619e7c060abcfed5a1a942b39fd9eadef96f565c602aac920579042dd0aeb

Observation 212bd76f-6364-41ed-ac58-8ef82c2f2f98 · outbound

This paper cites Self- playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Self- playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.796428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:05.458658Z digest=sha256:f942b8a1a0a9f31c64d721d55a2c55f347ec2e6897f7aa24305b7728c586e7f8

Observation 0e1d43cc-a800-41fa-9bd3-ba158b18055a · outbound

This paper cites Transformers as Soft Reasoners over Language.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Transformers as Soft Reasoners over Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.511121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.511121Z digest=sha256:cf4f24261453dee4e09904ad3f117367e9ec0da3e4f1c46ffbdf1a01f68749cd

Observation 0a96dd2b-7f59-414c-b992-b6f92d214188 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.571433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.571433Z digest=sha256:c157181d321f3e97944cb0ab2f3545677682d5cf15c8f4d558979affc23753fd

Observation 09d5c820-15d6-4f1a-9671-70f7e1198aab · outbound

This paper cites Gemini 2.0 flash thinking, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Gemini 2.0 flash thinking, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.669140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.669140Z digest=sha256:2d65ba3fa2ec2082ba8e542b2b6fc8356565596ff72a47aab507fd6ceda0087d

Observation 5269d331-add4-4aea-a46d-b61819dccfba · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.728700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.728700Z digest=sha256:4942c0fd0075d955ce3e8f6f4291b66da26d92a0ac6fc441b23f9379c8c55079

Observation 9a8ec324-7c9d-4337-b5a0-2ec0abbe9e96 · outbound

This paper cites Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361, 2021.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.798953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.798953Z digest=sha256:1b4064b175f2b00a021bd46146600ea01ff242e5aa27bbaa733b5154d08c95ad

Observation e546555a-3bc0-4ed5-b677-ec87a70695d5 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.864107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.864107Z digest=sha256:3fdef760e96f316502b0eaed474e122610e5a47462e38d5fb1f9f2a0fc1689c1

Observation bcd6ed91-71c9-4de7-b825-c40330044a9d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.932027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.932027Z digest=sha256:1497689f0c1c00fc0cb8694649f1e3f1204d68eb006422377961a1d28f0adcbe

Observation c1d3bab2-ec8e-4b06-b5ea-c4ff44d58900 · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL FOLIO: Natural Language Reasoning with First-Order Logic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.003231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.003231Z digest=sha256:954fd7b92b789199bbfa3237f177384d0a58652b041254830b131738a3bac8d6

Observation a19a564f-598b-4599-a236-031e91b38bdf · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Measuring mathematical problem solving with the math dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.068109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.068109Z digest=sha256:e1160659014769ab444f16fcfdd3cf3043eb9e7cba6f190d7c3da08ca74783c5

Observation 355d0c2c-b29a-472a-a57c-407d9fa9959e · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.155144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.155144Z digest=sha256:cf90cef26d9a0deba534f040854425fcbe33160b35b3226e045e97c2c9d82275

Observation 18da2e70-102c-4c7f-8c81-4fd3f7fa4496 · outbound

This paper cites Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.221502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.221502Z digest=sha256:5ee7c3fa459da09b3d73c77227385e5e1d75856007a3d3399ebdcbd062cff90f

Observation 6442c8e8-9681-4416-9224-33767f2bbff7 · outbound

This paper cites Boardgameqa: A dataset for natural language reasoning with contradictory information.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Boardgameqa: A dataset for natural language reasoning with contradictory information

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.734162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.241979Z digest=sha256:7eb66d446672944327d4dfbab2187725bdf9de5203947dd4f3e465f329db56b8

Observation 2922c9ec-3339-4875-bddc-725e62d6c14e · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6, 2019.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6, 2019

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.704465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.247220Z digest=sha256:7e8112772a67e6ed3c1923e204dc797e0a6b461f0bc37b2a572baa59de4a2f28

Observation ee6cc1d8-bbdd-4b9b-9e66-7843aa934929 · outbound

This paper cites Chain of code: Reasoning with a language model-augmented code emulator.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Chain of code: Reasoning with a language model-augmented code emulator

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.678037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.254413Z digest=sha256:67c6a61d0c13127fd33c6cc776df7ef0a92e4df1ff037da4722ee9ceafe83f07

Observation 03c8f9e5-1a42-40d5-b2da-16e5218e12b3 · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.261278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.261278Z digest=sha256:835683532477ace367b70e714c8648e20825848bdc39964ecbcd5e871c4d9a0e

Observation 0a7f6c22-c1a8-4074-a35e-e00a8c6b8587 · outbound

This paper cites Limr: Less is more for rl scaling, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Limr: Less is more for rl scaling, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.266810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.266810Z digest=sha256:5d662c95e52cc6aa6d07ddf62e51c1cf6c90ea47e8d43fa970406e498960ca77

Observation c4c61012-3a83-4b03-9eea-bf583e5e64ee · outbound

This paper cites From system 1 to system 2: A survey of reasoning large language models, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL From system 1 to system 2: A survey of reasoning large language models, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.644307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.271252Z digest=sha256:8cfa085ed425144c50768fc57777c1e57e2f028e0eb390c7ac74716c1aa20dba

Observation 29969c6a-139c-4e09-a049-55ed9bb828c8 · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.276239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.276239Z digest=sha256:3a789570caa407ed9cbf36b10a7f30c8b50cc004dc15ab8b1fdb29cd54d4e287

Observation 8423668f-70ee-447b-833a-9dee553df8a4 · outbound

This paper cites Natural language inference in context-investigating contextual reasoning over long texts.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Natural language inference in context-investigating contextual reasoning over long texts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.628502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.280844Z digest=sha256:92961a6cb6f97d2c7d87979768afebaaead8a5162f8c7f77a1b9d7cfc99806a9

Observation 3e1a4b64-531c-4654-89b3-9a8337c56d83 · outbound

This paper cites Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2947–2962, 2023.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2947–2962, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.610533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.285793Z digest=sha256:01bc90e17489ac2af88e2bc7959c7a8cd422e2a0396a1d534c395831ccb216bf

Observation 95a44259-4307-4cc3-8ac1-b9ac02df23cf · outbound

This paper cites 大模型逻辑推理研究综述 (survey on logical reasoning of large pre- trained language models).

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL 大模型逻辑推理研究综述 (survey on logical reasoning of large pre- trained language models)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.593555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.291168Z digest=sha256:534e988649a7726c19cbcf57ec6819f4f9ebd849108a7f8becfb9105e00b6e3c

Observation 273438f4-5775-49f5-885a-3fef33514fc5 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.296264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.296264Z digest=sha256:5a307ff9126970c2f3e8a56525449016b6574e35191f88122995098a58fc26fb

Observation 6b5f3609-3501-456a-943d-3574cbc9e3f6 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Understanding r1-zero-like training: A critical perspective, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.302217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.302217Z digest=sha256:f35d3ffe2b18a5aefc847b5574665c305e00674afb8ce9006b3af0000a9e4477

Observation 8e4789af-807a-449b-9d8d-a0c69a8ccd72 · outbound

This paper cites Reasoning with large language models for medical question answering.Journal of the American Medical Informatics Association, 31(9):1964–1975, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Reasoning with large language models for medical question answering.Journal of the American Medical Informatics Association, 31(9):1964–1975, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.560070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.310288Z digest=sha256:d89f893de2918ced761edff316469aa24bd41e6297e4f9aa4ae0579818b0fdf8

Observation 1e5db605-f610-426a-a1aa-ccd084413474 · outbound

This paper cites Improve mathematical reasoning in language models by automated process supervision, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Improve mathematical reasoning in language models by automated process supervision, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.543417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.316399Z digest=sha256:d220989b1542bc4d280c73ea36937bf0e6053e9e6e7730cd1e324275a253ca09

Observation 8ca36034-17c8-4859-8c4b-7a18a21a1959 · outbound

This paper cites Exploring the limit of outcome reward for learning mathematical reasoning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Exploring the limit of outcome reward for learning mathematical reasoning, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.320941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.320941Z digest=sha256:d7904813bc1f45f739025f90768b3ca78e534dc51a00530fd8ec3c386c016e83

Observation a64b2246-ff11-42ec-b052-3c2b371c4d6d · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Real: Efficient rlhf training of large language models with parameter reallocation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.325245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.325245Z digest=sha256:6d281af85116588f382206224c5c1a1dfc8bf286f630bf9be2d8dbde1220e6f3

Observation 89658dd5-f79b-4562-8d19-88c8603d5be1 · outbound

This paper cites Enhancing reasoning capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Enhancing reasoning capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.489665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.329775Z digest=sha256:8edb6b584b4780ff614425bd059996d19b1713fd8573fbf423604a69f819e912

Observation 9965ee69-3d72-4e05-9625-12449474ce55 · outbound

This paper cites Advanced semantics for commonsense knowledge extraction.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Advanced semantics for commonsense knowledge extraction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.465259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.334589Z digest=sha256:c77ab1b3c3c59391b1986a2e5b5b6924aaccfc8a55a5bf223bf9296c8fe8f13c

Observation a2abe552-b684-461c-8ee6-e02b32a85bb3 · outbound

This paper cites Learning to reason with llms, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Learning to reason with llms, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.340441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.340441Z digest=sha256:9c8f74e62004bcc21b596d9993b0e840e64f61a0970ed33109f63eaeda2b0510

Observation a124621e-0d8f-45cd-878f-6db8f00f9a15 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.345622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.345622Z digest=sha256:fdabd730313abbfa09493f58f37ddc7c7c0cb05481ba86a59503b1358007dbd4

Observation 36668151-2c8f-4385-8f61-4c85f13e2149 · outbound

This paper cites Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.421265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.350765Z digest=sha256:943f2c61dc6253057ac133bf4b38b45e5a33a642bc793052f8a861e82aed474f

Observation c84e8efc-0eb1-408a-b7e6-609ab6f1ef72 · outbound

This paper cites Qwen2.5 technical report, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Qwen2.5 technical report, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.355048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.355048Z digest=sha256:25a4b429cb9b9fbe11635162df04e756d763f997afd35bd8dba80bc2c17bf51b

Observation bfc4097e-25ec-4b96-a47e-faf186ca7e4d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Qwq-32b: Embracing the power of reinforcement learning, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.361130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.361130Z digest=sha256:d73fc7da8cbe47a23c41bc5442e5007e6c274faf1c57e8030fdee4248ae0ae69

Observation e8703507-c9af-4c01-8e63-1957d0e91145 · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.366029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.366029Z digest=sha256:5dcb32e2b83a913b4470f789a289e4e7efe4860b54958caa8b18a857c578211c

Observation b0f0895b-9acb-461c-9d78-6c8844fc20b8 · outbound

This paper cites Testing the general deductive reasoning capacity of large language models using ood examples.Advances in Neural Information Processing Systems, 36:3083–3105, 2023.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Testing the general deductive reasoning capacity of large language models using ood examples.Advances in Neural Information Processing Systems, 36:3083–3105, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.381247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.371099Z digest=sha256:f1172cb49f1725c54008531d1266f49b850aef804b9d15002303b73f842643fb

Observation 53e22eda-0e18-4252-9fd6-6c8b7321b991 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation, 2018.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL High-dimensional continuous control using generalized advantage estimation, 2018

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.375876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.375876Z digest=sha256:09b2c9c45d5fc1095e2238310941ba4d701a118a1ee5ecd6b10d5bfffd059b31

Observation 4665d72d-0ab1-45b7-9720-a7522d293317 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Proximal policy optimization algorithms, 2017

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.380438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.380438Z digest=sha256:a0d76aa666e46638954ca6ff84a3af0256056a6f400df7badc5408f221f6daf0

Observation 76bf7573-9660-4de6-b51b-193c6c08f8cc · outbound

This paper cites Automatic solutions of logic puzzles.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Automatic solutions of logic puzzles

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.332435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.384947Z digest=sha256:c91252db8842e1dce0315f68efde2658e483a9cc0d7b155253610a486d396f55

Observation e2a07ee8-ad48-43d5-bdb8-7c983b3c3e7d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.390455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.390455Z digest=sha256:2d5cdc02d3b79b390059f1714481b8ec00483bb8b028578bb363f0cf5c4901d8

Observation d3877bd5-b1fc-4b05-8349-1711b53a8446 · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.395257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.395257Z digest=sha256:7298971bbbb446d043c6498726f45a51874a29ea680a091556f88e3eb9b941be

Observation 8531279f-7ae6-4beb-933a-c0eb75e3c785 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.402593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.402593Z digest=sha256:184a0b5261a6302b18f197c7300d8fa4e85fd3daa2837de113dfd9c1b97f35d1

Observation f8b7f272-204d-465c-a1aa-7444d57cf1f9 · outbound

This paper cites MIT press Cambridge, 1998.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL MIT press Cambridge, 1998

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.408266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.408266Z digest=sha256:3909ec67eca73130a3ab76e92bc75e81da966a5f715a35d905d4aa8964f70b96

Observation c3f4937f-ea0b-4334-a053-bd4aaf3d7510 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.413231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.413231Z digest=sha256:b449d286d90b5e8f3a5b07b61aeb57f53ee093575b520153eba0cbedbbc76df6

Observation dad3b24a-8875-480f-9a86-f853f97c303b · outbound

This paper cites Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:42:06.625878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.419474Z digest=sha256:a12a24582fa52c86ff21713114289254e4dd25297cf2eb2676528ac00ec8d4df

Observation c0cf8e43-d49e-4adf-a9be-78609e07433b · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.302087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.424640Z digest=sha256:d0c344e94dc1f76a0dfbec439443e09a215c8715874829f3162e977b5f498e90

Observation d9c6894b-eabb-43b7-9ca1-1543ab1a513c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.429383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.429383Z digest=sha256:21d866d766d08c395196f1c4d5faf726d49f01a21f2ba71816d38cb63d3ea011

Observation 543ede4a-f228-40e4-8e54-4343af2b16e0 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.263273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.435873Z digest=sha256:7519c15550a83b24721099b788c4e766ed3b66d9fd9aaf34857ba10558fb6580

Observation 71bd38f1-573f-4554-bff1-6b107809ad40 · outbound

This paper cites Legalreasoner: A multi-stage framework for legal judgment prediction via large language models and knowledge integration.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Legalreasoner: A multi-stage framework for legal judgment prediction via large language models and knowledge integration

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.244773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.440662Z digest=sha256:7738d2857de0548a0ccc1249ab4876f04f62e3acf36ff8e0c55b532d666d3822

Observation 8ce955bc-2e1d-48be-a995-f1f6c09e22ea · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.445735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.445735Z digest=sha256:3a66cb1bd8665c9ff689ac68cbcca82445f36cc0254e50aad8c941c089578362

Observation 2c2e16b8-680b-44d5-8e29-414d7bb94455 · outbound

This paper cites On memorization of large language models in logical reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL On memorization of large language models in logical reasoning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.216188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.450159Z digest=sha256:b8b00827c915a0da5e1395a83bff8021dddf7f46c52755b560acf08bad24aaf8

Observation 36449521-a3f2-453d-a8b9-af3ee859ceb8 · outbound

This paper cites Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.455832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.455832Z digest=sha256:23f1b7eb813fb304f4051143fde0cf9df390641ebb1f1509cc3e7bdabddb76b6

Observation 94806828-92e3-4d58-bd00-648bf406b4d7 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.460031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.460031Z digest=sha256:194a00a16381c916e15233adc5db65ed7740b3a75ffaec15cdf698d32e68610c

Observation 95b6e9fb-a1bd-4f56-8651-e9a10cae643d · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.464464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.464464Z digest=sha256:4887d9984e44f1847b351aed0e5423fe6efcf3571e40f11b7b7c5c909b6a65f5

Observation 652cb997-1e1d-4751-bc58-5e4b95ddbf04 · outbound

This paper cites Lawllm: Intelligent legal system with legal reasoning and verifiable retrieval.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Lawllm: Intelligent legal system with legal reasoning and verifiable retrieval

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.174320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.469293Z digest=sha256:43b034896f9debbe56004966e0efbccfd10894f89b20942728a004ca780b1882

Observation a2e9bfbd-2315-4212-ae65-7a9fcd6943c1 · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.474025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.474025Z digest=sha256:87376465fd6d494a8f01cd8c33b2ee389c70062e403f45b131af7b661b4e3dbe

Observation 4404c5fd-2fa1-423a-a3b1-419b3d1c9fc1 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.478842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.478842Z digest=sha256:574b12597db0ae3a09a7a82c6a210aa85cedddd71644dd72fd8677e6e29a155f

Observation cb0d3350-691f-4d83-82e6-406ba2ede228 · outbound

This paper cites R1-reward: Training multimodal reward model through stable reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL R1-reward: Training multimodal reward model through stable reinforcement learning, 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.136777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.485081Z digest=sha256:c80410200e3eb375b9682cf1b82e26f52ab9cd6a4e36e57afa0d69ba23067a29

Observation 921924ed-5215-4ebe-ba12-99a806a60fa3 · outbound

This paper cites AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.489294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.489294Z digest=sha256:087b16966c29d63ea7e5eeaa8d40fed652edabf59614d53842eb7295d8b5d4ce

Observation d8179a4a-a3d1-4047-9b5b-e87933e7816b · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.115497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.495133Z digest=sha256:0ee2d256f66302f740eecacfe61bce167f64d189923764bbd5c438219a275b2c

Observation 031ed804-4040-4d9e-b669-bb1e4e812f50 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.090966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.500166Z digest=sha256:bf244106a6d87da8a4079e637f71652857747ad07ec589869f48ae9ff002a2b2

Observation 0c848a07-871c-470a-92bc-4a2048ff9b50 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.069307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.504954Z digest=sha256:bc2cfbb4ea43fd1e836ca090cb2aa5e4fe7043dff75166a6b9ec3117ade97fbd

Observation 63fdc848-2390-47ce-b156-8f6f6ed1cfdf · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.050266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.510163Z digest=sha256:2f072aa937a262a62bf07a80af35f8d65358a4ab000dbdecba07a4ba10f90df1

Observation e0dcc09d-4cc0-470b-9de8-7480d9989c85 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.031013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.515259Z digest=sha256:0ea63f548edc0bd403aef51bceb08f4babcd96e6af380eaa6dc1f3c0c0e50cec

Observation 5ceff70f-fa52-4110-b4aa-9ba3b33e5eb8 · outbound

This paper cites Alice studies.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Alice studies

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.011688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:42:06.520147Z digest=sha256:82cac3b60fba2e2ac6a0e6ec4ae833716c257f51cd3c49ab06b066e4d14da38a

Pith citing papers

No inbound Pith citation observations are available.