Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL

As of 21 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2509.06024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06024 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:42:06.520147Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22cd9d77-5bc3-43bf-9b6e-012dd366a55d · outbound

This paper cites L1: Controlling how long a reasoning model thinks with reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL L1: Controlling how long a reasoning model thinks with reinforcement learning, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.155669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.155669Z digest=sha256:d84f660a87bd6435bcb6da22b6dd3ac1743840a361c7ada87b97ab2ba79a604e

Observation 2c9f5930-0553-42f5-9e99-0b969903eda9 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.213407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.213407Z digest=sha256:dfdb8e295bd10344aac0fd02979097133e7744779e6a775962728df3790facd3

Observation c27a294e-71ce-4d0c-af54-188db47e8299 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Evaluating Large Language Models Trained on Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.231836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.231836Z digest=sha256:ceb9c43211d276202c261d67c3e0e7388ab4df1b25efb45186e10eb04cbbc237

Observation 69fdc157-bbdc-465c-b495-18e1571d5675 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.339231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.339231Z digest=sha256:42d97cf34819519c89c57574f945e395952b4fee7f2c5d3c736cc0d777d960f1

Observation 7e793c40-6723-4edc-a07b-443d5e24281c · outbound

This paper cites Empowering LLMs with Logical Reasoning: A Comprehensive Survey.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Empowering LLMs with Logical Reasoning: A Comprehensive Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.386445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.386445Z digest=sha256:2f740ad1b4f368203a26d3f93c9b41d11256799b052fe3fc0111838a4470121f

Observation 212bd76f-6364-41ed-ac58-8ef82c2f2f98 · outbound

This paper cites Self- playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Self- playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.796428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:05.458658Z digest=sha256:3bca6569625999c9a4c0d0cb73e80cecb9a7c30641a5ec34d7c6b8a870dd5833

Observation 0e1d43cc-a800-41fa-9bd3-ba158b18055a · outbound

This paper cites Transformers as Soft Reasoners over Language.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Transformers as Soft Reasoners over Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.511121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.511121Z digest=sha256:74863ec75943bf56e5637610412c20af8d607cdaa80a582bb00d1385591e9748

Observation 0a96dd2b-7f59-414c-b992-b6f92d214188 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.571433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.571433Z digest=sha256:4ed69cac059085e19043057717ed97dbb928a7d34002a703efa1bdd16665abc1

Observation 09d5c820-15d6-4f1a-9671-70f7e1198aab · outbound

This paper cites Gemini 2.0 flash thinking, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Gemini 2.0 flash thinking, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.669140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.669140Z digest=sha256:495ca01f8e186e976117a64b749206384f814239cdc8c1e45d003abae9a909db

Observation 5269d331-add4-4aea-a46d-b61819dccfba · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.728700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.728700Z digest=sha256:ba793b6c03c1475c5e58787bc69d91b995d8e1fdb4ebffa65b581eda4178a4b1

Observation 9a8ec324-7c9d-4337-b5a0-2ec0abbe9e96 · outbound

This paper cites Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361, 2021.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.798953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.798953Z digest=sha256:ab23e64345e868c33e741e0402195c6f7346817671de432c5097a08ac83fc7bd

Observation e546555a-3bc0-4ed5-b677-ec87a70695d5 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.864107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.864107Z digest=sha256:1ba7c586348d561a1a7f45a959823d0db7f28a1b3f0537fc49046e45df0a744b

Observation bcd6ed91-71c9-4de7-b825-c40330044a9d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:05.932027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:05.932027Z digest=sha256:0126dd49f1af10256c5b77e9c07a0e50185d1e8681f1e60d8193aa93966b2b91

Observation c1d3bab2-ec8e-4b06-b5ea-c4ff44d58900 · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL FOLIO: Natural Language Reasoning with First-Order Logic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.003231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.003231Z digest=sha256:4cdf14b13139ea6598b7c272a9f6fbff5d6555a84be572a7e3713d175dbc4eb8

Observation a19a564f-598b-4599-a236-031e91b38bdf · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Measuring mathematical problem solving with the math dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.068109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.068109Z digest=sha256:53ba34824194dcab528980ef74e20a8f01f8f9bed8bef6bc5c3480fe4824cb6b

Observation 355d0c2c-b29a-472a-a57c-407d9fa9959e · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.155144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.155144Z digest=sha256:ac36797b390e7f637df1a3cff3b5f303a0654502b3fcc8f7aacffe25f8d0f721

Observation 18da2e70-102c-4c7f-8c81-4fd3f7fa4496 · outbound

This paper cites Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.221502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.221502Z digest=sha256:6d21f7b0d665ce0b1c7d9d2675d7eda3c85a0c9803386f5c603f446d7d267af7

Observation 6442c8e8-9681-4416-9224-33767f2bbff7 · outbound

This paper cites Boardgameqa: A dataset for natural language reasoning with contradictory information.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Boardgameqa: A dataset for natural language reasoning with contradictory information

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.734162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.241979Z digest=sha256:46f435bd684c5795a3f8d75602128f602a9873f1dcb3d1885cb50bc5e18dbad1

Observation 2922c9ec-3339-4875-bddc-725e62d6c14e · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6, 2019.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6, 2019

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.704465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.247220Z digest=sha256:b66e75e0ce2d0d0ea34ef8af52c118a0912bdea8d9a27c1fb9b9375f14cf593a

Observation ee6cc1d8-bbdd-4b9b-9e66-7843aa934929 · outbound

This paper cites Chain of code: Reasoning with a language model-augmented code emulator.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Chain of code: Reasoning with a language model-augmented code emulator

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.678037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.254413Z digest=sha256:c033982be9911440b932529e9844a2eff521e2567b0d3b7b9162c727d6a1a5b8

Observation 03c8f9e5-1a42-40d5-b2da-16e5218e12b3 · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.261278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.261278Z digest=sha256:9654a2e84d1525cceed44d4691e35ae7fe6c9387a6194dab560824a761f201b6

Observation 0a7f6c22-c1a8-4074-a35e-e00a8c6b8587 · outbound

This paper cites Limr: Less is more for rl scaling, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Limr: Less is more for rl scaling, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.266810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.266810Z digest=sha256:bcf4528e1d03f4cd91359cd3d1b894cbe527f88a3ca0b2feaf1958d29c5497ac

Observation c4c61012-3a83-4b03-9eea-bf583e5e64ee · outbound

This paper cites From system 1 to system 2: A survey of reasoning large language models, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL From system 1 to system 2: A survey of reasoning large language models, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.644307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.271252Z digest=sha256:4496533026a27c8ed0a7b43653011e408af897f00a03fef6c26ca174f3ca6539

Observation 29969c6a-139c-4e09-a049-55ed9bb828c8 · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.276239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.276239Z digest=sha256:d757788d4ac12c4781f8678c560b73b6fd2020e4ce20731be25234bcaa440ea5

Observation 8423668f-70ee-447b-833a-9dee553df8a4 · outbound

This paper cites Natural language inference in context-investigating contextual reasoning over long texts.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Natural language inference in context-investigating contextual reasoning over long texts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.628502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.280844Z digest=sha256:ef08b38f5d18c098942c5c605b6de3be3059174bc4324172cf1f5db76d0c6dd7

Observation 3e1a4b64-531c-4654-89b3-9a8337c56d83 · outbound

This paper cites Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2947–2962, 2023.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2947–2962, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.610533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.285793Z digest=sha256:846506d3f35af2e0b910c81fe6d9050147b2273f9210eb25787a88d5a79ce495

Observation 95a44259-4307-4cc3-8ac1-b9ac02df23cf · outbound

This paper cites 大模型逻辑推理研究综述 (survey on logical reasoning of large pre- trained language models).

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL 大模型逻辑推理研究综述 (survey on logical reasoning of large pre- trained language models)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.593555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.291168Z digest=sha256:1a38ac9c07e3c57acef4a8ae5f77672d586aad2dac185ae1f7526c4674967540

Observation 273438f4-5775-49f5-885a-3fef33514fc5 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.296264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.296264Z digest=sha256:cdcdaa191efc53474c8e529d12ef67f7da53a3e3b9abca253b96acd43ac533c1

Observation 6b5f3609-3501-456a-943d-3574cbc9e3f6 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Understanding r1-zero-like training: A critical perspective, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.302217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.302217Z digest=sha256:5ec49f7103ffaf38c424339df0f54378106b067a910a431b339a5e2c5c417751

Observation 8e4789af-807a-449b-9d8d-a0c69a8ccd72 · outbound

This paper cites Reasoning with large language models for medical question answering.Journal of the American Medical Informatics Association, 31(9):1964–1975, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Reasoning with large language models for medical question answering.Journal of the American Medical Informatics Association, 31(9):1964–1975, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.560070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.310288Z digest=sha256:558eb083eece299e45cd6ebc990862910c4f67cdd6cc1783bbdb5ba16d918a36

Observation 1e5db605-f610-426a-a1aa-ccd084413474 · outbound

This paper cites Improve mathematical reasoning in language models by automated process supervision, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Improve mathematical reasoning in language models by automated process supervision, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.543417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.316399Z digest=sha256:8e63643a4365817fa05f8588e6f9b5189ac4f28906362e703efd991ddc9df995

Observation 8ca36034-17c8-4859-8c4b-7a18a21a1959 · outbound

This paper cites Exploring the limit of outcome reward for learning mathematical reasoning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Exploring the limit of outcome reward for learning mathematical reasoning, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.320941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.320941Z digest=sha256:5f19e345d83e628ec0e294ab1e2d21f6d2dec0b39f0af920a1a27218a2ab5c8a

Observation a64b2246-ff11-42ec-b052-3c2b371c4d6d · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Real: Efficient rlhf training of large language models with parameter reallocation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.325245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.325245Z digest=sha256:0553739d7bcbade53d352fd17e4fae4ce8881b71750402ba92ab388abc34ea4e

Observation 89658dd5-f79b-4562-8d19-88c8603d5be1 · outbound

This paper cites Enhancing reasoning capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Enhancing reasoning capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.489665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.329775Z digest=sha256:80f3fe49b52db67aeda21589bef61049867d694cd8263f646334892e444380e6

Observation 9965ee69-3d72-4e05-9625-12449474ce55 · outbound

This paper cites Advanced semantics for commonsense knowledge extraction.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Advanced semantics for commonsense knowledge extraction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.465259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.334589Z digest=sha256:93ea6c77609270dbdc4f1021bac88f6582a2cf03b88fac04928772308675abd3

Observation a2abe552-b684-461c-8ee6-e02b32a85bb3 · outbound

This paper cites Learning to reason with llms, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Learning to reason with llms, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.340441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.340441Z digest=sha256:e3391a8ec4ff4fd8d4366192df0de0740a00d73f00b563d6c1e02d86d4cf01f6

Observation a124621e-0d8f-45cd-878f-6db8f00f9a15 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.345622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.345622Z digest=sha256:578f0ce3d864ffcfd6858588a1ae5a321d4ee3512e90459023c97b47248a34ed

Observation 36668151-2c8f-4385-8f61-4c85f13e2149 · outbound

This paper cites Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.421265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.350765Z digest=sha256:8104727964652b5ac11a5cb5092704cb4370f722d7e31c6b5679a24162631af3

Observation c84e8efc-0eb1-408a-b7e6-609ab6f1ef72 · outbound

This paper cites Qwen2.5 technical report, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Qwen2.5 technical report, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.355048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.355048Z digest=sha256:774e2fae78dcb220e06b787f11adff0698d97f1ba947b54e01b04d6f3f813aa1

Observation bfc4097e-25ec-4b96-a47e-faf186ca7e4d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Qwq-32b: Embracing the power of reinforcement learning, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.361130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.361130Z digest=sha256:6ecb9d6ab01faf29289025f9783b136c9711f1c783466a9837475a9362ef384a

Observation e8703507-c9af-4c01-8e63-1957d0e91145 · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.366029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.366029Z digest=sha256:ffaf847675fd8fce7b077550d75b3fb07872e62225f1ac1b93dd5c66f066612c

Observation b0f0895b-9acb-461c-9d78-6c8844fc20b8 · outbound

This paper cites Testing the general deductive reasoning capacity of large language models using ood examples.Advances in Neural Information Processing Systems, 36:3083–3105, 2023.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Testing the general deductive reasoning capacity of large language models using ood examples.Advances in Neural Information Processing Systems, 36:3083–3105, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.381247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.371099Z digest=sha256:9593a661a6eb304cd12cf4a48c6391b260d356d137cf8d619848441528d8be18

Observation 53e22eda-0e18-4252-9fd6-6c8b7321b991 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation, 2018.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL High-dimensional continuous control using generalized advantage estimation, 2018

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.375876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.375876Z digest=sha256:2d9b17313399a8b632df98fa1d95eeae171fe1e4602cb27341636d4bdf76275f

Observation 4665d72d-0ab1-45b7-9720-a7522d293317 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Proximal policy optimization algorithms, 2017

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.380438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.380438Z digest=sha256:ad29ab6fc9f5f8d4f7168bbf06895d3dd356a061a30cf6784075e8d1e0be1b83

Observation 76bf7573-9660-4de6-b51b-193c6c08f8cc · outbound

This paper cites Automatic solutions of logic puzzles.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Automatic solutions of logic puzzles

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.332435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.384947Z digest=sha256:7ea99a8512cabb3dc6f610ff9ffe84d116cd86e7c546ae628dcc0ce81f41b459

Observation e2a07ee8-ad48-43d5-bdb8-7c983b3c3e7d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.390455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.390455Z digest=sha256:682dc7fd162b00f88f9922c290ab2ea16d955766736262b0f7791b37dbf1bafe

Observation d3877bd5-b1fc-4b05-8349-1711b53a8446 · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.395257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.395257Z digest=sha256:96ac226ff89bc2eb8ab3daf052d2a82043962ebfacae2878b3682d6eed797d86

Observation 8531279f-7ae6-4beb-933a-c0eb75e3c785 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.402593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.402593Z digest=sha256:e46baad15ff5cc3dd4d757a7820a87b88f1e8c93d36c5c6810c44252a12da789

Observation f8b7f272-204d-465c-a1aa-7444d57cf1f9 · outbound

This paper cites MIT press Cambridge, 1998.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL MIT press Cambridge, 1998

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.408266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.408266Z digest=sha256:9cce6934be2f8575734d9be499573bc2f64e7ea05e7db0513486325359800d60

Observation c3f4937f-ea0b-4334-a053-bd4aaf3d7510 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.413231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.413231Z digest=sha256:1abac07feb2ba3bc7ebbc59e14978b2adabd89cbf111d16e7a9b7ffb5856c8df

Observation dad3b24a-8875-480f-9a86-f853f97c303b · outbound

This paper cites Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:42:06.625878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.419474Z digest=sha256:d93e1f666f85359bcd9c8fb45f976f793f1f2dacf967adacbcbacf3cbd7ffa31

Observation c0cf8e43-d49e-4adf-a9be-78609e07433b · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.302087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.424640Z digest=sha256:109dbc83b15497d083d8618b262980f42f7794feb47d95095690015cc6a46cfa

Observation d9c6894b-eabb-43b7-9ca1-1543ab1a513c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.429383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.429383Z digest=sha256:a27254abcce105949523d50ce2d4217522d357ad769521a02d85471b890d70f2

Observation 543ede4a-f228-40e4-8e54-4343af2b16e0 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.263273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.435873Z digest=sha256:3a98e4f53abb7384d4c8d29adca4ba4d7df97d10126be9f5abf0265b81bf0a18

Observation 71bd38f1-573f-4554-bff1-6b107809ad40 · outbound

This paper cites Legalreasoner: A multi-stage framework for legal judgment prediction via large language models and knowledge integration.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Legalreasoner: A multi-stage framework for legal judgment prediction via large language models and knowledge integration

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.244773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.440662Z digest=sha256:5b24d8616e56dc4a3091e67b41a5510dec7f176291c1ea087ad32be0259f9219

Observation 8ce955bc-2e1d-48be-a995-f1f6c09e22ea · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.445735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.445735Z digest=sha256:8551d50057ef9b4621e38c27e363b20cc2ff7b8bcf0c52fc9cc3647dcbe53bb0

Observation 2c2e16b8-680b-44d5-8e29-414d7bb94455 · outbound

This paper cites On memorization of large language models in logical reasoning.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL On memorization of large language models in logical reasoning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.216188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.450159Z digest=sha256:adb7873c0fd5844490e7a2dcdc3d2d36b279cd17852464636ea2f4a916292d63

Observation 36449521-a3f2-453d-a8b9-af3ee859ceb8 · outbound

This paper cites Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.455832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.455832Z digest=sha256:637821a63e54845f0876a1f15891654cfd7e7986d01cd8609b5dfcf362931fb1

Observation 94806828-92e3-4d58-bd00-648bf406b4d7 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.460031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.460031Z digest=sha256:0e1c534120efc2477e03f3d0ba1a3a919f7d9dcc85b99a64d30d3e9e0eec6a4a

Observation 95b6e9fb-a1bd-4f56-8651-e9a10cae643d · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.464464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.464464Z digest=sha256:82865c2b7b98bd8ca53567f7d8409dbda54473d39ad62991138bf4dabb4bde8f

Observation 652cb997-1e1d-4751-bc58-5e4b95ddbf04 · outbound

This paper cites Lawllm: Intelligent legal system with legal reasoning and verifiable retrieval.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Lawllm: Intelligent legal system with legal reasoning and verifiable retrieval

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.174320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.469293Z digest=sha256:f80269c300901e3752b1f22f0453f6ae44bfecd2329f27af54ada9fa34c3978a

Observation a2e9bfbd-2315-4212-ae65-7a9fcd6943c1 · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.474025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.474025Z digest=sha256:89c85556290f212cdcf514e334747d5bbc62a4f678dfb5919fb9e63e7569f53b

Observation 4404c5fd-2fa1-423a-a3b1-419b3d1c9fc1 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.478842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.478842Z digest=sha256:428c0b428bec53b7115c3602fd134d805239cb8f328247708316c3481c8d4500

Observation cb0d3350-691f-4d83-82e6-406ba2ede228 · outbound

This paper cites R1-reward: Training multimodal reward model through stable reinforcement learning, 2025.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL R1-reward: Training multimodal reward model through stable reinforcement learning, 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.136777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.485081Z digest=sha256:10f75c468f7960fd873cc1d850c953c74705da6b2ee43aaefbc44b10c10b7747

Observation 921924ed-5215-4ebe-ba12-99a806a60fa3 · outbound

This paper cites AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.489294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.489294Z digest=sha256:df6a8d88973561e66d3865cc4d9603e0d893a120e12e1482f2c82b3c6f5f5821

Observation d8179a4a-a3d1-4047-9b5b-e87933e7816b · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.115497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.495133Z digest=sha256:94001df5aed77ecc9abb489e958159c88c96723767343828df4daaa54e86018a

Observation 031ed804-4040-4d9e-b669-bb1e4e812f50 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.090966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.500166Z digest=sha256:d06f4d31803f8b836663f5b3dda08e508321a6e859445553ec9ccbb2f2acfb87

Observation 0c848a07-871c-470a-92bc-4a2048ff9b50 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.069307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.504954Z digest=sha256:2aa59052560cb18327b4c92b0f14a3714d440adb6d8c1c60928b4eae18705b26

Observation 63fdc848-2390-47ce-b156-8f6f6ed1cfdf · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.050266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.510163Z digest=sha256:b31a7219114757e1d606e552ae70cfd52f59bc9234e673f5f805b744b81cd2cc

Observation e0dcc09d-4cc0-470b-9de8-7480d9989c85 · outbound

This paper cites an unresolved cited work.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:42:07.031013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.515259Z digest=sha256:accfe6625d115ac753dcae37f184d1c32e66ccbca829e1db03b7d3882cb8bb32

Observation 5ceff70f-fa52-4110-b4aa-9ba3b33e5eb8 · outbound

This paper cites Alice studies.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Alice studies

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:42:07.011688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:42:06.520147Z digest=sha256:e7913ffe7f6dfe76bf3805d87ca9ea73f526fddd858e3224254a65c590ec48d7

Pith citing papers

No inbound Pith citation observations are available.