Pith. sign in

Paper Citation Record · LEDGER

Unlocking Recursive Thinking of LLMs: Alignment via Refinement

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2506.06009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06009 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:07:02.121912Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246ca00c-68a2-4c84-bd15-805de75a1b75 · outbound

This paper cites online" 'onlinestring :=.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.972708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.972708Z digest=sha256:029372d1eafcc7b7268092e11419d8217d848c0a8feb4b4b5f00260e586f4c14

Observation 939f3c85-c716-43fe-821c-736597ed7817 · outbound

This paper cites write newline.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.976315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.976315Z digest=sha256:a3a1ce85607fe9e50c3a52acea95b9489447e6a71a5cca62d40c97e51fad4a43

Observation 6a23e769-4d86-4c64-9a8e-9603b5d2a836 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.979925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.979925Z digest=sha256:87d7b6d799ba240868bbca7339096076c0ac5beba1792784db5cef85ad696f02

Observation 47b801cc-3afb-4edc-95fa-edd60c593ab1 · outbound

This paper cites Critique-out-Loud Reward Models.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Critique-out-Loud Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.983009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.983009Z digest=sha256:acb736da2c20d875040a6be3f3c035afae82c40ea56d240ee7516ffafbbeb80b

Observation c55dfe77-2b6c-49ab-b566-62fab778e8b4 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.551702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:01.986290Z digest=sha256:c5ee77df417de738f26d38ef6696123a1e89c49f4e23061972b712f3b6bff1aa

Observation 6c04a003-9f19-46bf-b99d-54b7299b459f · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.989412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.989412Z digest=sha256:b1245deb202b9052ef328dd50329b3c8a38e70771ab40613c47183d6e640e165

Observation 6839484a-01ee-4961-8192-6cd9c52c977e · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.992586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.992586Z digest=sha256:9c25b71e551c870b7ebc33668cb3b013c83151438628f58c1275f4cc84e730f2

Observation 7813797b-1ee3-4912-9539-35b925349318 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:01.995800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:01.995800Z digest=sha256:370fe782f029bbc45b397d2c5b85ec5d153ef95e5d58cd4484a481af0af33661

Observation 380dbae0-c436-4b05-80cd-67d470103764 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.542765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:01.999001Z digest=sha256:ec94c217f6428cf56ffa957bf55237b448b1f17d39840fa93a735ca100a702e4

Observation 70eb2ecf-2649-4acf-a152-1efe251e6f61 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement KTO: Model Alignment as Prospect Theoretic Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.002272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.002272Z digest=sha256:61a5fe466d8b25353c7252cd879abfa821709cab7d7354d3f6e53bb63d57182c

Observation ffc07461-ffdc-44a3-8845-da654ca7584f · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.005898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.005898Z digest=sha256:7293a076f80511d4de9b6f2ed5eec6f228eac5d0b4cb34dd18910a04e389b99d

Observation 1deacb34-0eb5-4a6d-8403-2b305905c92a · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.533322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.009364Z digest=sha256:07a361824d1ab7a0538785828766f9528690e1f1b8829c995ced346934b4ce17

Observation 99d2a96c-a0a0-4b41-b672-bf9620c8ca44 · outbound

This paper cites SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.012071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.012071Z digest=sha256:11a66eecc754e99078f8c04c8458cf96ebe7ac47a54ed665177b3c5c0e11f70d

Observation 6509816e-6a6d-47dd-a93c-8d7edb8b66f2 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.015165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.015165Z digest=sha256:c41e612300c1489904d26e63442fc67494ccc7715173486a85cab3c8f0886892

Observation 204fcbc4-8020-47c9-b96e-6a8676ecf97c · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.517796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.018445Z digest=sha256:9db31fa6257e98f1b24d7e531df218ddbf5784716bacafdb614612f1270a9457

Observation 705a9ac5-f8aa-43a1-acfa-18a3e9b4fbe4 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Training Language Models to Self-Correct via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.021435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.021435Z digest=sha256:d271768c34676879a5cb322e734dd0918aac37161ca5ce771a0e55c7cb1b668c

Observation a7932206-d0a4-4c75-bb3e-9b990646251b · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.024636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.024636Z digest=sha256:eb302d81cb6dc314803e8923b34c8eb87b1df50aba08af403e40b3be1b03bb2a

Observation 9d099228-8a1c-4a64-9c96-40bce31657d7 · outbound

This paper cites Hashimoto.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.027559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.027559Z digest=sha256:8b0580e12b4ba3a7d1768db4638c3c43eaa8956707ce6020b32b3ce266fd82a9

Observation 9ff1999d-45be-4193-9751-ae2c244ffcd8 · outbound

This paper cites Fennec: Fine-grained Language Model Evaluation and Correction Extended through Branching and Bridging.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Fennec: Fine-grained Language Model Evaluation and Correction Extended through Branching and Bridging

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:07:02.326602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.030197Z digest=sha256:b80fec1e7f5a365bb661179d8b2c3a64c11a39424755e4c58ab9b1d9fc2536a8

Observation e9d85ba7-afeb-4fd9-a028-9647961ba4e8 · outbound

This paper cites Let's Verify Step by Step.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Let's Verify Step by Step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.033102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.033102Z digest=sha256:1e263ff5745ff1e996064d63c4a80531e039b547d54afe5d82e7665bba3a2c46

Observation 71669695-99c8-4e6d-8853-be1d52f87454 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.036124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.036124Z digest=sha256:8780036a2731a5dcf54a3e7c8e3e162601fca2e23b2713fa0a271ad3f405222c

Observation d449e7df-6801-4529-a0c1-68995e4c7361 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.039027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.039027Z digest=sha256:fd4e266b8f37e6ab62878e99beb0972777121abe1ba7c2290d9e8ec32fc16c0e

Observation cda1544f-8451-41de-9077-a474bf655e98 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.041738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.041738Z digest=sha256:09c73b4d5d6c04d331f6f859e3a9ab9c85ef9d0e6131a16adf64bbbae34d59e1

Observation 48c893e3-c9f5-446b-931b-6ef14232d5c1 · outbound

This paper cites s1: Simple test-time scaling.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement s1: Simple test-time scaling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.044561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.044561Z digest=sha256:faef7ccc376b07656885f4ec76afa9c692617d8bf97d2dc72cdb70eb111ba7e8

Observation dd036561-2dc0-4c25-ae10-2a2b251de0ae · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.497164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.047780Z digest=sha256:247a65168c695ea3e2eb95af384914e0d69f26572f7c195e43455834233a5c39

Observation 7faebdb5-bce8-4568-9819-96e8c32d9373 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Disentangling Length from Quality in Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.051648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.051648Z digest=sha256:ed3412baeda298ac746c744cf67cd28df80251df8a859963c5704b040b73bfe4

Observation 2d7ca1f7-6c45-49c7-8011-f26477e5e3b3 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.055033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.055033Z digest=sha256:42da885d1f79739fc893e3f0f33f277063ecab45d25ffe6b9356c916dc5ad865

Observation 240fdee3-e6fb-4e72-ab6c-d236343dbeff · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.058112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.058112Z digest=sha256:75d91e52b7c706a97e074dd7ba49f8709f90647f6a0157ea9832a15de8c2b033

Observation a0289a2f-f48d-4367-9d62-3fb1d7e8a19a · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.488026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.061385Z digest=sha256:8f386fa72b6e3f05d84fbdaba8f67623da4b21f5256a68e2ba16a6a9b6fb9abc

Observation e3c7f2da-8a66-47c0-b9b9-cb3b863d71f2 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.064199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.064199Z digest=sha256:2150903b7151001de8230e7950f5ffc38ea779062c513f13df43a184a7750be0

Observation 8d62d28c-59f8-4d8d-a541-ff514de9200d · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.066957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.066957Z digest=sha256:f701e55d7865fa0d05548f54433cbf16083e1a34a20cb39cd109f377241823fb

Observation 0f14bae8-f08e-4f62-ba2d-9cb358d8eb81 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.069605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.069605Z digest=sha256:bd117ef04ca6f783e10aee90c0384d9aafaa5ac689ca796c661bb2a7a81d14fc

Observation 429d58db-9a7f-433a-92e6-8cd9a2964ec3 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.072269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.072269Z digest=sha256:9acc8edcdc5c1dd5f78b09a18d53991409c7db824cf02a98ed71126c7df1a17a

Observation b414345b-aa49-4c5e-a267-65b1183c2d77 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.453762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.075154Z digest=sha256:c1f5a0a3bc153bc0e7a68d12fd9ffaaa90b5e2934650e6f49d24382ebfbf0644

Observation 7ca97041-5f81-428c-8287-3dc77c8239d5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.078218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.078218Z digest=sha256:c1e66e428449f697f55a9b8bd9bcc123fac8af85878a165b9d5361ab6e867478

Observation 3a8cd4f1-5e19-4997-bd32-5b4e9ae2bd3e · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.444742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.081251Z digest=sha256:8885ecdce4b8681268119e0e9eaa429c54344c01e29b7a539e679a0f47733f73

Observation ecba1e32-ca90-4164-998b-5c71dbe261b2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.084050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.084050Z digest=sha256:4a6d232ad51aad07cfa1ce30325c664096f2240fc1a95b386eb85e4245e0c14b

Observation 5e4046cb-f43d-485f-a977-3299fb9f0052 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:02.435618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.087091Z digest=sha256:fe532aabf3158b096c22e914f8215f674e1e846d9dde3af44bad705f22c1ba93

Observation e5792110-ad88-4671-964a-a5b8859a5d18 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.089755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.089755Z digest=sha256:6e2826a4d94eb56b04062f631a044be08cc1d29bf581dd6f8b91d30c0f066823

Observation bfb59995-fcf9-4ed4-a7c5-2584175072a5 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.092644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.092644Z digest=sha256:c23cdca94d1505db4786bfcfae99fefad3b43bc72b2970ee39839870e515f23d

Observation 617291b0-867b-4816-97f9-e09c6666a867 · outbound

This paper cites DRT: Deep Reasoning Translation via Long Chain-of-Thought.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement DRT: Deep Reasoning Translation via Long Chain-of-Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.095155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.095155Z digest=sha256:bfc2c26de857a3bf1e0f2c1fa61c5baaa28177a7ad55746ebaf788de6dee93c9

Observation 491b5ac9-a129-4783-ba7c-6607f705e9b3 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.098175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.098175Z digest=sha256:457bba20cbee014ab24843db3ce032d78401f9c1d873d29ff2c542c8a4d414c9

Observation 95bf17a7-838b-45b7-9e5f-4411dfae52a6 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.101272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.101272Z digest=sha256:1d6776e35eeecb74f76e3ea4a7ca234bb6940c207f8810395ee01c0d5e219fa6

Observation 76841248-9d69-4b5e-b366-47d44e28e2b8 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.104060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.104060Z digest=sha256:fbc7222c00e203645dde322d584f5f4623e3c93895b5969a0c961ab54fc673dd

Observation 861754ea-92b0-4b77-ab47-7d256420d134 · outbound

This paper cites an unresolved cited work.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.107069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.107069Z digest=sha256:afde622679a0e4ffdf1ed3c3aab6adbfcd77b5414d57e1f28408418f9e805aa7

Observation 922c3bed-baff-4497-b9b8-297001d79be1 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.109745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.109745Z digest=sha256:0f70d522e681eaaefd22583b3ce279083a72ad01cba0fecbc99b6c9a924dc31e

Observation 25ec39e4-7325-4bf2-93cd-bd83a1ca71c3 · outbound

This paper cites Self-Rewarding Language Models.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Self-Rewarding Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.112916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.112916Z digest=sha256:1973475d5176561af4c6f43b701577874b0d00f8487ae060e485458fb9951c68

Observation 33364e88-0301-43cb-82b7-a273cb02dc70 · outbound

This paper cites Understanding the Dark Side of LLMs' Intrinsic Self-Correction.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Understanding the Dark Side of LLMs' Intrinsic Self-Correction

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:07:02.175915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:07:02.115821Z digest=sha256:c0228f811548a9d0788f4f1e841952bc8921aeaa032756e96f1c8f7eacf9a5ec

Observation 19cb7ee1-9951-4eca-8ed0-a9302ad030b9 · outbound

This paper cites o1-Coder: an o1 Replication for Coding.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement o1-Coder: an o1 Replication for Coding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.118750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.118750Z digest=sha256:85494e301b30c0d0e21072ea370f4092a2f1361d908c6be55a5e78fc90e8def0

Observation b1d0106d-abe3-4a86-b6f0-07b718b24a5d · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.121912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.121912Z digest=sha256:5d57e9bc485ced7214792bd05312d1987b798a58c0882baac935d05f9e1fd385

Pith citing papers

No inbound Pith citation observations are available.