Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Enhanced LLMs: A Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2412.10400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10400 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:49.495121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.174737Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03a804e4-8548-4dcf-a4d6-4073a43b8907 · inbound

Adaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, and Graph Structures cites this paper.

Adaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, and Graph Structures Reinforcement Learning Enhanced LLMs: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:49.495121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:49.495121Z digest=sha256:a3aa03ad919f40509699d84e97fb9ac422564dcbbc13845e9ee3a47eab37264d

Observation 3b203e7a-088c-4dbc-b3f4-92ca1cebad40 · inbound

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers cites this paper.

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers Reinforcement Learning Enhanced LLMs: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:12.936464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:08:12.936464Z digest=sha256:2e0af9db51363cde0e40bd5d871d28a27eedb882fd56b528303d9a8c86996fda

Observation 74c35b06-33bc-4769-a41d-ee2412b33a4e · inbound

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models cites this paper.

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:39.153862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:39.153862Z digest=sha256:188ccc56881c43a8be7b2c3e85e314c15f2a681e2c9df0e8deb76cf4ed2947a4

Observation b3f198b2-d25f-4a79-afe5-066bbddecdb4 · inbound

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance cites this paper.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Reinforcement Learning Enhanced LLMs: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.611764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.611764Z digest=sha256:22f0318ce1c9520b06b409583ce5b40e223414e5396d8559d73dc67913cdc7f6

Observation 271c3ef1-143a-4feb-a781-27cf896f307c · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:40.631750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:40.631750Z digest=sha256:d63f5190d8a29a2faaa84fcc93313dbe70bcabc13bd692543ca8b3fcface5942

Observation 92ac68ef-15bb-45df-83c8-f1635466a8a7 · inbound

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities cites this paper.

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities Reinforcement Learning Enhanced LLMs: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:45.293601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:45.293601Z digest=sha256:20b00c58e005457295bff2bb0a8f2fcb24d6dc63cfb1b55989606bee3936c999

Observation 8d61cdcc-f5b3-4f50-ac9f-63d1b5addaf0 · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:35.032860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:35.032860Z digest=sha256:eafab6eb2defd23ac4b9bcc64f6d4111201d174abe20176f2d05010381079bf8

Observation 4ac75bcc-2484-4701-8ee5-7d2205ab938f · inbound

UrbanMind: Towards Urban General Intelligence via Tool-Enhanced Retrieval-Augmented Generation and Multilevel Optimization cites this paper.

UrbanMind: Towards Urban General Intelligence via Tool-Enhanced Retrieval-Augmented Generation and Multilevel Optimization Reinforcement Learning Enhanced LLMs: A Survey

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:56.908088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:56.908088Z digest=sha256:d08901dfc587d87f532724a2db8de352a20f50f8bd2406492c2a5a99fc1eb9be

Observation 8454d102-a1f9-4778-ac7f-ca29f592db15 · inbound

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning cites this paper.

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning Reinforcement Learning Enhanced LLMs: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:45.167344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:45.167344Z digest=sha256:f645398734330fc0e0ad07ba4ae138a8879042ee6e92c089dff081d5b91b5c10

Observation d8a03833-3369-49cf-b682-b4d1928c1d0b · inbound

CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning cites this paper.

CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning Reinforcement Learning Enhanced LLMs: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:15:03.669635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:15:03.669635Z digest=sha256:d6ba4e731b5ea1c8a996e46f501f62baaddb935fbe97169de09f90bb05591148

Observation c802a854-f512-4e56-88f1-ed39a82c88db · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning Enhanced LLMs: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.875324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c185f829f0e09b44e8f4e8088d3d3cb789fb2a2d9078f6c4ac282c387a15f251

Observation 4337c025-72a6-4375-9146-b13ac84b2cc5 · inbound

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models cites this paper.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.811953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.811953Z digest=sha256:745cacdbc512b4e5271f7e12b3ee55ce9c25567e6de7f3e9d6b88c4d528c7673

Observation 962487d1-4a09-403b-9316-9085a2f737fb · inbound

Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing cites this paper.

Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing Reinforcement Learning Enhanced LLMs: A Survey

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T20:49:37.841831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:49:37.841831Z digest=sha256:68c4514f94692782f4f387469a762d4c8f2dc8a986c3a0cb16b2acc5fda5edbb

Observation 924608cf-7715-43ae-b943-b3fbfcb71f8a · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning Enhanced LLMs: A Survey

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.090072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.090072Z digest=sha256:db481ab32a0ab043046f7fe1c9248fc730af92fdb32b3b81e5b9755add5a5666

Observation 1872373b-1b53-42ae-a144-2f468c9ec62b · inbound

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning cites this paper.

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning Reinforcement Learning Enhanced LLMs: A Survey

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:41.576325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:29:34.145855Z digest=sha256:ce9fc892fb8eb71d3c8cadc9225b7d474bec35d8ec813435ee216bc9733d7b8f

Observation bce1f1bb-2c66-4943-afcb-3754b3fa4315 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.688214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:4070deb82e38a7dcc73c75c8b40f2b96ea7f071b9b477ce6b81700abc89a8381

Observation 97ae5a59-580d-44ff-afe6-7d7aa208166d · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.839688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:c7e24f79cdbe49ce06e2458df45853a61c137a26dfbe0b7790432f2b2ef9c6c1

Observation 48e78007-5e78-49b1-a1c7-4e38121d42b0 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.669713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:16a5fe16aeb5aa2b4ebb764b44b5a581d872ca7d086233c86e269a2a27c60d94

Observation 331828b8-536a-4a04-a4b7-0e41a4240678 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcement Learning Enhanced LLMs: A Survey

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:48.880543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:1873ef0e08ad6efb644f82ece440242d04480a044ef23f377009554e6fe40d5c

Observation da193571-0479-4764-91d0-ae53fc401e22 · inbound

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective cites this paper.

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective Reinforcement Learning Enhanced LLMs: A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.836626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:52:51.509014Z digest=sha256:92d3384ae1f48631d3314ac6ec8f9246fb98051ebd9a32ec114d920c66db715b

Observation bcf3352b-0cc0-411e-83c8-bc518ee2c6d2 · inbound

Distributed Direct Preference Optimization cites this paper.

Distributed Direct Preference Optimization Reinforcement Learning Enhanced LLMs: A Survey

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:40.940560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:56:54.996304Z digest=sha256:bad718ac0d31a65ac36de917e901062254b6618f902ae29676fcee8ca7eddb72

Observation 51c7f75c-ae03-49f3-b4df-8b02fbed5ada · inbound

Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs cites this paper.

Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs Reinforcement Learning Enhanced LLMs: A Survey

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.399088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T20:48:59.539626Z digest=sha256:085a9a0f556ed71c3e072bea1e0256ca2f2086630fbd96c1b07e8d53c5023298

Observation a585e21d-8c8a-485a-9f16-1ca05051fd09 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Reinforcement Learning Enhanced LLMs: A Survey

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.176238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:b84403d4622941ceac572db49da98374678173477db07eac9fed6fb71ca2e8a5