Pith. sign in

Paper Citation Record · LEDGER

Deep Reinforcement Learning: From First Principles to Reasoning Models

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.00133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00133 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:16:05.392554Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cf5dcd9-b8a1-419b-aa3c-91e86c17e31a · outbound

This paper cites Genie: Generative Interactive Environments.

Deep Reinforcement Learning: From First Principles to Reasoning Models Genie: Generative Interactive Environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.032259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.032259Z digest=sha256:0c62078740d3989164aa7bf22ed2eb44f10725609f63b062e79cad4cc1170f92

Observation 5d729623-6732-4a19-b31b-0a77d438c5f2 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic for Discrete Action Settings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.255389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.255389Z digest=sha256:2d432040b33bcab2b75a418533b4ea8fc8a9a11234511e54b29e195f9cc99099

Observation 19a6c79a-b59f-43f6-8274-d9d792531765 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.904068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.904068Z digest=sha256:2f3f1571d7ba0cfe2095b6f9d2a28c1ec33e03f37e7d68e7ce71ac794ab77907

Observation 31225c62-a18c-46fb-88a0-af4f988b016a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Deep Reinforcement Learning: From First Principles to Reasoning Models Improving alignment of dialogue agents via targeted human judgements

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.233701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.233701Z digest=sha256:322d10d6aab881d403d2a613a124f69e48adddb1352515d1851df1f81d5fccdc

Observation a08dccf8-c088-4d9e-b847-916002338995 · outbound

This paper cites Learning Control Barrier Functions and their application in Reinforcement Learning: A Survey.

Deep Reinforcement Learning: From First Principles to Reasoning Models Learning Control Barrier Functions and their application in Reinforcement Learning: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.383193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.383193Z digest=sha256:704e4ca1e68598efb97d6a39ac49aa8631a9a4549bd00954d0ad93996e093d43

Observation de2b33a2-0b57-466f-8614-b71b6d803c74 · outbound

This paper cites World Models.

Deep Reinforcement Learning: From First Principles to Reasoning Models World Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.503990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.503990Z digest=sha256:4804ab4ed7c3dea99df3bb2a4d94d56f149de28859edaa3d62473a21159e4a03

Observation 660c4eec-04d9-462b-9a86-a13ec9e8fdda · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic Algorithms and Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.634487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.634487Z digest=sha256:bd2c948ceace96a7ba9ad7e43eed2186a08643d62a55ac1744beb8828271243a

Observation 939e08b6-7e87-40a3-b90f-c4bbb335b98a · outbound

This paper cites Mastering Diverse Domains through World Models.

Deep Reinforcement Learning: From First Principles to Reasoning Models Mastering Diverse Domains through World Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.795109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.795109Z digest=sha256:d24735f822c5a7d3b4f168db5f075ec9fada46359d33b7e1bd1dcf58a4778e1a

Observation ac4b6f60-c66a-4eaf-96f4-0e00bcbea1f6 · outbound

This paper cites Nicklas Hansen, Xiaolong Wang, and Hao Su.

Deep Reinforcement Learning: From First Principles to Reasoning Models Nicklas Hansen, Xiaolong Wang, and Hao Su

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.917358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.917358Z digest=sha256:c77c6f8ccc955f38996df5ae85aa64f0013894cedda61ab97cc78c650958962d

Observation d8ce6f9a-7e3d-42fd-89cc-5cd12670f129 · outbound

This paper cites Dropout Q-Functions for Doubly Efficient Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.055354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.055354Z digest=sha256:dbf533c90feed746c6424b6f633643223be2ad2a1653046031bf92cb1b501b0f

Observation e0381c50-8cf6-4358-b515-fd7660eb2a3e · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Deep Reinforcement Learning: From First Principles to Reasoning Models ORPO: Monolithic Preference Optimization without Reference Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.181961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.181961Z digest=sha256:156e5c01e01ef62f1430879234e02f619bd83964699111897aa5200fd0c995a9

Observation b90705c3-4260-4570-8d66-8578008d5b04 · outbound

This paper cites A Control Barrier Function-Constrained Model Predictive Control Framework for Safe Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models A Control Barrier Function-Constrained Model Predictive Control Framework for Safe Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.348797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.348797Z digest=sha256:e1f40d81f142a56f343baa45919d97a9cb4013473c46b57bf40948377f0cc698

Observation 003ce360-d3df-4b6e-a8b5-bfc5c88f9d4b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Deep Reinforcement Learning: From First Principles to Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.479731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.479731Z digest=sha256:cdff494eceb240cf0063dabbf4eacedb218555689b5f7c54026ad61073976873

Observation c36079bd-0e77-47aa-8dcc-6f668a865610 · outbound

This paper cites A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions.

Deep Reinforcement Learning: From First Principles to Reasoning Models A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.605676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.605676Z digest=sha256:c3be493314e5a58cd12fa0acc5a08f0bb36372824f22c58b8b45e1579f410a10

Observation f032c323-925b-4c94-a2a7-1ac27b426398 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Deep Reinforcement Learning: From First Principles to Reasoning Models Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.842764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.842764Z digest=sha256:ab5f22232caa15919dea81c39370707ee3281cc6b38b54274bd37b0745bb458b

Observation 2447f6b6-a2e2-4152-bfa4-e84abf25a818 · outbound

This paper cites Continuous control with deep reinforcement learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Continuous control with deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.984462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.984462Z digest=sha256:d385741b91a998507aae4974ca5be86ac4b1dacdcfe5d93f9cd2c37430d913e3

Observation 37e71f77-bbbb-44d0-91f7-57e680ea3b4e · outbound

This paper cites Online finetuning decision transformers with pure reinforcement learning gradients.arXiv preprint arXiv:2601.00167,.

Deep Reinforcement Learning: From First Principles to Reasoning Models Online finetuning decision transformers with pure reinforcement learning gradients.arXiv preprint arXiv:2601.00167,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.103819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.103819Z digest=sha256:5ea7712588050e237220499ae854bad8bf5716191392fa70ea5bb2c05c7c1e05

Observation e3ccb85a-90fc-470d-be34-1ff702fb6851 · outbound

This paper cites Efficient soft actor-critic with LLM-based action-level guidance for continuous control.arXiv preprint arXiv:2603.17468,.

Deep Reinforcement Learning: From First Principles to Reasoning Models Efficient soft actor-critic with LLM-based action-level guidance for continuous control.arXiv preprint arXiv:2603.17468,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.214093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.214093Z digest=sha256:3af1e9cddd6a75baa263a1dc87ce4a151e8f70ae7a406db51b2d7e04c6f5ea49

Observation 1982bf56-417a-4cf6-96dd-b3a3fcb8088c · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Deep Reinforcement Learning: From First Principles to Reasoning Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.318050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.318050Z digest=sha256:fe02930fa92b24fc0b6716fc1c0ee7200e4359c2a21a359e53ef60d1a825dde0

Observation 0be68d9f-2ae8-4ba6-ae08-d67bf38cf296 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Deep Reinforcement Learning: From First Principles to Reasoning Models AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.481502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.481502Z digest=sha256:b19d8cd517cec37215179f4bbfd406aef1521aee07a51f8d453ea1320010f3c7

Observation 0fb2f258-a3a1-47d9-bf4b-d0edba7f3b23 · outbound

This paper cites How to train your latent control barrier function.arXiv preprint arXiv:2511.18606,.

Deep Reinforcement Learning: From First Principles to Reasoning Models How to train your latent control barrier function.arXiv preprint arXiv:2511.18606,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.603115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.603115Z digest=sha256:56bdb1e800af00deb36e0431e04b7ff73bdfa3bd733c3c5a1c718a09d4f4c2ba

Observation eba33d7a-9c7d-4e48-911d-9ecc8dcb641b · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.719913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.719913Z digest=sha256:06369e91e2fc03c472fb68da6a523a8dbe61166a6862072787d44db7126d44e2

Observation 20b1dd58-bc4b-4dfb-95d6-1c4d4426f72d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Deep Reinforcement Learning: From First Principles to Reasoning Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.861262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.861262Z digest=sha256:3eb6f7a5b9566a3cdf0eae794f9de9aab0ae6bb596fd2ba86abde36900a1b070

Observation c2111a3b-d4cc-4667-9bf6-50479691f8eb · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Deep Reinforcement Learning: From First Principles to Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.047614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.047614Z digest=sha256:d1bbe35c466f8e1970672d25d3f3e31952a818743069b21ed54d9b51f3dc7224

Observation 7d9d1167-8c6e-4b9e-ad31-8ac7a273542f · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

Deep Reinforcement Learning: From First Principles to Reasoning Models Towards Understanding Sycophancy in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.213789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.213789Z digest=sha256:b11aadff90fc01ff8d719f205209ca2517837b1da3031da10c5d4327318afd2a

Observation 81b3cc78-0585-4d7a-b578-47d9d18788a3 · outbound

This paper cites Defining and Characterizing Reward Hacking.

Deep Reinforcement Learning: From First Principles to Reasoning Models Defining and Characterizing Reward Hacking

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.347683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.347683Z digest=sha256:a192e4c49c6d2c13d2de04efd5597148ac20f4111ebed58d0eadd0fa758a7275

Observation f028ebc4-b306-4de3-b1a3-e632691dbc5f · outbound

This paper cites Target Return Optimizer for Multi-Game Decision Transformer.

Deep Reinforcement Learning: From First Principles to Reasoning Models Target Return Optimizer for Multi-Game Decision Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.426467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.426467Z digest=sha256:24a20b8fa802391d2cd4c3e5a4077f70f8c5ae8cd921038b558eb8ba79fcfd7b

Observation cd40a34f-c338-47fc-83fb-c2ffc2a069b5 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Deep Reinforcement Learning: From First Principles to Reasoning Models Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.504223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.504223Z digest=sha256:414f0c106501902862123c965484ba0a22596fd25b9ab356e893a616b8c516e4

Observation 86f53e1a-e55a-4873-8bf7-d53c20bd42a8 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Deep Reinforcement Learning: From First Principles to Reasoning Models Deep Reinforcement Learning and the Deadly Triad

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.573522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.573522Z digest=sha256:8b29dece01e1d41b6b464f30071f3c732e4ed373fcdfd3d866721435e7747784

Observation 32df2db8-2384-4a95-ad7f-c374e6cba6e4 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Behavior Regularized Offline Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.732902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.732902Z digest=sha256:1a2242487e011829347122bf34772c538f98b6bab2137b6a229ec38efcf267cb

Observation 0f70470f-965c-4888-9572-6e5548bbf852 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Deep Reinforcement Learning: From First Principles to Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.833754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.833754Z digest=sha256:dbbda57fd5bf77b942365ada900c5fc206c2bf2af8858d6178cb43f38996a699

Observation c55b42e7-558f-4eba-b3d0-6d259ae5a02c · outbound

This paper cites CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions.

Deep Reinforcement Learning: From First Principles to Reasoning Models CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.938739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.938739Z digest=sha256:1a7e261faa1922718bdd9764c81b6a410b914897335e5f61df92c2aa4f1b218c

Observation 143bfbb0-2a4a-4c98-989e-e08eb895f6c5 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Deep Reinforcement Learning: From First Principles to Reasoning Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.018389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.018389Z digest=sha256:dab0447a68e3697510d013bd6210230074efa55513f88b0c4462670a5bf43d14

Observation 48e765bb-755c-4a4d-b144-67fca8c8c5da · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Deep Reinforcement Learning: From First Principles to Reasoning Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.118218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.118218Z digest=sha256:b668b64992bb86c6e6e19f3a126b223c33669452565a9f765971162700dcd2a8

Observation 4c85bfbf-931a-4c06-855b-513c0439af37 · outbound

This paper cites A Deeper Look at Experience Replay.

Deep Reinforcement Learning: From First Principles to Reasoning Models A Deeper Look at Experience Replay

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.188515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.188515Z digest=sha256:f8ed4d5e2afaf2ce9d00a07107a24fc95fce5c674e001ed8f25e2399ac9a8907

Observation 24175e85-901a-43af-9389-604fd30d8ee7 · outbound

This paper cites Revisiting Discrete Soft Actor-Critic.

Deep Reinforcement Learning: From First Principles to Reasoning Models Revisiting Discrete Soft Actor-Critic

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.266717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.266717Z digest=sha256:6d0c57173430099a1f1e4d8476412219f8f3cdda080c6c17fcb6445c7f802855

Observation c24103c3-e6bd-4024-95f5-f708b29fdf6f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Deep Reinforcement Learning: From First Principles to Reasoning Models Fine-Tuning Language Models from Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.392554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.392554Z digest=sha256:06e465a4004299e90d2c76139eca978f6261eec9d6692662fcc1e394728b6486

Observation 92e11276-cd1e-490d-8e1a-aedfd8134572 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.626146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.626146Z digest=sha256:903b7ce9f531a0d701f9a00011e5725feb33f9b00fe521c6dc00dc36a1b2100f

Observation c61d01a7-d41e-460e-8973-022a07b7e244 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Deep Reinforcement Learning: From First Principles to Reasoning Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.721065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.721065Z digest=sha256:c73e5b3ab6ad8300b85a2b0152e6bef6d371edde175af5ba14acd0341e689e30

Observation aadab0d4-40bf-47b7-b095-e1c44389b842 · outbound

This paper cites Delgrange et al.

Deep Reinforcement Learning: From First Principles to Reasoning Models Delgrange et al

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.759217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.759217Z digest=sha256:b45325f9912951c76e2a2659c15f5ec8c1ea42df8f1b3f46d4f99c2375589960

Observation 2988e363-41f8-479f-90ad-9503baae70d0 · outbound

This paper cites Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:01.059094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:01.059094Z digest=sha256:99ab90ea58dadd0e838e10693c2644f9ee3774b886ccf57d468ccfd3cac830af

Observation 4fea863f-6f42-45cc-b4a4-e1fa5487bc6d · outbound

This paper cites Concrete Problems in AI Safety.

Deep Reinforcement Learning: From First Principles to Reasoning Models Concrete Problems in AI Safety

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T01:15:59.412376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:15:59.412376Z digest=sha256:20dc9944798815d5f5342a14559c0ebcf36f6db683fe455c001deb667f875258

Observation 99bc4337-08ae-461a-9bd6-1a2396b92e0c · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Deep Reinforcement Learning: From First Principles to Reasoning Models Constitutional AI: Harmlessness from AI Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T01:15:59.641823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:15:59.641823Z digest=sha256:0cbffb9ec13c9a32412092898860246c265b34c00d1483dbee38c907c099c8bf

Observation dd6d564e-4faf-4ee3-81ba-5f8d7d12e922 · outbound

This paper cites 2024 ACM a.m.

Deep Reinforcement Learning: From First Principles to Reasoning Models 2024 ACM a.m

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T01:15:59.525068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:15:59.525068Z digest=sha256:54e8eec69fb67687e19d6f1cd3805b2b6a70747a781cf9ce453d6f07be51287b

Observation 1888d178-3b4b-42cf-ad81-6866b1b56b89 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Deep Reinforcement Learning: From First Principles to Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.362479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.362479Z digest=sha256:1113f41bcb34f6fcd6aa816c52de5d6883048230eb95028afac73c0ce2d804c3

Observation 5a6e3e87-58e6-4cb1-a27d-90494072aec9 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Deep Reinforcement Learning: From First Principles to Reasoning Models Process Reinforcement through Implicit Rewards

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.507444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.507444Z digest=sha256:45ed54520a5742e774470e1ac8620f54834b18060d43ad78c21596d6abb17a79

Observation e73967b3-8b6f-4d3a-a30b-9f3235461112 · outbound

This paper cites OpenAI Gym.

Deep Reinforcement Learning: From First Principles to Reasoning Models OpenAI Gym

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T01:15:59.863844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:15:59.863844Z digest=sha256:bff837a8e1dce647475fc88af4296107a40cb02b50834aac8473bb55b5c9b85e

Observation 65176c4b-f0e2-447f-8ebf-363b34b5ffdf · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.653728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.653728Z digest=sha256:289b278610e4091a2c28ef2220569b16fc816d364e6d9ca8d54f7cd7a580209c

Observation bb622e9c-e10a-4359-bdc1-e8dd456c3d5e · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Deep Reinforcement Learning: From First Principles to Reasoning Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.139959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.139959Z digest=sha256:5bf124492553d8e442afa21f4f8b9647ece2add309c85c25752457fe498b1591

Observation 009c66c3-36ad-4803-b635-9cdf1538379b · outbound

This paper cites arXiv preprint.

Deep Reinforcement Learning: From First Principles to Reasoning Models arXiv preprint

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T01:15:59.757963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:15:59.757963Z digest=sha256:89ed5ef9d6a671ebefcc82eae1973e1b4cdeb5afbbd824242ba5658a710ab11c

Observation b13b507c-def0-489c-8653-a2b950219aed · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Deep Reinforcement Learning: From First Principles to Reasoning Models Playing Atari with Deep Reinforcement Learning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:03.400506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:03.400506Z digest=sha256:289a1e2a4991ff50881596156562573e1cbbadb4adc84b4c5374643070e19dbe

Pith citing papers

No inbound Pith citation observations are available.