Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:57.673384Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2505.11081.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:57.673384Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T15:55:45.589124Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.634737Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 96bc527e-1924-4ed5-899f-78b34614b53e · outbound
ShiQ: Bringing back Bellman to LLMs A learning algorithm for boltzmann machines.Cognitive science, 9(1):147–169, 1985
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022836f6-d4e3-417e-bedf-3e606660425c · outbound
ShiQ: Bringing back Bellman to LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78145be-ec0b-4216-ad7d-683137e28b51 · outbound
ShiQ: Bringing back Bellman to LLMs A general theoretical paradigm to understand learning from human preferences
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3055a8d9-9ab5-4f18-a28d-9fe7d258d290 · outbound
ShiQ: Bringing back Bellman to LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb6a178d-74de-4e94-9dd6-e05e1e01a526 · outbound
ShiQ: Bringing back Bellman to LLMs Residual algorithms: Reinforcement learning with function approximation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbc3035-ec37-4dd1-8425-63e714c4e2a6 · outbound
ShiQ: Bringing back Bellman to LLMs Linear least-squares algorithms for temporal difference learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1108bc0-169c-41a0-854e-0abcb65ef2ba · outbound
ShiQ: Bringing back Bellman to LLMs Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf27859-184c-430f-851d-a484ccfc99ed · outbound
ShiQ: Bringing back Bellman to LLMs Command A: An Enterprise-Ready Large Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82fcb595-98d5-40a0-bef5-ac256bebfadd · outbound
ShiQ: Bringing back Bellman to LLMs Ultrafeedback: Boosting language models with high-quality feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39953d5b-35fd-4628-af56-85097a7b5ccb · outbound
ShiQ: Bringing back Bellman to LLMs Off-policy actor-critic
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1b13d439-6bce-4319-af00-4e9ba47dddf9 · outbound
ShiQ: Bringing back Bellman to LLMs KTO: Model Alignment as Prospect Theoretic Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23227b70-862b-4e9c-a2cc-d948c9e3472e · outbound
ShiQ: Bringing back Bellman to LLMs Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1a4af39-ee1d-404d-ba1b-9a38153ca2c2 · outbound
ShiQ: Bringing back Bellman to LLMs Addressing function approximation error in actor-critic methods
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03cc769b-3d6b-4332-be6e-40c227e85f9e · outbound
ShiQ: Bringing back Bellman to LLMs Is the bellman residual a bad proxy?Advances in Neural Information Processing Systems, 30, 2017
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a9ab6012-21e4-402a-945d-2ee1c13f9281 · outbound
ShiQ: Bringing back Bellman to LLMs A theory of regularized markov decision processes
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 926f8381-b10b-47f2-b092-ae0fc4c8be2c · outbound
ShiQ: Bringing back Bellman to LLMs Averaging log-likelihoods in direct alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7841342c-d83f-4e51-8577-1a5e7aff03a6 · outbound
ShiQ: Bringing back Bellman to LLMs Efficient (soft) q-learning for text generation with limited good data
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c8d8304b-f413-4196-b281-102636062b26 · outbound
ShiQ: Bringing back Bellman to LLMs Rainbow: Combining improvements in deep reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f9ae8ed-8cc0-4579-98b0-39027f78708e · outbound
ShiQ: Bringing back Bellman to LLMs The Curious Case of Neural Text Degeneration
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e731ea-e33b-4d1a-b2ac-aa81f4b0a581 · outbound
ShiQ: Bringing back Bellman to LLMs Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8b128c-b30f-4025-8bfc-aca1ec19eb6b · outbound
ShiQ: Bringing back Bellman to LLMs Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f8d21c-88ef-48f9-8f51-a387710ca392 · outbound
ShiQ: Bringing back Bellman to LLMs Buy 4 reinforce samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction (ICLR Workshop), 2019
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 18b626a5-f6de-4ed2-a6b5-75372c95bf39 · outbound
ShiQ: Bringing back Bellman to LLMs Offline reinforcement learning with implicit q-learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d01fb1bf-eab7-4558-9fa1-8110c78f4171 · outbound
ShiQ: Bringing back Bellman to LLMs Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328, 2022
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8ee31d-38fc-4f4e-a2a0-dde7bd755e87 · outbound
ShiQ: Bringing back Bellman to LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19866da0-4419-4c82-9b34-aacf251dae46 · outbound
ShiQ: Bringing back Bellman to LLMs Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d67656b-899e-45f6-8477-1b55cc5f9f80 · outbound
ShiQ: Bringing back Bellman to LLMs Safe and efficient off-policy reinforcement learning.Advances in neural information processing systems, 29, 2016
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d7aed369-8a7f-4d2d-b39c-59cbe960d7f5 · outbound
ShiQ: Bringing back Bellman to LLMs Bridging the gap between value and policy based reinforcement learning.Advances in neural information processing systems, 30, 2017
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2f0cecf-46b6-4c59-85dc-37550bbca7e5 · outbound
ShiQ: Bringing back Bellman to LLMs Policy invariance under reward transformations: Theory and application to reward shaping
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7b96bb31-8a46-446e-a0ae-24b5043a707b · outbound
ShiQ: Bringing back Bellman to LLMs Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4789803-cb77-42f6-b5a6-86680bec27ff · outbound
ShiQ: Bringing back Bellman to LLMs Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7dfce9-c30a-4e81-9231-aef6ca36efa0 · outbound
ShiQ: Bringing back Bellman to LLMs Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a54fcf35-98e9-4c86-a287-e57a05b36e2d · outbound
ShiQ: Bringing back Bellman to LLMs From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc0837a-f55f-43ba-b31e-9718bc8d6d7d · outbound
ShiQ: Bringing back Bellman to LLMs Offline Regularised Reinforcement Learning for Large Language Models Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fdd4b4d-02c6-4e73-b278-efb41499f604 · outbound
ShiQ: Bringing back Bellman to LLMs Factually consistent summarization via reinforcement learning with textual entailment feedback
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bcb36d6c-bbef-405f-b719-f0586d4e4ccf · outbound
ShiQ: Bringing back Bellman to LLMs Approx- imate modified policy iteration and its application to the game of tetris.J
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 213c66aa-c1af-481d-9019-40c64fbe60cf · outbound
ShiQ: Bringing back Bellman to LLMs Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85991ec-24c4-476d-b750-246c340cbfee · outbound
ShiQ: Bringing back Bellman to LLMs Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb689e5-5810-4219-a172-3681e5b10f94 · outbound
ShiQ: Bringing back Bellman to LLMs Offline RL for natural language generation with implicit language q learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3b0adb19-4dbd-4282-aa01-6293c0b7be72 · outbound
ShiQ: Bringing back Bellman to LLMs Generalized preference optimization: A unified approach to offline alignment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d4e4e0e5-f2b5-4ddd-bbf6-eee1db7c913d · outbound
ShiQ: Bringing back Bellman to LLMs RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2927fd-5aa3-4949-8f0f-b7a9482c7e55 · outbound
ShiQ: Bringing back Bellman to LLMs Leverage the average: an analysis of kl regularization in reinforcement learning.Advances in Neural Information Processing Systems, 33:12163–12174, 2020
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 077620c5-8f80-437f-aeaf-373d6e7e9b1e · outbound
ShiQ: Bringing back Bellman to LLMs Munchausen reinforcement learning.Advances in Neural Information Processing Systems, 33:4235–4246, 2020
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5d2d18e3-512f-4d79-8f00-ef19195b33f8 · outbound
ShiQ: Bringing back Bellman to LLMs Dueling network architectures for deep reinforcement learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58f1446-8375-4031-87a3-1d57a67489d5 · outbound
ShiQ: Bringing back Bellman to LLMs Potential-based shaping and q-value initialization are equivalent.Journal of Artificial Intelligence Research, 19:205–208, 2003
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5b979ec3-bb05-4a3c-a3e8-793c27b5aa84 · outbound
ShiQ: Bringing back Bellman to LLMs Function optimization using connectionist reinforcement learning algorithms.Connection Science, 3(3):241–268, 1991
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a82447f1-a91f-42cb-aa73-3b7e5c467616 · outbound
ShiQ: Bringing back Bellman to LLMs Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0bb9dc-e5d6-43fd-a9a8-d6514018fe9f · outbound
ShiQ: Bringing back Bellman to LLMs InThe Twelfth International Conference on Learning Representations, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6372540e-0a55-4552-b6ba-acefa4250a03 · outbound
ShiQ: Bringing back Bellman to LLMs Adam-mini: Use Fewer Learning Rates To Gain More
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbe1492-80e1-46c9-90d7-5d7cc68d3708 · outbound
ShiQ: Bringing back Bellman to LLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c88e6f-9758-4e5a-ba60-8ab220d3031b · outbound
ShiQ: Bringing back Bellman to LLMs ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc62e57-a085-422d-a0f6-ddefbeed1118 · outbound
ShiQ: Bringing back Bellman to LLMs PhD thesis, Carnegie Mellon University, 2010
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 940fde9c-4619-400c-a29f-6a51a3203b6f · outbound
ShiQ: Bringing back Bellman to LLMs Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebcac538-19f2-45bf-bd52-700034ba3dd1 · outbound
ShiQ: Bringing back Bellman to LLMs Initialization trick
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 683665a4-f333-4b52-a187-eda081d5cf04 · outbound
ShiQ: Bringing back Bellman to LLMs This contrasts with other RL-finetuning approaches, such as DRO [ 34] or CoPG [ 12], that involve a square term per sequence of the batch
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 638e2006-da73-4454-b1aa-9c417e26ecd1 · outbound
ShiQ: Bringing back Bellman to LLMs These tools use machine learning algorithms to generate realistic voices and faces
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 108adfb6-2a4c-4e93-9832-893323e94382 · outbound
ShiQ: Bringing back Bellman to LLMs The more reference audio you have, the better the result
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 21c39961-af15-46bf-8967-c76c9f35c729 · outbound
ShiQ: Bringing back Bellman to LLMs This process may take some time, depending on the complexity of the task and the 35
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93209482-94e8-4983-a010-6f5035bca00a · inbound
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction ShiQ: Bringing back Bellman to LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 97d5248d-0f2e-4647-ad0c-b1edf88dc07f · inbound
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction ShiQ: Bringing back Bellman to LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b304f723-78b5-4404-8aac-290f8627d5f8 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ShiQ: Bringing back Bellman to LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.