Pith. sign in

Paper Citation Record · LEDGER

ShiQ: Bringing back Bellman to LLMs

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2505.11081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11081 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:57.673384Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T15:55:45.589124Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.634737Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96bc527e-1924-4ed5-899f-78b34614b53e · outbound

This paper cites A learning algorithm for boltzmann machines.Cognitive science, 9(1):147–169, 1985.

ShiQ: Bringing back Bellman to LLMs A learning algorithm for boltzmann machines.Cognitive science, 9(1):147–169, 1985

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.455667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.455667Z digest=sha256:a9f52afcc1925fd7710927a430bb5663d57202c048e077e6da95c781f25e0633

Observation 022836f6-d4e3-417e-bedf-3e606660425c · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

ShiQ: Bringing back Bellman to LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.460970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.460970Z digest=sha256:99bd8912ccf40e4793a7da3d38434a6cfe0bd006b83bd448fad0fcbf9accdc96

Observation b78145be-ec0b-4216-ad7d-683137e28b51 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

ShiQ: Bringing back Bellman to LLMs A general theoretical paradigm to understand learning from human preferences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.300968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.465428Z digest=sha256:392c0c6899490e0860fdec71ff9cd3da6dc841aa9ffcf588857768f392153569

Observation 3055a8d9-9ab5-4f18-a28d-9fe7d258d290 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

ShiQ: Bringing back Bellman to LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.469525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.469525Z digest=sha256:fb7a481e8586a05b269c2b9752daa7cdf9a34cf82948029721e049727f026030

Observation fb6a178d-74de-4e94-9dd6-e05e1e01a526 · outbound

This paper cites Residual algorithms: Reinforcement learning with function approximation.

ShiQ: Bringing back Bellman to LLMs Residual algorithms: Reinforcement learning with function approximation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.473289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.473289Z digest=sha256:c1fbb93ddbb2f5d4a0edccb58694ecdfc06ec4c7e9dc393d94114ddbf409d751

Observation bbbc3035-ec37-4dd1-8425-63e714c4e2a6 · outbound

This paper cites Linear least-squares algorithms for temporal difference learning.

ShiQ: Bringing back Bellman to LLMs Linear least-squares algorithms for temporal difference learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.281123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.477125Z digest=sha256:094c535eed57183666b2281426c8c2d76962f4683c9c4b478856bafee72c63a5

Observation a1108bc0-169c-41a0-854e-0abcb65ef2ba · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

ShiQ: Bringing back Bellman to LLMs Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.481630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.481630Z digest=sha256:076b79cfa2e90bd63c2c9a8f4dfbd794a7d1ce4e609071b149c2b1efb99b11db

Observation baf27859-184c-430f-851d-a484ccfc99ed · outbound

This paper cites Command A: An Enterprise-Ready Large Language Model.

ShiQ: Bringing back Bellman to LLMs Command A: An Enterprise-Ready Large Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.485815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.485815Z digest=sha256:97a52c643f8bf9079fc8a84a2d70cc46ac4ebcb83b12b4c2bf63e48a4cdfe3b2

Observation 82fcb595-98d5-40a0-bef5-ac256bebfadd · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

ShiQ: Bringing back Bellman to LLMs Ultrafeedback: Boosting language models with high-quality feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.489789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.489789Z digest=sha256:96aa3ba44e93b99fd5f658cd0a9f8cf762c576bb65d493a87a049cf52f85a3c5

Observation 39953d5b-35fd-4628-af56-85097a7b5ccb · outbound

This paper cites Off-policy actor-critic.

ShiQ: Bringing back Bellman to LLMs Off-policy actor-critic

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.257084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.493246Z digest=sha256:bfe329f272fa0bfb0c3b7ffe7084c0deacad15fc6c89bdf84a92875219329ba3

Observation 1b13d439-6bce-4319-af00-4e9ba47dddf9 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ShiQ: Bringing back Bellman to LLMs KTO: Model Alignment as Prospect Theoretic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.496903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.496903Z digest=sha256:f6d00e54be955c8854e43428b10880969f1cbc2d4fe5e0f96ca577879a76a723

Observation 23227b70-862b-4e9c-a2cc-d948c9e3472e · outbound

This paper cites Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion.

ShiQ: Bringing back Bellman to LLMs Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.244639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.501190Z digest=sha256:36683af06ff98360be97474272b39affb1d7269f4eff78305a348763d86f5876

Observation c1a4af39-ee1d-404d-ba1b-9a38153ca2c2 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

ShiQ: Bringing back Bellman to LLMs Addressing function approximation error in actor-critic methods

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.504944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.504944Z digest=sha256:ba6f5f8479d57937e0c4126c6e9c677087a44124da2bffd1acf839b76c04a2ef

Observation 03cc769b-3d6b-4332-be6e-40c227e85f9e · outbound

This paper cites Is the bellman residual a bad proxy?Advances in Neural Information Processing Systems, 30, 2017.

ShiQ: Bringing back Bellman to LLMs Is the bellman residual a bad proxy?Advances in Neural Information Processing Systems, 30, 2017

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.225965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.508495Z digest=sha256:fdc11efcc87434d1c59fb331ed75bbd49278d4e7f59e988d4dfa91010482bf48

Observation a9ab6012-21e4-402a-945d-2ee1c13f9281 · outbound

This paper cites A theory of regularized markov decision processes.

ShiQ: Bringing back Bellman to LLMs A theory of regularized markov decision processes

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.213433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.511648Z digest=sha256:9b0cf7910e542e9b3545d45677cb00619a41a826e74e44c16c631c199bd19e66

Observation 926f8381-b10b-47f2-b092-ae0fc4c8be2c · outbound

This paper cites Averaging log-likelihoods in direct alignment.

ShiQ: Bringing back Bellman to LLMs Averaging log-likelihoods in direct alignment

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:10:57.843680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.515108Z digest=sha256:6798e16944f15e5311de7936e4e3af4fb76445e6f9b2c1559646dc81f71264e4

Observation 7841342c-d83f-4e51-8577-1a5e7aff03a6 · outbound

This paper cites Efficient (soft) q-learning for text generation with limited good data.

ShiQ: Bringing back Bellman to LLMs Efficient (soft) q-learning for text generation with limited good data

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.202042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.518966Z digest=sha256:2082811e591083802e06852c5582414abc96d5ecc30b51e1074a25c46b9fc1cf

Observation c8d8304b-f413-4196-b281-102636062b26 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

ShiQ: Bringing back Bellman to LLMs Rainbow: Combining improvements in deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.190200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.522717Z digest=sha256:0d6a87a9759ba5fe102f39672bff3ae5a47ec1c3ebd0a8c22fd729e1385eaf59

Observation 8f9ae8ed-8cc0-4579-98b0-39027f78708e · outbound

This paper cites The Curious Case of Neural Text Degeneration.

ShiQ: Bringing back Bellman to LLMs The Curious Case of Neural Text Degeneration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.525931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.525931Z digest=sha256:b865c437cecaeaa7c41ba9f38cc894e14274b9655e8c5e094a93f5d05fb09a74

Observation a5e731ea-e33b-4d1a-b2ac-aa81f4b0a581 · outbound

This paper cites Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning.

ShiQ: Bringing back Bellman to LLMs Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.529911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.529911Z digest=sha256:1fdb61a4a7f32e26a52939b07210b791b189260773a209a52f10ac3090766a99

Observation 7e8b128c-b30f-4025-8bfc-aca1ec19eb6b · outbound

This paper cites Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization.

ShiQ: Bringing back Bellman to LLMs Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.533374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.533374Z digest=sha256:53db5a0ac468d8e0367646ecc749927a8ad91f38d640de17a6ba66a81a6b5865

Observation b0f8d21c-88ef-48f9-8f51-a387710ca392 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction (ICLR Workshop), 2019.

ShiQ: Bringing back Bellman to LLMs Buy 4 reinforce samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction (ICLR Workshop), 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.178590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.537151Z digest=sha256:5199fed86eaf3387e4704886692a14e3ed148475478240e01d0ca52373ab5f72

Observation 18b626a5-f6de-4ed2-a6b5-75372c95bf39 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

ShiQ: Bringing back Bellman to LLMs Offline reinforcement learning with implicit q-learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.165611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.540652Z digest=sha256:7f8a060b46da9eadbc6d172ffe0c5888703f6143b7e9b6bf9497ff90f9fdb37c

Observation d01fb1bf-eab7-4558-9fa1-8110c78f4171 · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328, 2022.

ShiQ: Bringing back Bellman to LLMs Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328, 2022

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.544320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.544320Z digest=sha256:904888d4ad10793270c7ead0f2b3e49328bbfe0e41cbe60cf5826817a3f5ad52

Observation ee8ee31d-38fc-4f4e-a2a0-dde7bd755e87 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

ShiQ: Bringing back Bellman to LLMs SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.547830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.547830Z digest=sha256:67156b9f8481d8ea4621bc54098501c9a5abbb20b79ecf8fadc601b8440db1dc

Observation 19866da0-4419-4c82-9b34-aacf251dae46 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

ShiQ: Bringing back Bellman to LLMs Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.551783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.551783Z digest=sha256:8df5ad91a3336ffb5ddd077246f337d0c275f1dd1db069406f13a14189b12810

Observation 5d67656b-899e-45f6-8477-1b55cc5f9f80 · outbound

This paper cites Safe and efficient off-policy reinforcement learning.Advances in neural information processing systems, 29, 2016.

ShiQ: Bringing back Bellman to LLMs Safe and efficient off-policy reinforcement learning.Advances in neural information processing systems, 29, 2016

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.137184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.555210Z digest=sha256:44a7e0bce98a8d5accb652c098527d84d1b5f00d810e3b35c84c27d9dae20c1c

Observation d7aed369-8a7f-4d2d-b39c-59cbe960d7f5 · outbound

This paper cites Bridging the gap between value and policy based reinforcement learning.Advances in neural information processing systems, 30, 2017.

ShiQ: Bringing back Bellman to LLMs Bridging the gap between value and policy based reinforcement learning.Advances in neural information processing systems, 30, 2017

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.124138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.559081Z digest=sha256:85bd6b568a4c74702aa30eadf72d2a85f7998db622ea2d14a372ac2931d242fc

Observation e2f0cecf-46b6-4c59-85dc-37550bbca7e5 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

ShiQ: Bringing back Bellman to LLMs Policy invariance under reward transformations: Theory and application to reward shaping

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.111655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.562577Z digest=sha256:ac0d5351bc41d472dc2f687b0aa7b428abf1e1bbbf7552ae92a74c61b1b7a311

Observation 7b96bb31-8a46-446e-a0ae-24b5043a707b · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

ShiQ: Bringing back Bellman to LLMs Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.566791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.566791Z digest=sha256:1b49dea577c7055164c153dcbf2299ab3abbfcddc2c5d7cc6312a78ee083b5bc

Observation b4789803-cb77-42f6-b5a6-86680bec27ff · outbound

This paper cites Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024.

ShiQ: Bringing back Bellman to LLMs Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.571105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.571105Z digest=sha256:49ab0101358f5b333fa4b233fb0565847bafffb4163721f828dd8c346b854775

Observation 8c7dfce9-c30a-4e81-9231-aef6ca36efa0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2023.

ShiQ: Bringing back Bellman to LLMs Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.086669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.574896Z digest=sha256:534777346f9ea60662928dcd9c7ec621452a8d81b399dd27bb6e45d7ac33de42

Observation a54fcf35-98e9-4c86-a287-e57a05b36e2d · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

ShiQ: Bringing back Bellman to LLMs From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.579584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.579584Z digest=sha256:ac780310c795688e7a3a7610bf2af1b7ca6590c5c724b5cdc1e5c33e5a0131e2

Observation fcc0837a-f55f-43ba-b31e-9718bc8d6d7d · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

ShiQ: Bringing back Bellman to LLMs Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.583738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.583738Z digest=sha256:7e15b1a0fbec46c2d74f980e5ce6389a33f48698e1ed765e4595904af979833e

Observation 0fdd4b4d-02c6-4e73-b278-efb41499f604 · outbound

This paper cites Factually consistent summarization via reinforcement learning with textual entailment feedback.

ShiQ: Bringing back Bellman to LLMs Factually consistent summarization via reinforcement learning with textual entailment feedback

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.075496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.587091Z digest=sha256:3c60610ca16753e1392354081546695a6376f8bec77c1a3ebb9ed64a847f9ab2

Observation bcb36d6c-bbef-405f-b719-f0586d4e4ccf · outbound

This paper cites Approx- imate modified policy iteration and its application to the game of tetris.J.

ShiQ: Bringing back Bellman to LLMs Approx- imate modified policy iteration and its application to the game of tetris.J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.063154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.590663Z digest=sha256:74368a0a374cb13fde7a23b149e1634f66bdbd0d9ac2dc850a71cebe619784b5

Observation 213c66aa-c1af-481d-9019-40c64fbe60cf · outbound

This paper cites Proximal Policy Optimization Algorithms.

ShiQ: Bringing back Bellman to LLMs Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.594034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.594034Z digest=sha256:2f024ca2bcb3fbb9d772ea01b63f0cd48f7d642d43accff617c4e2d32ade3b51

Observation d85991ec-24c4-476d-b750-246c340cbfee · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

ShiQ: Bringing back Bellman to LLMs Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.597496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.597496Z digest=sha256:2963841b9f27a36344cf83a1bd24809f15c2cf5292dc736c066a074bcc18d296

Observation 9cb689e5-5810-4219-a172-3681e5b10f94 · outbound

This paper cites Offline RL for natural language generation with implicit language q learning.

ShiQ: Bringing back Bellman to LLMs Offline RL for natural language generation with implicit language q learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.051009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.601286Z digest=sha256:afbbce5c5ce22814f7142c46f92a0fc10c7f376f256f415e35de70c46b5c9a18

Observation 3b0adb19-4dbd-4282-aa01-6293c0b7be72 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

ShiQ: Bringing back Bellman to LLMs Generalized preference optimization: A unified approach to offline alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.037824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.604883Z digest=sha256:15d215e9ea8e9a04154b5ed8a11c28bc8fbe378b8a56df5ff5e6ac257be244f6

Observation d4e4e0e5-f2b5-4ddd-bbf6-eee1db7c913d · outbound

This paper cites RL-finetuning LLMs from on- and off-policy data with a single algorithm.

ShiQ: Bringing back Bellman to LLMs RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.608629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.608629Z digest=sha256:1d22297d6e1703f4cad8bdeecc4038c65efac2e7e3b064469c952f4481ecb3b5

Observation fb2927fd-5aa3-4949-8f0f-b7a9482c7e55 · outbound

This paper cites Leverage the average: an analysis of kl regularization in reinforcement learning.Advances in Neural Information Processing Systems, 33:12163–12174, 2020.

ShiQ: Bringing back Bellman to LLMs Leverage the average: an analysis of kl regularization in reinforcement learning.Advances in Neural Information Processing Systems, 33:12163–12174, 2020

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.026707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.612498Z digest=sha256:aa7f841d89f44cbfd5c0412b1f41353bba6a1769303811e8f0dac30db27e6bf0

Observation 077620c5-8f80-437f-aeaf-373d6e7e9b1e · outbound

This paper cites Munchausen reinforcement learning.Advances in Neural Information Processing Systems, 33:4235–4246, 2020.

ShiQ: Bringing back Bellman to LLMs Munchausen reinforcement learning.Advances in Neural Information Processing Systems, 33:4235–4246, 2020

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:58.015588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.615592Z digest=sha256:b50c195c1a83a78e4a8b7ce0c897f9b63c59106f936e6290b3a0b8358dfb3f0c

Observation 5d2d18e3-512f-4d79-8f00-ef19195b33f8 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

ShiQ: Bringing back Bellman to LLMs Dueling network architectures for deep reinforcement learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.619137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.619137Z digest=sha256:49ffb13d5c754321bea23c3e94bee7640611d3bf4f7769dfbe8b03d6359dd115

Observation a58f1446-8375-4031-87a3-1d57a67489d5 · outbound

This paper cites Potential-based shaping and q-value initialization are equivalent.Journal of Artificial Intelligence Research, 19:205–208, 2003.

ShiQ: Bringing back Bellman to LLMs Potential-based shaping and q-value initialization are equivalent.Journal of Artificial Intelligence Research, 19:205–208, 2003

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.997686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.622900Z digest=sha256:29f22fe880a3f4184ec295276504f21fb4d2afffe36f9b152a77d9b4483683dd

Observation 5b979ec3-bb05-4a3c-a3e8-793c27b5aa84 · outbound

This paper cites Function optimization using connectionist reinforcement learning algorithms.Connection Science, 3(3):241–268, 1991.

ShiQ: Bringing back Bellman to LLMs Function optimization using connectionist reinforcement learning algorithms.Connection Science, 3(3):241–268, 1991

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.987572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.626223Z digest=sha256:98166cc6c2f5b9f205fc37e3b2707f41ebdd6e4c3b6c46d12300643f9ffdc742

Observation a82447f1-a91f-42cb-aa73-3b7e5c467616 · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

ShiQ: Bringing back Bellman to LLMs Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.629840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.629840Z digest=sha256:7910f002c4f8fbbda14b7e3761352978331c7eb5009917e69be5953b64aa1248

Observation 0a0bb9dc-e5d6-43fd-a9a8-d6514018fe9f · outbound

This paper cites InThe Twelfth International Conference on Learning Representations, 2024.

ShiQ: Bringing back Bellman to LLMs InThe Twelfth International Conference on Learning Representations, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.976771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.633448Z digest=sha256:17359de0a8cae0474a14c805645d2b23dc04376cb38451d0db7217baf8013e7f

Observation 6372540e-0a55-4552-b6ba-acefa4250a03 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

ShiQ: Bringing back Bellman to LLMs Adam-mini: Use Fewer Learning Rates To Gain More

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.637034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.637034Z digest=sha256:12c375c270c1c0d03d27617780af9c3a04bf2cc126b5e8a30db0b8cf5976a500

Observation ebbe1492-80e1-46c9-90d7-5d7cc68d3708 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

ShiQ: Bringing back Bellman to LLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.640889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.640889Z digest=sha256:e212eee7840adb7724c30ce7daa2a7b44e0a7fe9b8d0c37eed4942c0704cb237

Observation e9c88e6f-9758-4e5a-ba60-8ab220d3031b · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

ShiQ: Bringing back Bellman to LLMs ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.644647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.644647Z digest=sha256:0f76516c2388550cdacd5091d04e9f757c97ca77161779ab7e9141a889cae33e

Observation 0cc62e57-a085-422d-a0f6-ddefbeed1118 · outbound

This paper cites PhD thesis, Carnegie Mellon University, 2010.

ShiQ: Bringing back Bellman to LLMs PhD thesis, Carnegie Mellon University, 2010

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.966526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.648893Z digest=sha256:78386799ff30bc4ee3901b94e6df7a8dedc935bdd36474422de7c8df4c611723

Observation 940fde9c-4619-400c-a29f-6a51a3203b6f · outbound

This paper cites an unresolved cited work.

ShiQ: Bringing back Bellman to LLMs Unresolved cited work

Reference 53

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:10:57.954395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.652994Z digest=sha256:0edfef09c2addad8cca345deac15d216f6d4188e3095fe6c2aec12c45c4eab5b

Observation ebcac538-19f2-45bf-bd52-700034ba3dd1 · outbound

This paper cites Initialization trick.

ShiQ: Bringing back Bellman to LLMs Initialization trick

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.942715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.656959Z digest=sha256:e7cdf11a37713fda0a049cdb9214717ab46ac9ac63cc4357523013809eb67e83

Observation 683665a4-f333-4b52-a187-eda081d5cf04 · outbound

This paper cites This contrasts with other RL-finetuning approaches, such as DRO [ 34] or CoPG [ 12], that involve a square term per sequence of the batch.

ShiQ: Bringing back Bellman to LLMs This contrasts with other RL-finetuning approaches, such as DRO [ 34] or CoPG [ 12], that involve a square term per sequence of the batch

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.932010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.660461Z digest=sha256:7e21f30a3ad08bbc1ebfb3979a9f862386e0730762074f91efd8c5cb1d468d11

Observation 638e2006-da73-4454-b1aa-9c417e26ecd1 · outbound

This paper cites These tools use machine learning algorithms to generate realistic voices and faces.

ShiQ: Bringing back Bellman to LLMs These tools use machine learning algorithms to generate realistic voices and faces

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.921229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.665274Z digest=sha256:3dcdd501822ca8da6a5b8c58ae93d643d2c1f0f8ac4a4f6ebc8d78b60c92bfed

Observation 108adfb6-2a4c-4e93-9832-893323e94382 · outbound

This paper cites The more reference audio you have, the better the result.

ShiQ: Bringing back Bellman to LLMs The more reference audio you have, the better the result

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.910171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.669412Z digest=sha256:67782eecdda5c82304c1bb8140c8c8919e10a28976aacc872303861ac1795fe6

Observation 21c39961-af15-46bf-8967-c76c9f35c729 · outbound

This paper cites This process may take some time, depending on the complexity of the task and the 35.

ShiQ: Bringing back Bellman to LLMs This process may take some time, depending on the complexity of the task and the 35

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:57.897515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:10:57.673384Z digest=sha256:17fa7423d01c102f918a6fbadb186f9ae81fdef81ef53097976544ab08ceadd4

Pith citing papers

Observation 93209482-94e8-4983-a010-6f5035bca00a · inbound

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction cites this paper.

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction ShiQ: Bringing back Bellman to LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:33.501043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-16T21:28:20.658539Z digest=sha256:1e5e729b02f830aae27091fd9490d9b4b78ea6df4ed180aff683fa5cbd333f6f

Observation 97d5248d-0f2e-4647-ad0c-b1edf88dc07f · inbound

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction cites this paper.

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction ShiQ: Bringing back Bellman to LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T15:55:45.589124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T15:55:45.589124Z digest=sha256:330493222c4a2b1514141e93ea9cc3948866b790b05d9119b391a26d26382c30

Observation b304f723-78b5-4404-8aac-290f8627d5f8 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ShiQ: Bringing back Bellman to LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.636098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:c8a39133da84a7d78925875066c3818c2d8bc0af8d8df26d8a09e04e336dbbf9