Pith. sign in

Paper Citation Record · LEDGER

Offline RL for Natural Language Generation with Implicit Language Q Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2206.11871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.11871 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:31:18.563950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.700056Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb309334-bc7c-46be-b08e-4600205c2f23 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 176

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:04:10.772296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:d9b7829eb95326a5348c05a1ead619ed3eb2c85746119076c77bbc05f4cb85f8

Observation 6671bc15-ff5f-4021-9b52-8f235c747b07 · inbound

Towards Cost-Effective Reward Guided Text Generation cites this paper.

Towards Cost-Effective Reward Guided Text Generation Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.563950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.563950Z digest=sha256:fae1fddf210af695f41e207d5be0d498b42c5b6a49876ce6b2cf1ec3b7ec6db6

Observation 1eda3411-9bf8-4505-aa04-a1c7d6d2ef62 · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.509615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.509615Z digest=sha256:d60c4e8c59121a5d2120d23c09650d6385c6842569b51f4bd2dd42e71ed0e5b8

Observation 85e3a1ec-6c4e-408e-ad20-e2303f14f675 · inbound

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs cites this paper.

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:52.155945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:16:51.288077Z digest=sha256:afb8ebd5e624505ebe17a7d1cbd6efcc04fc45e325bf82d5b994553d0f66085a

Observation 39dfec57-eb54-4149-9c3a-0761c4223173 · inbound

The Fourier Spectral Transformer Networks For Efficient and Generalizable Nonlinear PDEs Prediction cites this paper.

The Fourier Spectral Transformer Networks For Efficient and Generalizable Nonlinear PDEs Prediction Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:27:46.666843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:27:46.666843Z digest=sha256:b1c5ccfecd767ab5e7d9143961cbf170c37e0a2196bc9f902e4404106db83e84

Observation bd9f13a9-47f1-4a00-b732-6889e6f75555 · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.365511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.365511Z digest=sha256:67cb877d7051c852202fef795c38502ba240d43fb105d7c0be0b0a4ad3833285

Observation ab96ca14-e472-445b-9314-59e560c0698e · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.866671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:a1a824ed62e8d5f5c5bdac114f5a2bfde34414b3f043117816ee647bd4e6ae09

Observation c16dbc95-2234-4efb-9b6f-eb9bf839e206 · inbound

Conditional Attribute Estimation with Autoregressive Sequence Models cites this paper.

Conditional Attribute Estimation with Autoregressive Sequence Models Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:50:04.428105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:49:49.475183Z digest=sha256:8b22a50ebe18b978d42cc2ba801991ad7f5d626feb5c3dd2d886edda61b01d20

Observation 62ee2879-334b-4228-a994-d53873996787 · inbound

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning cites this paper.

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.438948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T11:54:07.211895Z digest=sha256:522f837b84297504b54242ab43c052ba6247a7dba9d96a47207532f2b94bc4a7

Observation 4f8fdb1b-ea3c-4569-860f-962af4b104c4 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.701597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:a82c0915bcc896074e2da796e3e4e64c41c3e40330cea9d086eeeff26837f786

Observation cd313023-93eb-4b29-8827-8a00b2a9bb0f · inbound

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents cites this paper.

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:29.353152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T00:48:01.918482Z digest=sha256:e4131efbff8ebbaebc562270a1639fa00167c76dd9a2aef83e5791e4d8f52356

Observation 07eaa869-fd18-4474-94a8-7c26c86b8527 · inbound

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents cites this paper.

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T08:43:11.507334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T08:43:11.507334Z digest=sha256:e6e46ee8cbe85c1c6f17b3202b7b2c7b4993edf1df6352239c8c76e4af39ff8b

Observation 4d815492-c4a5-4744-8f5d-7e26e2dc2823 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:40.732858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:40.732858Z digest=sha256:e2fdd3c13b8b2366fe0c5a5044b045638cf145207ce5819b40da5ea97a5151e0