Pith. sign in

Paper Citation Record · LEDGER

Natural Language Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 7 inbound Pith citation observations for arXiv:2411.14251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14251 v3

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:26:45.325950Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:27:23.283216Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:16:44.965693Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a368fc41-0a9e-44b2-a09d-f5b4429e93fc · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

Natural Language Reinforcement Learning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.935259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.935259Z digest=sha256:88b61e7927f12ce10eb884671287f0c29e599aa8b67d9c8cec23637a647646be

Observation edbff655-b780-4f2c-acc6-4781cb896451 · outbound

This paper cites PaLM 2 Technical Report.

Natural Language Reinforcement Learning PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.940575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.940575Z digest=sha256:a481cb6b17000839b20034a4c85b8b6e107a42750b0cadcbf3b9231a47f0dfda

Observation 8e6354cf-3945-4821-9e95-e221f3b5d51b · outbound

This paper cites Barreto, W.

Natural Language Reinforcement Learning Barreto, W

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.945172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.945172Z digest=sha256:f8f0fccae0d24a2a259d02a154c38b5692bb2f9de715d18fda104a1ef2d6770d

Observation cd3505fa-6d02-425b-b25f-5ec9ffe43bc7 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.949269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.949269Z digest=sha256:16b7668ccd8521d5651a0e34732651827e6fdbb24df1ad426adbb0196c48a9ad

Observation 82eedd35-baa7-4241-ae48-ae7c2f7ffc24 · outbound

This paper cites Bellman, R.

Natural Language Reinforcement Learning Bellman, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.953506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.953506Z digest=sha256:8e8d015945ccab238a20d561bf82781402c97c54379ded87aff7b64399d88225

Observation 027d1516-beae-4a97-b028-10aad4e27b5d · outbound

This paper cites Brown, B.

Natural Language Reinforcement Learning Brown, B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.957770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.957770Z digest=sha256:b059f50c34a3901317bb6862509e3469d9a242424c478dd103ab01a2e501ddf8

Observation b898905b-6b13-42b8-b6bb-57be078bb2a6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Natural Language Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.962409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.962409Z digest=sha256:4c730fc2cec5f1330227f8f21f6122aa5f19a81be30ed38daae3438f93e85faa

Observation 968957f6-3aeb-4a60-91c8-f50ddcfe376b · outbound

This paper cites LLF-Bench: Benchmark for Interactive Learning from Language Feedback.

Natural Language Reinforcement Learning LLF-Bench: Benchmark for Interactive Learning from Language Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.966685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.966685Z digest=sha256:4994baf2d30b86ae4d3234a37ac3fe34bc9d28d1ca63000df266f26e8e834ae3

Observation 215f34e5-cb78-4abd-bae6-8d664bfbc452 · outbound

This paper cites Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs.

Natural Language Reinforcement Learning Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.971047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.971047Z digest=sha256:4184437b05b041a2b2eee489c9fa656a61ff707e63a209890c2cc4f44a1bda46

Observation e0d72d1f-5f2b-4c08-9ed6-c2e6c6c1f16b · outbound

This paper cites Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning.

Natural Language Reinforcement Learning Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.975500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.975500Z digest=sha256:698cb63391bdc8885519b97ac1fd2473871d8aef64a9c99fc727603f17608777

Observation e30ebab1-2d78-405c-95ac-2df10dac9828 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Natural Language Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.979752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.979752Z digest=sha256:b53f179e52eab7f998a09630f0c0eac199321b8a2891a2cfcf2a66ddfa6c976f

Observation 6eec9846-cdb4-4846-97f4-fa977e383a24 · outbound

This paper cites State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User Understanding.

Natural Language Reinforcement Learning State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User Understanding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:26:45.777439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:44.983956Z digest=sha256:fde43d04be7c20decb25d98dc341b086f57bd4f354bc22cc599b2015a96617fa

Observation 68263461-c87b-4c98-a219-73e774b388ce · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.988684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.988684Z digest=sha256:c1f49abe9cee1f6b6d85bb315e879371b3013d44daa3cebc53ba75ba19a4061d

Observation 90799239-f6f4-44ef-aa2b-5760d2e0ee17 · outbound

This paper cites The Llama 3 Herd of Models.

Natural Language Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.992608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.992608Z digest=sha256:43fdd03ab7f69e4f0a4ade482154afa702edf2dda954ec5d08d5dd36033ac550

Observation 340b0613-f11d-4dde-bd22-1ef2c30e3f63 · outbound

This paper cites ChessGPT: Bridging Policy Learning and Language Modeling.

Natural Language Reinforcement Learning ChessGPT: Bridging Policy Learning and Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:44.996629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:44.996629Z digest=sha256:563559791c4d1b8970583b9176873b8833bed10c9e21bf9718ea6ebea3cae618

Observation cf78593d-89e2-49b2-a804-19c8160fc570 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Natural Language Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.000653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.000653Z digest=sha256:01203d0cb7a014af1df2c1fdc4fed27dd264d9645c6267eaf773ecf71e363866

Observation 022ddc83-bea1-4f4b-98cf-1d39cb98a159 · outbound

This paper cites LLM-based NLG Evaluation: Current Status and Challenges.

Natural Language Reinforcement Learning LLM-based NLG Evaluation: Current Status and Challenges

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.005039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.005039Z digest=sha256:b7e5ddc033b82caa62394b42bcff9ebc2498f0abbf87510cfc9f8f57bfdff45c

Observation ed42830a-4390-451e-b83b-c2c48bb67aa9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Natural Language Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.009187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.009187Z digest=sha256:b0b12c3db3b7ac440de1d60f41c46e62794479e0e6326f8c701b2defd5348667

Observation 21297ea2-61f3-4a41-92c8-9f173e7a7b85 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Natural Language Reinforcement Learning Reasoning with Language Model is Planning with World Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.013350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.013350Z digest=sha256:eba6c5d3d686a1515620f9f477d843a35c9ac3492f40f233aa8a894c34fe4378

Observation c8c35a29-850e-4ef6-a3ba-6711af4651e6 · outbound

This paper cites Hayes and J.

Natural Language Reinforcement Learning Hayes and J

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.017508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.017508Z digest=sha256:622ddcca0dcf0c32d2bf8129554813210951a822dc0818ac62e28b9b20aede30

Observation f79122b0-303a-4c7a-adfa-0994a33a1e7f · outbound

This paper cites OpenAI o1 System Card.

Natural Language Reinforcement Learning OpenAI o1 System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.021348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.021348Z digest=sha256:f6b2575ff878034ca2b927242c246589505b1ddcb4f201e75d12b4e88a1397f8

Observation 5623c0bb-efbc-4590-9539-14819357dfe2 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.025502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.025502Z digest=sha256:b5d0bffb694f00b4644dee756a38b39a0850b545d3f6e1dee3dd62679fd5ba46

Observation e39943cb-360e-4690-9ee5-1b85c31dd478 · outbound

This paper cites TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks.

Natural Language Reinforcement Learning TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.029330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.029330Z digest=sha256:fc6c46ad4333c4c7073ac1e01bb65f202d3ed00b2a8746451822aaf4dcc1876d

Observation 50e91fdc-b438-433e-9c09-a0b23aa0b2c3 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Natural Language Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.033453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.033453Z digest=sha256:5ea33745c7d1feeba52c6a1044041cda84933769b48aee34190571ee9b162d5d

Observation a2a2e805-cc57-485e-8a56-0d3276fd2a5f · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.037941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.037941Z digest=sha256:477ffb2e62e2643e4a3c843c9bb91ac82dc1f44bcb28f950173a5cec43fb7254

Observation 090d6e21-77de-4093-9ceb-52a911f540c7 · outbound

This paper cites Kocsis and C.

Natural Language Reinforcement Learning Kocsis and C

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.041894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.041894Z digest=sha256:32775676b68557d3c73b1fe3f60781d48b82199bef6ec38ae1d25606e43a7f77

Observation baea8e53-557f-4d0a-a45d-660c4bede6b9 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.370574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.045777Z digest=sha256:8ef10e6e972970c529ccddab0409bb66e0a3103a76d9fe64225257093487d132

Observation 15555ad6-9319-414c-af79-3b621ccafb93 · outbound

This paper cites OpenSpiel: A Framework for Reinforcement Learning in Games.

Natural Language Reinforcement Learning OpenSpiel: A Framework for Reinforcement Learning in Games

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.049600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.049600Z digest=sha256:2280401a5d90e43cc9deffaa2d1ecae46b0a9b1d897c7273818623497510831a

Observation 0ef26c8b-17ad-43a0-a63c-bc4178e0e717 · outbound

This paper cites Generative Judge for Evaluating Alignment.

Natural Language Reinforcement Learning Generative Judge for Evaluating Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.053988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.053988Z digest=sha256:bc252312cb644ebc5182d7df3beb20bea1a0d48bccc6eae7e2ba82040e43279a

Observation 32b09d38-f9ab-467a-b377-a323a2b11b45 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Natural Language Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.058370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.058370Z digest=sha256:20eae98f9f04a73c96c3722a173d3fe689b87573839e89df544c3a686207ebd2

Observation 3c3292ec-4609-4c37-a898-c828cc6d0528 · outbound

This paper cites Generative Reward Models.

Natural Language Reinforcement Learning Generative Reward Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.062487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.062487Z digest=sha256:fa94fc110a46c52e41d93e4fd6d71cb0afdda54a955f48f467cfb2401606420e

Observation 5c7a0cc6-029a-497c-99b0-9b0581e26e10 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.066454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.066454Z digest=sha256:63bcefda1f184e4cf3bef12798ca21755f55aa7ed703138f96d07fc2857ead5c

Observation 3edd602c-77a4-4631-892a-97efadcdf350 · outbound

This paper cites GPT-4 Technical Report.

Natural Language Reinforcement Learning GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.070775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.070775Z digest=sha256:d44816bba84e5ed4a08971d2c5e68456398e5e112f09d31292ff196364661e65

Observation 74c14bfd-11a0-404a-a664-d71479fa398f · outbound

This paper cites Ouyang, J.

Natural Language Reinforcement Learning Ouyang, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.074842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.074842Z digest=sha256:584690c0d3463dfc95a6623fe19ce4d04fbadf31896109d91536a97dd975f217

Observation 46cb65de-18f8-4efb-a797-9dad6c1706b0 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Natural Language Reinforcement Learning Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.078805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.078805Z digest=sha256:97d892f6054d1288dce2e1c71e29449915d5006cc61080effeb87c4b5eed746a

Observation fd94541d-a95b-4d81-ac3a-83ac8faae2d8 · outbound

This paper cites Saffidine, N.

Natural Language Reinforcement Learning Saffidine, N

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.340851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.083077Z digest=sha256:722676e67fd2c5078d86e33943b1b08f5790b41cc1d22ec85bb983224d2ce0f3

Observation efdae428-298b-40be-a3cc-7eb771a76251 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Natural Language Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.087007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.087007Z digest=sha256:c50eb1d0726029f1f846f79976f54653010673a11d490d3d520b114c9d488c6f

Observation 26cac532-b3c5-415e-9b01-039864b7dbe5 · outbound

This paper cites Rethinking Reflection in Pre-Training.

Natural Language Reinforcement Learning Rethinking Reflection in Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.091249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.091249Z digest=sha256:7dfc255c58374305cdb3113bf57dc99bcdea33fadfb9a904117055fbef7b9177

Observation 2a937d47-c854-4991-98d6-34686234fe20 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Natural Language Reinforcement Learning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.095327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.095327Z digest=sha256:3f9077577cdb13f9ccee9b4e4ddf25ad21efedc841def08ce2914d6360f0963a

Observation c71c5c5e-437b-45dd-ae5a-f7301efda2d9 · outbound

This paper cites Silver and R.

Natural Language Reinforcement Learning Silver and R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.327548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.099656Z digest=sha256:a420538e8e07968f5cfd716b50bb92555250bb0c57e3a736b4c1fda6feb770e1

Observation b181ebfa-afc3-498d-8ae2-8c55cfc9d7a7 · outbound

This paper cites Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations.

Natural Language Reinforcement Learning Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.103803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.103803Z digest=sha256:53fbf5788c903d85ca2eea682549732655ec9e0cbfd6e37a86e82117a3519997

Observation 40e41c35-8a1f-44ab-b114-adeb6c0703c4 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.313417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.108413Z digest=sha256:cab324f07730d69fe0ebaad75380c510c724b0757963c416022a7ea01260266a

Observation 516ac680-a738-4587-8af6-7fa5091b44e9 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.112377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.112377Z digest=sha256:4dabe13826344e62cb425550bc35beed7b29ab63a192f8de5103cd800394e876

Observation 433e3daf-fa10-40c6-969b-efde65600741 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.291282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.116293Z digest=sha256:27a24b2a053a08b007fbc6bfd68c39dde24bfdbba3b475b79db372ac472574e1

Observation fd17f4e5-b3c1-44ec-9148-8e5ad73b8e39 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Natural Language Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.120438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.120438Z digest=sha256:1eb4c79519f9753c95211028fcb2f4cc387205bada0f23cc86e52b62acd93808

Observation b80ef602-fd7d-4a90-bcaa-9466c20ee407 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Natural Language Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.124562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.124562Z digest=sha256:ae8d5a9c1ab5918c75b9962624e4530f5903862312f8569a70a2893ccc046770

Observation 9fe5d0bf-c40b-462d-a63d-2e5ab29e633f · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Natural Language Reinforcement Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.128901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.128901Z digest=sha256:27d120423464e1b4aeedda3503154307ae306cc735b8c9076bb8fdd215b16a89

Observation fc44de71-934f-4cc1-a34b-022d6a17a79e · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Natural Language Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.133299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.133299Z digest=sha256:06a49f21320ae86229128e791a06558ce578ebb088707a46cb01708793a3dcd1

Observation c163c745-5944-48c7-a55b-e3608bae8c10 · outbound

This paper cites Emergent Abilities of Large Language Models.

Natural Language Reinforcement Learning Emergent Abilities of Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.137667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.137667Z digest=sha256:3ec46b44b3a8bb3cdc347d282501ab2497d837d689ac83bde9296e85ea3fac0d

Observation 9604b669-5dc9-4309-8af6-c501d77ac55e · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.141791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.141791Z digest=sha256:b61bd1607970ba1783e086959f43da4ba46ab5b25964d92e6db7d8489b8309d3

Observation 09b81976-abb5-40cd-9058-8b833ba79809 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.145606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.145606Z digest=sha256:06bde4a36246eaa4310f3e7542e3b7a6399d78fb59d592644b6a591270adf7d2

Observation a9f26d6e-8fda-4938-9a86-b1d7e0d39b94 · outbound

This paper cites Large Language Models for Generative Information Extraction: A Survey.

Natural Language Reinforcement Learning Large Language Models for Generative Information Extraction: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.149473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.149473Z digest=sha256:f24172c7461c12f41fedde508ec52c58ffbc77d8d96da005ae3a311fb844826c

Observation de3cae2d-6219-4bcf-8e05-f411311920cf · outbound

This paper cites Large Language Models as Optimizers.

Natural Language Reinforcement Learning Large Language Models as Optimizers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.153633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.153633Z digest=sha256:77080dbc8ad36a2f89059647e6c57b85050837f9c8896063112133c0d0c4fd51

Observation 69d78044-50c8-4cac-b9c0-aa02ce107df7 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Natural Language Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.157743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.157743Z digest=sha256:147a4bd41c7868e3cca6869ef343d0211be62725e4256961a7d3fb7d841bede8

Observation 275d478a-7246-49b5-a490-64807a15e1a4 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Natural Language Reinforcement Learning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.161908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.161908Z digest=sha256:a9bfbe6f7597a35dfd83788511c8b88dea1c95b9015983d05a7e8d70863819b0

Observation c955fc7f-bdb2-4801-bee7-dd28cff0eafc · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Natural Language Reinforcement Learning TextGrad: Automatic "Differentiation" via Text

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.166125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.166125Z digest=sha256:a99878ff5ce14c68f98e89c3721412396a5a5c594dc1d72cfbf5908e39aa7047

Observation 1385d4df-1b9d-4dd4-9239-1ba1dda1cbca · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Natural Language Reinforcement Learning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.170891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.170891Z digest=sha256:11854ea45e842169870b3a0136f4f3ba43dc0167e011a451ef697d928db31ed8

Observation 83a6326b-e7e6-42c3-ab74-7a1d3a5f38bc · outbound

This paper cites Benchmarking Large Language Models for News Summarization.

Natural Language Reinforcement Learning Benchmarking Large Language Models for News Summarization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.174979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.174979Z digest=sha256:73fc7aa20f7a71bfe2144335bf0607467e0dd2cf6b9c8719f40287cd7ecf2770

Observation 3161a2f7-e05a-433b-b13d-28948673740d · outbound

This paper cites Policy Improvement using Language Feedback Models.

Natural Language Reinforcement Learning Policy Improvement using Language Feedback Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.179105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.179105Z digest=sha256:bbcd4be836902ed067d8bea11e04ad5ee046cfe7ce49f185a4d82aaff6f655fc

Observation 315e4d54-7606-40cf-9c90-2262fe82b012 · outbound

This paper cites Use <white> or <black> to represent the winning side.

Natural Language Reinforcement Learning Use <white> or <black> to represent the winning side

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.223330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.195886Z digest=sha256:f0c3f0ad38938094199af2c0f0dedb5b6eb8d092b065ece7218a898ce1e9b227

Observation 73decc35-f457-4762-b451-80870bd6cc29 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.210511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.199956Z digest=sha256:aff4373dc7ac14d2738b5014f5bf7792c59716f7e19172c3736f819ba4742e44

Observation 77302053-7a92-47b6-8402-ea19c56002fb · outbound

This paper cites **Advantage:** <white> Overall, White has a slight advantage in this position, with multiple ways to break through Black’s lines and gain a significant advantage.

Natural Language Reinforcement Learning **Advantage:** <white> Overall, White has a slight advantage in this position, with multiple ways to break through Black’s lines and gain a significant advantage

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.197407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.204249Z digest=sha256:d67bc5badddbef60844f0edffea57f0ce68d9111fe28fc4e83bb33d85f1f82ac

Observation 6742d035-c1cd-4feb-a4c2-ed68866ed586 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.184618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.208754Z digest=sha256:8210a52df723c0987632dc7a1a733794fdef012f5e52454cf626e1df716d05b1

Observation 49b80d2d-ecb4-4a07-bb18-ccd3349b0257 · outbound

This paper cites **Advantage:** <white> Overall, White has a slight advantage in this position, with multiple ways to break through Black’s lines and gain a significant advantage.

Natural Language Reinforcement Learning **Advantage:** <white> Overall, White has a slight advantage in this position, with multiple ways to break through Black’s lines and gain a significant advantage

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.172127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.212686Z digest=sha256:bba346dc35d32d0593aea688c8c86c8d1d99e54ddc77860c573db2057a36ba7a

Observation c49dc1fc-d99e-4de9-988b-aa5338f9d038 · outbound

This paper cites This move has the goal of preparing to defend and possibly create a barrier against Black’s pieces.

Natural Language Reinforcement Learning This move has the goal of preparing to defend and possibly create a barrier against Black’s pieces

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.159468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.216831Z digest=sha256:d0cee17b99cdacd2c849eab352119174f35a4374412bdc3ad3dbd1291b17f2ba

Observation 0cedb6ca-ddaa-4013-9e28-ef5290690248 · outbound

This paper cites This move aims to create more space and put pressure on Black’s pieces, which will make it harder for them to maneuver.

Natural Language Reinforcement Learning This move aims to create more space and put pressure on Black’s pieces, which will make it harder for them to maneuver

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.146899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.220690Z digest=sha256:3644d89d7cb9ef04c41c55289667527b3e67809c564fa3b324e913c5b1d9ba70

Observation cb177108-f87d-4e0d-91b4-08ae93c7ec17 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.134238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.224565Z digest=sha256:fdefc1ce077fbbda1ec86c481901479bf423202fd272a3cb28aaffec0d45f887

Observation f23cfb63-092e-479d-a9ec-f4bde217cc0b · outbound

This paper cites final_evaluation.

Natural Language Reinforcement Learning final_evaluation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.121483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.228376Z digest=sha256:1792a748a6eb54973ee5853919a08bcedaf1ea432da895e4a96607b5da0f0f7b

Observation f5bbee96-1118-454a-b70d-e1b3cfc39820 · outbound

This paper cites "" EVAL_USER_PROMPT =.

Natural Language Reinforcement Learning "" EVAL_USER_PROMPT =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.108186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.244623Z digest=sha256:d139624943cc7e6a96edc3a79923e6382760329dd8bf1af389e74a37769e4d62

Observation fcba6d14-aea2-49bb-bd2f-1c12f89e5eb0 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.262116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.248527Z digest=sha256:f1f4186b640f57a38689631ce2f896699e6791319780b6fed3f29093bed87a7e

Observation 6922645e-48b9-4b95-9df2-3ef9a5369764 · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.249725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.253113Z digest=sha256:1a9716f76d393da8bed84f0023b72674b4e685253329c99aec006e7dae105796

Observation 11b82be6-f8ee-4836-b06d-18703db65bca · outbound

This paper cites an unresolved cited work.

Natural Language Reinforcement Learning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:26:46.236565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.257173Z digest=sha256:6f9156b93fee4a65fbc77f8eeb9b2518644f16a670e6716ade3617573b19a2ad

Observation 08198664-f4ef-4fab-83a8-6cc9d48fb95c · outbound

This paper cites "" TD_USER_PROMPT =.

Natural Language Reinforcement Learning "" TD_USER_PROMPT =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.095362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.261418Z digest=sha256:5fe73601ef1d069152e612380b229a49cac5555fe63edac1cb43a39dfb11e0dc

Observation 768c5c5a-a3a9-4aa7-b0a7-aa2b90aa875e · outbound

This paper cites The conclusion supports the paper’s contributions and scope.

Natural Language Reinforcement Learning The conclusion supports the paper’s contributions and scope

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.082659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.265874Z digest=sha256:873f0a306c16e20d6c95fcbd18597135fb79e7af71c955e7dc867cf982acf791

Observation d9b6f142-05d6-4266-9eee-cf085f843313 · outbound

This paper cites Limitations.

Natural Language Reinforcement Learning Limitations

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.069435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.270329Z digest=sha256:d95e14d9fbf97f7091278fa18d2e5c385d34f7819d68189206b4e2a4d74f0159

Observation 9a5f5e7b-22c5-4a7e-b387-a17dc885f5bd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.057280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.274235Z digest=sha256:4d4dea683065fec4bbf3b5f0f4fd154025f3189fb8e4c46e42a42a48d78f7a2e

Observation 800b0abe-8c36-4fa3-a400-a44c47e56794 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not include experiments

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.044107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.278113Z digest=sha256:c5dce28b800feccc502e79e5a4ed6533371d5ea1e5dade472102b39788b339c8

Observation 96587e05-3452-41c3-95d1-28ad959637d0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.031338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.281992Z digest=sha256:28c21c25c4ba9be9045a4e36391f3d9cbb703c26090961e11ed91646c121cfea

Observation 49a75a41-5a31-4439-bd1f-82f37f7bcd4c · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not include experiments

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.018065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.285813Z digest=sha256:b37e5ad9602d50e0882a763a6bbada3b8979eb43054dc0a1e04d9122566dfda1

Observation 01f27543-9ee9-4199-9b2c-bd3a2ee81389 · outbound

This paper cites Due to high computational burdens, the paper does not include error bars for all of our experiments, but the paper contains every detail necessary for experiment reproduction.

Natural Language Reinforcement Learning Due to high computational burdens, the paper does not include error bars for all of our experiments, but the paper contains every detail necessary for experiment reproduction

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:46.004196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.289778Z digest=sha256:59e1d82ff21409b22b2bda452ba5c8bc3481e020cd5813673d7f20aee72d248b

Observation f61acb50-017a-4a4d-8e66-fda4bd71858f · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not include experiments

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.990850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.293671Z digest=sha256:6075fb18e5c0aab54d711ed62c29169bd67526c5c78416f5d7eb7d53c5a4fd34

Observation d4448829-c2b1-488c-b45d-4f4ec9d08db4 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.976256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.297528Z digest=sha256:7b5827f11c4487e51015a8de85fdae0119f3de42d516d63fb132a9d772f2e817

Observation 334c8107-35b1-4106-9487-7a823278ab2c · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.962991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.301643Z digest=sha256:a3e5a35a85e9dddd2c7df59c606d33d35e9168940743c3e6946125b0b517f424

Observation 53c71d70-a38e-40d7-b635-d8b8cb40d1ba · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper poses no such risks

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.949541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.305587Z digest=sha256:39fea018fc2b8ae061c537de4cefdbff4651036424576ebb3d6e251d5569e685

Observation a5cc73ca-7410-4d9b-826a-cff21d26a8a3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not use existing assets

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.935765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.309592Z digest=sha256:97adbb23f2277dffacda577d9d3c7c8aa79a0a8b0d4e1c8159408b23dd90571b

Observation 86107130-3554-4d29-b039-416562879423 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not release new assets

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.922556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.313504Z digest=sha256:ee8cfd0383c40c59dee5da79d3aaf79fdd3ee50f5ad9f566d6e2ab1f05ba1fac

Observation 3db7b08e-9894-421d-8a8f-8004f128d4d9 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.317297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.317297Z digest=sha256:42494e56090213eafc5ffe7277b46d321feb8e509cfa3e1cca2faba6f56d1d52

Observation e202afd8-c6e8-4284-b00f-ca07ba2aae17 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Natural Language Reinforcement Learning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.900941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.321149Z digest=sha256:5d6b8fc97f26c9db93242af62a8992751ca9f17ae73602c24907a541f56002da

Observation 775af9aa-6b5e-4ce4-95e1-46dec77421c2 · outbound

This paper cites Answer: [Yes] Justification: The paper is about foundational research on LLM and has described the usage of LLMs.

Natural Language Reinforcement Learning Answer: [Yes] Justification: The paper is about foundational research on LLM and has described the usage of LLMs

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:26:45.887647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:26:45.325950Z digest=sha256:78b1052a3d01319f37bdf8b749fb103efd8fa513f60306c3fb10a952055a71f6

Pith citing papers

Observation 4a8ea224-44d9-4c42-bbe4-61d127990c06 · inbound

A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks cites this paper.

A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks Natural Language Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:23.283216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:23.283216Z digest=sha256:885e409b99c45e884209d417d76a77ac5ed005261f9f756ebd9a57d916d56765

Observation 1d0bc4b8-304c-47b1-b80c-88712380399e · inbound

elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings cites this paper.

elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings Natural Language Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:17:51.412053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:17:51.412053Z digest=sha256:ec3a9da4c44c843508648438bef73e7de7211fea9eb3a9388eb2c8430f11cd71

Observation 6c032603-e570-4a06-8d29-73791bff4a0d · inbound

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra cites this paper.

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra Natural Language Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:42.368027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:42.368027Z digest=sha256:eb0659c1414dd8cc553fd7ec807f85dbb7147b1eff7f28615821dd280fccef79

Observation a793711b-2543-4159-9e11-a314f0d7daa6 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Natural Language Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.368600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:4affaa1173500829a15254057de94a150b74e24d8b0b90d4f1f026077d001850

Observation 62af02a2-3793-4ca9-9930-1bf60acf34a8 · inbound

When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling cites this paper.

When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling Natural Language Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.967417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:01:20.272455Z digest=sha256:6a00bfd3a238f496c4797dd6b6d20461bc11eebb0d5ad743d47cb393ad4e4cae

Observation 0a90098d-cd39-4f92-a217-89194b133d30 · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Natural Language Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.260514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.260514Z digest=sha256:2df641449b0d891f744b174ee6a74c3a87ac603aecd524c91a807cc72784d73e

Observation 9be8e652-2fb8-4dc0-ad9b-9062167e5bbf · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Natural Language Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:15.885559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:15.885559Z digest=sha256:9331314a660cacf90f285c974b8ccb8eb74712338cb5e3c86aaf78ec475c4b86