Pith. sign in

Paper Citation Record · LEDGER

Reinforced Language Models for Sequential Decision Making

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2508.10839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10839 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:20:56.723685Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:44:28.552513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:47:27.135925Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc974d6f-2720-497d-8179-7b6803162c93 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reinforced Language Models for Sequential Decision Making , " * write output.state after.block = add.period write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:00.378253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:53.559913Z digest=sha256:452d12a14b77afd4172265d63f4a4e574621db39d91c7c2bc189b4d6da1933bf

Observation 6786b4fd-cb1d-4dfe-8e02-4f79bbaabca2 · outbound

This paper cites write newline.

Reinforced Language Models for Sequential Decision Making write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:53.598487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:53.598487Z digest=sha256:aec831fe6068ab46afd25fa52ea489a39bdd3bc9b715ed10110242eed439a38d

Observation 9f7e173b-2edd-4965-82cb-0502505ecd5f · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:53.664054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:53.664054Z digest=sha256:04097f4fe5fc0bff165ff5e81338ad2d1d7c9b461e138a94edaae3f5f60fbe5e

Observation a85edd25-c227-4cdc-a376-494d2f734fb4 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:00.200912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:53.739760Z digest=sha256:71a1bcbac8a8103bf2b14b2ed26ab76f6a1fb69998c0ff38e4726fcb4d9027e1

Observation fe43891f-efe5-4332-88fd-ba8d3b1a22d2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforced Language Models for Sequential Decision Making DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:53.799169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:53.799169Z digest=sha256:bc790a21bb5f6a0d26371921eb1983a39df75189529af95873c4a41d29be1c61

Observation 211b3b15-0822-47f1-a202-816c7a2cd241 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:00.069380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:53.864907Z digest=sha256:400827c5594fafbf945bc60e1438efb6debeb556a109d51f93138ea4cc467b14

Observation 6331ee09-c3a0-4477-8cb5-3373dce81642 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.923308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:53.936107Z digest=sha256:f2b3645bd8a8d8e5f30e932663c3f04041a960c32b95a30ed1b13a0d1327f7fe

Observation a35cd3c0-c1b6-4e4b-bdbb-680e44d431d2 · outbound

This paper cites A.; Bernstein, D.

Reinforced Language Models for Sequential Decision Making A.; Bernstein, D

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:20:59.819666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.029675Z digest=sha256:68efae4ae4026483fa4aaa57601895b0ff0510796f600fbee21ef6c68b0bc20c

Observation c16c6a00-d1fe-4f2f-a93f-369528e211ca · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Reinforced Language Models for Sequential Decision Making T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.083020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.083020Z digest=sha256:7586a330198d7b35fcb70892fafcd79d0e677a3f3df8f6ee1d1bba976a1faffd

Observation 8b7cefc6-3ed0-4b60-8b31-63ef294136ee · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.658999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.126520Z digest=sha256:9e4f73784fe49b46854e48d12b9b8cfe8d245fda0ef1c4a4a9038d72beef1ca1

Observation 24b0cb18-55f0-4650-b89c-637c4f8ce916 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.551565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.236148Z digest=sha256:7a69126b2e975a78928a07ed924278a93db7d4a2f942e11fe5951865df3bd77d

Observation 928fb816-b6c9-4941-9b71-324f10ca7d27 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.435228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.301771Z digest=sha256:35cb033a87f27a34475684c1e26bf4afe9b69f27926b1f887969290d216c88c4

Observation 8a8570bd-8a71-442c-b15f-4fe34a1b0797 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.288591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.361663Z digest=sha256:db69a5b7348ce98845917e9939d439dd32f17db8e2c05dc4f1b9c2106776c85d

Observation d9191b20-4cca-4422-824f-2d124ffdc844 · outbound

This paper cites HMCF: A Human-in-the-loop Multi-Robot Collaboration Framework Based on Large Language Models.

Reinforced Language Models for Sequential Decision Making HMCF: A Human-in-the-loop Multi-Robot Collaboration Framework Based on Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.486456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.486456Z digest=sha256:4b54b2a3f5daa9e855deee4f47e2cb727793b614a535f91acc131ad895a1ffeb

Observation a12a504d-2e2d-4c62-81a3-dd81429bcd31 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Reinforced Language Models for Sequential Decision Making Playing Atari with Deep Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.531721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.531721Z digest=sha256:1b48bdfc1446b29b6321db2b8971e667015e65d99567bc51bc5f9da8fbca20d4

Observation a2b58967-429c-4a92-99c2-94d717356c35 · outbound

This paper cites ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning.

Reinforced Language Models for Sequential Decision Making ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.604148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.604148Z digest=sha256:997cc6cced801c9e2967af4486ba3c9c38953fa70bd5f7b73d9ca120d0908a01

Observation e3cdc1d5-b022-4c21-801d-868b5f0f3c3a · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:59.108432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.684439Z digest=sha256:fc3416452359511fc63fa027ad142c6007e9f6148e5929ade071a39ce2c256ee

Observation 5d520b1a-4fe6-4431-a27a-8e6ab58e5c4c · outbound

This paper cites GPT-4o System Card.

Reinforced Language Models for Sequential Decision Making GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.757877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.757877Z digest=sha256:0a450c0ffe25e1c09c4d91104e8a58ee9793f1573d47fbaa5b7a5a67d79553ac

Observation 292b3ed9-5318-4fad-9e5d-8a22f21c988e · outbound

This paper cites L.; Mishkin, P.; Zhang, C.; et al.

Reinforced Language Models for Sequential Decision Making L.; Mishkin, P.; Zhang, C.; et al

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:20:58.945154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.857337Z digest=sha256:127d543bef684b5f099bbc13b56c15df8064a30be2f2d7dd8640089901f0ff6c

Observation 2086a6f9-418d-40f4-b5a4-18d3fbaad7c4 · outbound

This paper cites E.; Zhang, K.; and Kim, J.-K.

Reinforced Language Models for Sequential Decision Making E.; Zhang, K.; and Kim, J.-K

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:20:58.761059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:54.904478Z digest=sha256:6747fc41594fc9d69c8525b1b23ce5f71093bd682f59310f92b2606453830210

Observation 71ab2df3-6d4c-4faa-acd8-83a998f305dd · outbound

This paper cites Qwen2.5 Technical Report.

Reinforced Language Models for Sequential Decision Making Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:54.970183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:54.970183Z digest=sha256:f6214906f90c0f4d9a8d0fe2b6d1461d7dfe57625e3b03c5a2c7d8f82649c9f6

Observation 0f2007e4-d43b-4ece-837b-84e0da47fcc6 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

Reinforced Language Models for Sequential Decision Making D.; Ermon, S.; and Finn, C

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:20:58.602951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.087607Z digest=sha256:37732a330ad309fd8daf052e108ec3e79b3bb49f0849b9b5fdce7d8652fd823a

Observation a9c1cf07-8a3e-405c-80ad-0554d94eb40b · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:58.427076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.134546Z digest=sha256:9d48639beb0d9397469fe7b1f6d847d4827e217075bdd202ea58938651abdc23

Observation a4e8f578-d9bd-4c4e-9317-a8bbb40a3c05 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforced Language Models for Sequential Decision Making Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:55.288861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:55.288861Z digest=sha256:72778afbc6045b91364c6a0f4938385804478bac6a3ac48eac3e2a89565285f4

Observation 01358c47-70f4-42e0-908f-653b73d2922d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforced Language Models for Sequential Decision Making DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:55.391055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:55.391055Z digest=sha256:d802dc2bf5a6fb94c099ae14b908e37abb52ba47019d15c3f9f5cc480499bcde

Observation c3d96199-7a37-41cf-8670-755e8fc51ade · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:58.232517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.492438Z digest=sha256:2ab1374fc3485528be0b04901071d61007e8212988a1e4cac39bdfdcf6b7afc9

Observation 9bbcb265-1e03-4565-aa61-613f49a04724 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:58.085462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.589226Z digest=sha256:8a11a23698bc3c29ab17695c426b530fbc2dbbd5fc09c806e57fbbe2b75a5059

Observation 7685cd1e-7d05-46b8-90dd-410be04f9488 · outbound

This paper cites S.; and Barto, A.

Reinforced Language Models for Sequential Decision Making S.; and Barto, A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:20:57.893391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.665050Z digest=sha256:dbf35a8920c4d00fda8188e40265a4d4c4777b29db3f4afce94566e499e3a9da

Observation abc864c5-5297-4f6a-9b4b-d9f0d577812a · outbound

This paper cites Evaluation of Large Language Models for Decision Making in Autonomous Driving.

Reinforced Language Models for Sequential Decision Making Evaluation of Large Language Models for Decision Making in Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:55.719918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:55.719918Z digest=sha256:d8f0d81281b2107caa740ee70fc001f95286a6196e189d6e97ea1f59c24790ac

Observation 89f3a3cf-3e89-4dfa-88ac-74780d480c2f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Reinforced Language Models for Sequential Decision Making Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:55.795118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:55.795118Z digest=sha256:19fa1ea63f1abad315fb44cf3e3f0056a798634a2ae6c909f29ce6db160c980d

Observation 7dd9dd30-5aab-4cc1-9ac1-a8b1a95dda1a · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:57.761463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.912306Z digest=sha256:38b94a76b33f7e0de8b34288aeddf637bd3d27627294e414940ef68d913e1020

Observation dc262115-c184-4fe5-b0d6-36fa21ffa1a7 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:57.616584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:55.962220Z digest=sha256:501f8e60a9c8b9a8f581af07fa3897fe6113195f73d8485fea9930c6b8caf0ce

Observation 5f42ce11-bc78-47e2-af30-ec53f906a553 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Reinforced Language Models for Sequential Decision Making Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.054801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.054801Z digest=sha256:ba726bc73765db009a9eded954efd1156d372df74a9ddff6bb89710831a7ee7b

Observation 8aaccf3a-20ba-495a-88f2-cf358bd7790b · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Reinforced Language Models for Sequential Decision Making RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.119890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.119890Z digest=sha256:654b6daae80cf550a85d642c93e04a81feeddc1cdb2a68a850fb5e0235995bc9

Observation 7aa594cb-4c8a-478f-ae57-3b1519b1ff55 · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:57.427388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:56.176557Z digest=sha256:c5519ea18a525c5cc6c27725153341ecf5fdafeaaaa051d21702853291d2e6a8

Observation 3486c8ef-1a95-45c4-9c0e-dbe2f1a955d4 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Reinforced Language Models for Sequential Decision Making ReAct: Synergizing Reasoning and Acting in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.253277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.253277Z digest=sha256:9bf6185825d14c9c643b58b7c932fd04d8a758f3374213110d2e0cd06add41ce

Observation c039a950-0021-4e08-ad64-583bfbda2dcc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforced Language Models for Sequential Decision Making DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.318703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.318703Z digest=sha256:316673da3b8c78cbe8ca6745397e76abb2ccd6fac228eae9782a7637911dae4d

Observation 480a7a2d-b987-4bcd-b671-16d83a504fa6 · outbound

This paper cites Building Cooperative Embodied Agents Modularly with Large Language Models.

Reinforced Language Models for Sequential Decision Making Building Cooperative Embodied Agents Modularly with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.360454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.360454Z digest=sha256:7d8b1db3e4ccfbda3298debc3d62ffd912943d8f11c3f8d40b1834b1b97bb2c7

Observation f19ef7d3-1b2c-4a0d-9504-ef08ea9118b9 · outbound

This paper cites Group Sequence Policy Optimization.

Reinforced Language Models for Sequential Decision Making Group Sequence Policy Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.475462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.475462Z digest=sha256:77f284897111f0535383971565c0bb96b35e0b3d4907c7b15e72bc422fa56b6e

Observation 5bed92c2-63fe-4f30-be1e-1285c0b08239 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Reinforced Language Models for Sequential Decision Making DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.530570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.530570Z digest=sha256:dad9844504947ba87d0a93c414d86041d89d730fdbdcad08dfcae4c8caa64120

Observation 50ee1f6f-2450-4e74-b775-ef523dadca71 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Reinforced Language Models for Sequential Decision Making DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.616695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.616695Z digest=sha256:60372f191f36875ced2c4140e0f0d099d35372a92d560200a9abbc9d52110495

Observation 34b45a5f-aca1-4bfe-a42f-8cc1569ff09b · outbound

This paper cites an unresolved cited work.

Reinforced Language Models for Sequential Decision Making Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:20:57.294600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:20:56.693887Z digest=sha256:c486104819066236f66e58608de1332a903ceabbc4c860502b6b14fde6987828

Observation 30982ed7-f0ed-41fd-8af0-906022070f59 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reinforced Language Models for Sequential Decision Making Fine-Tuning Language Models from Human Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.723685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.723685Z digest=sha256:5468323602ffc48421adc6404864bd05afedc4c7de496c3bc4af25834a810935

Pith citing papers

Observation 759d0bba-280a-4393-9ad9-cc3baaeb6ec2 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Reinforced Language Models for Sequential Decision Making

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.011453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:3b47976b91d357d7180f28ccfd2cd8f66617fb27358e9c25dad014ee1e38828c

Observation 8bf19e1c-060b-42d3-80d1-189e0a0ccc6c · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Reinforced Language Models for Sequential Decision Making

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:27.138826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:dfb7a0501cd95035219c2b6d64e7830afe8c9633c61315af7cd7113918b42fd7