Pith. sign in

Paper Citation Record · LEDGER

LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.18232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18232 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:57:48.357851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:58:02.852348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7ef4e4a-c011-4d6e-bb97-19aad46099f6 · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.357851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.357851Z digest=sha256:b59d397caf355e65c1e0d17b6578c909394925c32580c5aec4cd4768d690aead

Observation af1a829c-1c05-4013-b77c-5f4d2339970f · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:16:16.572195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:3ff61d0da4c771bde0bcb91ce4bf557218e19aa06ee9ee5241f518a9728808a3

Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.541246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.541246Z digest=sha256:d6f6e6a2c763beab8971a9b371449dcf45e8b8cfe8420e2338c8e7dbb95c2fbd

Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.490860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.490860Z digest=sha256:b18ebf1495949bb50c7176ea6bc23011b2c734117e434a1bbd138f514efb813c

Observation ee65c8b9-1501-46a0-842b-8bbb96af0cdf · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.917864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.917864Z digest=sha256:d8638a22031145e57dd033edf4f86e55ba153486531d3d571fd42dff9ee1db62

Observation a9c0a123-1458-4ec2-ad46-403029fe2937 · inbound

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models cites this paper.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.028658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.028658Z digest=sha256:0ffa287224e8781cfde829739a7953a75ff3dca15acc79578ad645a421b32336

Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · inbound

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison cites this paper.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.916722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.916722Z digest=sha256:07c1d8d8f3bcd4a3c7e006adee2e3628ac027f7181154ea55d649d1462bf18d1

Observation e37befca-cad7-4099-b173-2bd61a4c94f2 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.965011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:022662c4d2ecf27e8589010d95e034d44def6598352ff19b0b0b95b2000122ac

Observation c880f21b-7a36-4f16-9e8d-d49f9c34a51f · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.915040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:1e986b85db3b1103da43b9e579a3e6a21e2ae4d3dbb3b5eced3781089e95917c

Observation 2e7e27cf-c304-4a0e-9a5a-831f8cd370be · inbound

Alignment has a Fantasia Problem cites this paper.

Alignment has a Fantasia Problem LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:31:07.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T21:37:54.545865Z digest=sha256:a095c6dafe0fc4c2f38eb6105c611a39e60c03292d3416555194f877554db345

Observation 7e80a263-9c4e-46dd-9812-ae94e2c3f787 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:11.550271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:7e12c42f575cd60ca7a840af6d826d4873490cc229f34a2e00ea97a3cc6ad7d3

Observation af114a8d-6009-4eed-91da-26b6023ed2bb · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:48:28.538526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:4e57391852f84d2459f636bf46a92664158f5f13e8302ad1df6a3fd81b6d6726

Observation 1490a1fc-f5e1-4023-b042-81771be9314e · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.854176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:bb2026156918f67b0b9ecd9bd9746de1cff593c8e753daf1705321daa5b4ed75