Pith. sign in

Paper Citation Record · LEDGER

LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.18232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18232 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:57:48.357851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:58:02.852348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7ef4e4a-c011-4d6e-bb97-19aad46099f6 · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.357851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.357851Z digest=sha256:a2ee0757222deff315bc668cc7ea8b09b2877372e968bbdaed46d8e99635fc68

Observation af1a829c-1c05-4013-b77c-5f4d2339970f · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:16:16.572195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:47cf9ce8be47ba25d4948c94bd5247539951d2700133ddb3bd624156505ea874

Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.541246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.541246Z digest=sha256:d6f6e6a2c763beab8971a9b371449dcf45e8b8cfe8420e2338c8e7dbb95c2fbd

Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.490860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.490860Z digest=sha256:b18ebf1495949bb50c7176ea6bc23011b2c734117e434a1bbd138f514efb813c

Observation ee65c8b9-1501-46a0-842b-8bbb96af0cdf · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.917864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.917864Z digest=sha256:d8638a22031145e57dd033edf4f86e55ba153486531d3d571fd42dff9ee1db62

Observation a9c0a123-1458-4ec2-ad46-403029fe2937 · inbound

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models cites this paper.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.028658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.028658Z digest=sha256:0ffa287224e8781cfde829739a7953a75ff3dca15acc79578ad645a421b32336

Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · inbound

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison cites this paper.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.916722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.916722Z digest=sha256:07c1d8d8f3bcd4a3c7e006adee2e3628ac027f7181154ea55d649d1462bf18d1

Observation e37befca-cad7-4099-b173-2bd61a4c94f2 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.965011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:783f154d8598719a3b37b302a4d85e8ead38ffb686932f9169cdff5b7627387c

Observation c880f21b-7a36-4f16-9e8d-d49f9c34a51f · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.915040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:84542ba7fba553f020cfd5ee9406037b6bdca12e6ad241835a8c5ae695df6ad4

Observation 2e7e27cf-c304-4a0e-9a5a-831f8cd370be · inbound

Alignment has a Fantasia Problem cites this paper.

Alignment has a Fantasia Problem LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:31:07.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T21:37:54.545865Z digest=sha256:abef7a2ac62fd90486d042558263ad1159b59051d6abde1d45e07b1da44d3f46

Observation 7e80a263-9c4e-46dd-9812-ae94e2c3f787 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:11.550271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:3ddd381724fcea2ea8e2d0c4be6fc80be4a0dc88255feed43623b35b4ecdf0ca

Observation af114a8d-6009-4eed-91da-26b6023ed2bb · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:48:28.538526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:5b72ea65b4650c5e2df1ad2d88b97dc4a26346fccd57962a7dfe05c3ea941305

Observation 1490a1fc-f5e1-4023-b042-81771be9314e · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.854176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:e71a90970d5719cf6c2d3d86e548e0462d012f90691db44a0bb9443a9a063e07