Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.18232.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:57:48.357851Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:58:02.852348Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c7ef4e4a-c011-4d6e-bb97-19aad46099f6 · inbound
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af1a829c-1c05-4013-b77c-5f4d2339970f · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · inbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee65c8b9-1501-46a0-842b-8bbb96af0cdf · inbound
TextAtari: 100K Frames Game Playing with Language Agents LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c0a123-1458-4ec2-ad46-403029fe2937 · inbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · inbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37befca-cad7-4099-b173-2bd61a4c94f2 · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c880f21b-7a36-4f16-9e8d-d49f9c34a51f · inbound
Specificity-aware reinforcement learning for fine-grained open-world classification LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e7e27cf-c304-4a0e-9a5a-831f8cd370be · inbound
Alignment has a Fantasia Problem LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e80a263-9c4e-46dd-9812-ae94e2c3f787 · inbound
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation af114a8d-6009-4eed-91da-26b6023ed2bb · inbound
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1490a1fc-f5e1-4023-b042-81771be9314e · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.