Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2310.10505.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:57.409103Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.661453Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 262fd3c7-4766-406c-81c5-c18a13cbfd8b · inbound
HybridFlow: A Flexible and Efficient RLHF Framework ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cacf610b-102d-46f0-8143-b28a42d80033 · inbound
Process Reinforcement through Implicit Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0776b47-d595-408d-8199-0e20cbb831ad · inbound
Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b7d1d8-ef1f-423c-9de4-6e08705465b3 · inbound
DeepForm: Reasoning Large Language Model for Communication System Formulation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce80491-f06c-4686-80c5-003108961082 · inbound
Formalizing Learning from Language Feedback with Provable Guarantees ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · inbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6ead8d-9d7c-45c8-babe-ddda1bd24138 · inbound
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · inbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6447613-bc92-4abc-8e7b-a10de149a9dd · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 236
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b485999e-b648-4a9f-9df2-dd6221a6a049 · inbound
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c232298c-300c-4099-974f-f48cb920172d · inbound
A Survey of Reinforcement Learning for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 297
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 144fb14d-6962-4659-acae-7b7eda5d96b0 · inbound
Inpainting-Guided Policy Optimization for Diffusion Large Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b9192e-4db1-478d-89fb-1d29d469c7c2 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dbfeab3-23f4-4704-9446-35d3bdb5c8ca · inbound
Image Diffusion Preview with Consistency Solver ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f8af7fc-181a-40bd-8b32-1bbb3796e112 · inbound
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b776141-93a9-4d6c-83c9-d12f843224da · inbound
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13264c45-0831-4b7d-90fd-ce0557c6a033 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b66ac405-3a0a-4ced-a0a4-2cbd2d320290 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0162721a-59df-465d-87c7-96a8042457cc · inbound
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3d2df48-282d-404f-9325-17173f4a6ea6 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cdce20e-1b2a-48c3-8dca-d2c56cc51c9a · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68130dc2-706a-4799-bdea-c49286b38394 · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a798f6dd-ee22-49f2-83d9-6ee7af1dd054 · inbound
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead685cd-698e-49f8-9719-4c5930606f63 · inbound
Self-Supervised On-Policy Distillation for Reasoning Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f2f5fa2-f0a5-45b5-bd3f-952a6fc66c4d · inbound
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb06da29-d7a0-4cb1-a23a-27508f0ed17d · inbound
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a1d2f8a-47b6-474f-b7d9-c4d8ea84497a · inbound
Explicit Critic Guidance for Aligning Diffusion Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87373631-0bff-4bec-8c62-2e845eab17ec · inbound
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64710dcb-d09d-4c00-99ab-0a71fbe94a6c · inbound
Trust Region On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b77b3d6-6413-4afb-be35-9b8cd231d91f · inbound
Rethinking Groups in Critic-Free RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2618604b-2e5e-4b78-b562-3485ea2422f2 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5db4b711-2172-4457-8992-12fff20a36e2 · inbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa5c229-72ee-42f7-9314-b035d470078f · inbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · inbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.