Pith. sign in

Paper Citation Record · LEDGER

ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2310.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.10505 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:40.843022Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.661453Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 262fd3c7-4766-406c-81c5-c18a13cbfd8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.919778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:af5ab3ed21084de4eac497542195b2df9e149493dd641728bc1ee589721e5abb

Observation cacf610b-102d-46f0-8143-b28a42d80033 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.123445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:d7a843ce6ecf9ad44f58c0cb8c2512ecb27f53ccb2f0dcdd4d5221ef85e68975

Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.409103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.409103Z digest=sha256:f7e5eb2160e0a685e2bd0ac27da9fb5ad2a953cd28ccf3ac41d3d475ab2d6262

Observation f0776b47-d595-408d-8199-0e20cbb831ad · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:5b58357b04168e9f0d31f49a7baa39dcfbe58fcf29b9e1a416b0921dc843bf5d

Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.011230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.011230Z digest=sha256:5380bb6eebca4a18c9e4dc5cde64ad352364bb1df782b46565194fb26c8c6104

Observation d6b7d1d8-ef1f-423c-9de4-6e08705465b3 · inbound

DeepForm: Reasoning Large Language Model for Communication System Formulation cites this paper.

DeepForm: Reasoning Large Language Model for Communication System Formulation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:47.066674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:47.066674Z digest=sha256:b44dbb59969a47accc946ed0b825e58a7150b7ea8314f0d70ef17ed655510452

Observation 8ce80491-f06c-4686-80c5-003108961082 · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:18.715700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:18.715700Z digest=sha256:1dd3392dc4d66ce9aaaf234521f14ced34ed7a95bc89bbf572032d5308a00583

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · inbound

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy cites this paper.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:aa5d14377311f9813e2bdfc65d296db31d8a9da8e2fbee248c35eb1c99ea6804

Observation bf6ead8d-9d7c-45c8-babe-ddda1bd24138 · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.344567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.344567Z digest=sha256:86a13b920e9686981e58df8410e9e01690ccc9dab69143357d6e0fbba6cc9166

Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · inbound

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) cites this paper.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.084746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.084746Z digest=sha256:8d1427dc6bb4ed939462395e521a933c303f6125e0ff8c364d4775a8fc0b7a3a

Observation c6447613-bc92-4abc-8e7b-a10de149a9dd · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.156651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.156651Z digest=sha256:40900229c2055156eaa2129fc2f36b65ed14784e84da08b2f38341ca32f6db4e

Observation b485999e-b648-4a9f-9df2-dd6221a6a049 · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:23.364289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:23.364289Z digest=sha256:8504e867cb6db5b0b43ce233055d7059c820b3791e8767373778c876aff2851d

Observation c232298c-300c-4099-974f-f48cb920172d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.813875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:c384874a6440722f7907acf45d1a5f3334ab528620653249177781eb79a8f42c

Observation 144fb14d-6962-4659-acae-7b7eda5d96b0 · inbound

Inpainting-Guided Policy Optimization for Diffusion Large Language Models cites this paper.

Inpainting-Guided Policy Optimization for Diffusion Large Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:46.988775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:46.988775Z digest=sha256:7ad14ba0dc435e0ae1a4831e7b00de914cd5e1a50d745b4c34822c948aa7b6ce

Observation 43b9192e-4db1-478d-89fb-1d29d469c7c2 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:21.049607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:21.049607Z digest=sha256:2aa3ea1d3f9d1dfff1e358f90e3f2d46cc4803a802f220d9d9483f6d1cdc7eab

Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.815569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.815569Z digest=sha256:52ebe74f22b0629f4fd780cccd98b21cd26915b429830c3baf2c92120b2ebea3

Observation 5dbfeab3-23f4-4704-9446-35d3bdb5c8ca · inbound

Image Diffusion Preview with Consistency Solver cites this paper.

Image Diffusion Preview with Consistency Solver ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:43:34.676931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T21:42:24.098676Z digest=sha256:ec6eb06e4e1780d0d2d63971841ef5b6c80b4ad9b2cb6e96c648d9d3cdbe0b11

Observation 3f8af7fc-181a-40bd-8b32-1bbb3796e112 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:49.367579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:87e68caecbb933babfaa8c8a4258144d1f3e2de56311c0d08a2a1d4d53e73984

Observation 5b776141-93a9-4d6c-83c9-d12f843224da · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:39.414831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:39.414831Z digest=sha256:6990d4b7ca6682754d7ea11d4fcba8a622bdfd5c2b96a3967ae9f9d472d1b5ab

Observation 13264c45-0831-4b7d-90fd-ce0557c6a033 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:08.766184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:65a44fae891d3ea15cb698612680f3b282d4fcff3e5061bb0ee7b90b1edf107b

Observation b66ac405-3a0a-4ced-a0a4-2cbd2d320290 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:46:00.902942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:7c11ecfc78f44890ec58ca2520bcd7ff795b176d2186aadd20df21c73902ab5b

Observation 0162721a-59df-465d-87c7-96a8042457cc · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:42.470028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:82b4bd920d635c9c3e30e092d465fa2c6e2a1e9d8e92526c4272686f951c706c

Observation a3d2df48-282d-404f-9325-17173f4a6ea6 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.372171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:5ffcd7daaaa592edb75ac807f21a2a7f4934f975daab874669cbaf272c062627

Observation 8cdce20e-1b2a-48c3-8dca-d2c56cc51c9a · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:04:04.452971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:a8c60e49f54fa0911da46976b2b59a0ca9da485d80bff05fd333d70e68c3217e

Observation 68130dc2-706a-4799-bdea-c49286b38394 · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:29.028245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:d039c2f80440e6787fc087365d6479176a58efa7b02b0eec01f21d530125d587

Observation a798f6dd-ee22-49f2-83d9-6ee7af1dd054 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.818446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:67c0d4828b7715097e866b9c0c5ba6450c23a89d11a5b51e66d489320c5311da

Observation ead685cd-698e-49f8-9719-4c5930606f63 · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:43:22.208303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:036f58d1dbaba9a3f6c1e52cbde50a72c85837ece754dc6396af04168272e836

Observation 2f2f5fa2-f0a5-45b5-bd3f-952a6fc66c4d · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:18:07.233735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:2c48f4f94c113e48229e8e6f29703a6c2aa2105b943b7fca47f3f793336973c3

Observation fb06da29-d7a0-4cb1-a23a-27508f0ed17d · inbound

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis cites this paper.

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.805818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T08:06:02.176934Z digest=sha256:94f499ef19f4fd9e4c192c6177b338e8626182e26fc19b2627bb789dc1bfb881

Observation 0a1d2f8a-47b6-474f-b7d9-c4d8ea84497a · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.787110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:f026d6868204cf9d0a632a7ff541cfca77da9b4e9c5f1ed4429e4060e85b4ae8

Observation 87373631-0bff-4bec-8c62-2e845eab17ec · inbound

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs cites this paper.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.588012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:a6cc22b91bb65f9f6f64b6cd9ae9390bb15bd0e71e68080c3aa4c8b45542cf25

Observation 64710dcb-d09d-4c00-99ab-0a71fbe94a6c · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.678360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:5fb986f266cce05071876ec9a77f772c91a4c588416d1a4f680376f1b9c99e91

Observation 3b77b3d6-6413-4afb-be35-9b8cd231d91f · inbound

Rethinking Groups in Critic-Free RLVR cites this paper.

Rethinking Groups in Critic-Free RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.902442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T03:22:14.586529Z digest=sha256:620236066519e539e194a26593300ebdef32db1e7e3e8711141eeb43aed27eef

Observation 2618604b-2e5e-4b78-b562-3485ea2422f2 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.663139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:5a7d45410ade2330e75cc3e01ba5d850fc4a0427a2feae9397fe7e944361fe5b

Observation 5db4b711-2172-4457-8992-12fff20a36e2 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b3f017e8791ab114820402ca25c2ebf405c6ecefb9ce2293dffe842477271720

Observation caa5c229-72ee-42f7-9314-b035d470078f · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:53.020270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:53.020270Z digest=sha256:c093f75fc6269bf41f778c36871663b75172eaa4e8b3c988a348b5eee4aadf87

Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.374494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.374494Z digest=sha256:dd90dcb320b7985524aafa3179fe556222e7c420f39b5dc92517a4dddc89bfca

Observation 21b7eeaa-3378-4cf0-9f30-494b318e321e · inbound

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning cites this paper.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.843022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.843022Z digest=sha256:7643656c86ee776c9eb4a16e29a8f61afc519f9449f295974a4ceea8767c7957