Pith. sign in

Paper Citation Record · LEDGER

ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2310.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.10505 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:57.409103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.661453Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 262fd3c7-4766-406c-81c5-c18a13cbfd8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.919778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:5c563ca5caa8a59892850c312e8b006dea620d04fb48ea2b7e5162999ff4ac3d

Observation cacf610b-102d-46f0-8143-b28a42d80033 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.123445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:2014aa70ab6de2010af2c7fe5981d4ef5aa4dcc79f5893fbbcc8ef9dc3974bae

Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.409103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.409103Z digest=sha256:ec5a1726d5239803f499bd685b67f6847847da1c3d2a12cc41a1b4fa8a2ea0f0

Observation f0776b47-d595-408d-8199-0e20cbb831ad · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:720b809aa4bcc7e6843f5994efb1a44169bc4d490a62c8e46c6bb01cff9e6738

Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.011230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.011230Z digest=sha256:5fae6af789e08fd7692a2f19cf5f3f98b87aff4feb08961e373c4110dade651d

Observation d6b7d1d8-ef1f-423c-9de4-6e08705465b3 · inbound

DeepForm: Reasoning Large Language Model for Communication System Formulation cites this paper.

DeepForm: Reasoning Large Language Model for Communication System Formulation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:47.066674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:47.066674Z digest=sha256:8ba160c2a215de6fa6dfe795c92494cb949f7ce00e66512ccdb1c10c9ed52ef3

Observation 8ce80491-f06c-4686-80c5-003108961082 · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:18.715700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:18.715700Z digest=sha256:edbc76306e69db0771008b44569ba95b9f9c23eab172f609f375345cc2f10bd6

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · inbound

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy cites this paper.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:7a9d846b800ee62c55daf8bcb608d627fce53ed6860a5936d704afd8b3d30e21

Observation bf6ead8d-9d7c-45c8-babe-ddda1bd24138 · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.344567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.344567Z digest=sha256:003778672df2b245570ba1eb99bd5c149b26071a59b9e35698438c60f37ea2aa

Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · inbound

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) cites this paper.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.084746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.084746Z digest=sha256:4d9d3b066292f3348f39d92daeb7c3310986fa6c034061142816113757e46d40

Observation c6447613-bc92-4abc-8e7b-a10de149a9dd · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.156651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.156651Z digest=sha256:82a5c8e10e02120a9ae9fc529abc5debce5556cdce62860702bbe68beb238420

Observation b485999e-b648-4a9f-9df2-dd6221a6a049 · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:23.364289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:23.364289Z digest=sha256:511a46f8f9f531143e27cc5bfc59db9ad789f086db1c6d4caa09eddf48bc2d79

Observation c232298c-300c-4099-974f-f48cb920172d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.813875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:3bf6643e7f8f8868c4745ae63ccf42f847cb858f357235d7d9f429e25529be39

Observation 144fb14d-6962-4659-acae-7b7eda5d96b0 · inbound

Inpainting-Guided Policy Optimization for Diffusion Large Language Models cites this paper.

Inpainting-Guided Policy Optimization for Diffusion Large Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:46.988775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:46.988775Z digest=sha256:a56b4047212e4d007f366f00f4016d5ed11e62b471bd02e5bd8447a1069335fd

Observation 43b9192e-4db1-478d-89fb-1d29d469c7c2 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:21.049607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:21.049607Z digest=sha256:544bb0df4681a54d0c7343038f1eb3a2598b562ceb5ee9faf15719cbf82be77b

Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.815569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.815569Z digest=sha256:619e7b9adb69b618a55989e0a9e64e2e35428689cfdfea9429f4233b4c8550c4

Observation 5dbfeab3-23f4-4704-9446-35d3bdb5c8ca · inbound

Image Diffusion Preview with Consistency Solver cites this paper.

Image Diffusion Preview with Consistency Solver ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:43:34.676931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T21:42:24.098676Z digest=sha256:0cf1eab535f884f149ef039c9f0127ffdd22a58ab896df708d368d2c0c347be8

Observation 3f8af7fc-181a-40bd-8b32-1bbb3796e112 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:49.367579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:464659fec943a732f7b825a9806ca393e0d8da12b676d2d88adefcef77513935

Observation 5b776141-93a9-4d6c-83c9-d12f843224da · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:39.414831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:39.414831Z digest=sha256:02f66eefcb58deabbe7728cd7f1c9c8038788f193dd025aa5c690cbbb1c07199

Observation 13264c45-0831-4b7d-90fd-ce0557c6a033 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:08.766184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:fb9c8ffcb29a1ea50a55ba03b3fc9fefe4cb09e92c0fce3d28b6c19276800a44

Observation b66ac405-3a0a-4ced-a0a4-2cbd2d320290 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:46:00.902942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:867fad7c517101051746796e8c0bf414737079d885daf4d050333f110eb5af7f

Observation 0162721a-59df-465d-87c7-96a8042457cc · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:42.470028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:607ee8de00c946ccda6316379379c9780556322a9b0f7be4838fdd45a4fe428a

Observation a3d2df48-282d-404f-9325-17173f4a6ea6 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.372171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:65c4b76995b54bf12363485307e2f733e8427e1b98b67a2dd91936c601b8eb85

Observation 8cdce20e-1b2a-48c3-8dca-d2c56cc51c9a · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:04:04.452971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:d43dc8a8afd18631d29eda63fba9fdec8cf260c8f089fdaf47c0cb5efe9fe765

Observation 68130dc2-706a-4799-bdea-c49286b38394 · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:29.028245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:dc0941cc3375d92c7ac321e8eb09d38367dd665e5fc556baa8740602316eac4d

Observation a798f6dd-ee22-49f2-83d9-6ee7af1dd054 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.818446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:8f60fc166e69995f4ae2006d2a47b4d00de4711bbbcafbcd420a57d0351c5192

Observation ead685cd-698e-49f8-9719-4c5930606f63 · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:43:22.208303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:3f5b61e901e8f1f57b2180c5026117002670d923c70559aadf33dcc52c32e861

Observation 2f2f5fa2-f0a5-45b5-bd3f-952a6fc66c4d · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:18:07.233735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:045b4bfe1f635e9128e81f449c6ee351073edd48b28707bf1c38be2ec0955f2e

Observation fb06da29-d7a0-4cb1-a23a-27508f0ed17d · inbound

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis cites this paper.

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.805818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T08:06:02.176934Z digest=sha256:2c8855a7b32f4e0fffd2fa91c9b5d9e2e8b2ecaf2be6c8eb0bc60d98ad0305dd

Observation 0a1d2f8a-47b6-474f-b7d9-c4d8ea84497a · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.787110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:5b85753e0007f075a779ca81a08fc6282081a13b4868bf1fc162fffa43d2febc

Observation 87373631-0bff-4bec-8c62-2e845eab17ec · inbound

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs cites this paper.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.588012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:3129f05ba71c0b9767c4bc2f771663b3b9e5280d877e05227b603dedc58f66c8

Observation 64710dcb-d09d-4c00-99ab-0a71fbe94a6c · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.678360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:dd26c7cf43ce50e85530adf9080069fbefb63cbc1139617b7e2ad030743bf988

Observation 3b77b3d6-6413-4afb-be35-9b8cd231d91f · inbound

Rethinking Groups in Critic-Free RLVR cites this paper.

Rethinking Groups in Critic-Free RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.902442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T03:22:14.586529Z digest=sha256:073fcc6fa1fe1cfe2f809f87264b0cdfa48d38ae903ebf4024cd58102c5ea328

Observation 2618604b-2e5e-4b78-b562-3485ea2422f2 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.663139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:424716e4e10df155502dae4cc7768c6f7f867d885ab5a1a9d00855d47e195ca3

Observation 5db4b711-2172-4457-8992-12fff20a36e2 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:713d926c0f15259481a478822188b7af55f116da14596ccdbb111437a982985b

Observation caa5c229-72ee-42f7-9314-b035d470078f · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:53.020270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:53.020270Z digest=sha256:afd4bfe709f9abafe234a81f036e1e05a2a0deef5ef589b1528496e7c151dbe7

Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.374494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.374494Z digest=sha256:e017b4c2bee229556a38998fbf9cb330ea288031f604d8063461ffce4f7f7934