Pith. sign in

Paper Citation Record · LEDGER

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 82 inbound Pith citation observations for arXiv:2505.24298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24298 v5

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 132 of 132 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 82 of 82 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.537525Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact15
  • verified fuzzy15
  • unresolved14
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f1b5f65d-5401-46eb-a59d-e09f50458e72 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.034522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:4216dcf516f2bfcb5ef93f542fedc668c2b68223a81e012f2a4c7b04b4ccab65

Observation 04c98360-1a8c-4d4b-9851-275de73d2520 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.024153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:8151a13481e8b752c332b21a6eadb0f35123ec4d14962279c3c93d8244935a35

Observation c68c561a-5cbb-46c8-b5b1-78b1f6320f9f · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.092893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:ef5e2ee660435dbad2cfcdcc71aaee4b06fe4a8e39fcbc0d7eb2c9a1af1d764e

Observation f6726aac-2bc3-4aa1-867d-592e16976e21 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.061360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:d1ad16872f8a29b46872b9e11547b2cf0d2cbfaf6dac4880a1dbaccb9e5ed181

Observation 907e8104-ed56-4057-aa42-6ef156b66393 · outbound

This paper cites Espeholt, H.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.095148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:aa95596018af90d7ca2cea4cb28bf924afe55690f8f2ac4f601f0ddcfcf711f9

Observation 09668f0b-c48b-454b-8315-75fc0a8cb846 · outbound

This paper cites Espeholt, R.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.097530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:2801c0aee08eca33b94b14a87a65a13de7fe2d62155b272c318914c7f18e8d2d

Observation 65c5e111-101c-433e-b470-277b310c206f · outbound

This paper cites Hendrycks, C.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hendrycks, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.099511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:fc73e186d1f14b7a4ef8a20174653f862d94cde6eade6d493195158f205c1a4f

Observation 50b91d8b-0a5f-45b3-913e-904d135b37b4 · outbound

This paper cites Hilton, K.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hilton, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.102020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:c945cd881afd5c689a1ced0407454f662407eb535205e0ebb09ceadf4f7d23ba

Observation ddd1dcba-6b8a-4ca0-9e75-122efaeeb268 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.069097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:8b2dd349810c980addf736fed6946bf8b2c1dff8302e080e8a4d1bb64c9283b4

Observation 1a8ba6c0-7b7e-4f1c-a6ae-8e5caf4218e9 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.104566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:79006550fa25c2c9fb49751d3af714c1c44e297875e115b6ee1904f6fae40f7b

Observation d379ce49-8768-4e9f-824c-cb32e3bfe0d9 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.107087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:8086240bd8ace3db0b58e8df0509977a888f96e0ab9aa0c0c771566488202b9c

Observation 842f1a27-327f-4e04-9398-54f70b48c754 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.109287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:83101d9ebe7da4c416cf5dae1ae1651ee645aa31717783894b7062a7f8e4a90a

Observation 006a0879-9795-4d9f-a0a4-1c4ed76a61c7 · outbound

This paper cites Kapturowski, G.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Kapturowski, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.111431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:20e79bb98e238c8eef1a568fae66efd50887bd45f1c5f9a9fee29937814459ef

Observation 2bcfd707-7ecc-45ae-9639-e65d83ae8a57 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:21.973774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:b4f1a6784c3537d00939f5384c71e264e87f6a2570207a028e288369c1596cd3

Observation 3afe3235-b742-4fdf-9ba2-163c2ef1226f · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.113383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:9789fba9202cc90f11be81b97c1d207ef545671bd5184184c537043bd85577b4

Observation a83167a8-ffea-4f7c-80bb-72c2b906a51b · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.115514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:6a00525d7594b705e97822a8c2bc28ed3bd5530ace5f662b17b7596790273cab

Observation cd3d6ccb-9aac-4db8-8979-46e7f573bb14 · outbound

This paper cites Liang, R.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liang, R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.118256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:390132649c8d6a7941b319e5015ce645b5b0a969022699abb35d1c90e46ca5ff

Observation 00c3bd1a-c9f2-4983-9911-f4b4947d08cd · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.051843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:36b519f72c965cda5e073e36f596003aa32e1d48e17f88217dcafcd090cc6a56

Observation 4c66119e-5d22-4e30-b656-2bc41e0ceaef · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.120199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:25e0887863bf6e4bf2fcf66cabf1836f387fb00f78394a9be52ccaa7e3bc49e5

Observation 5b2a3e38-cc70-4796-8b56-a8489d7d4d67 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.122299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:e5460e4c2d1939c23138d4a61d43a8a01cbb28186d1f5ccb3d6113b06708a26e

Observation bc842cf3-3971-4022-b753-e1754b602810 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.124185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:1c3799a77b3a06f3b1dc2c1562f0cf2688d7a61cb999e2a3025643b792c84909

Observation 5187a572-3c6a-49dc-b575-ee6e3fe22140 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.126219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:cf99c32b5a97ae8525100fdeabbb5b8cd30fd3d107d40d0b3f8a40ce0614be9b

Observation 2fe33030-686b-40dd-a9d0-a8732103d5fa · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Specinfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.009165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:803fdd3b6d0e22f846661b13853349222fb3771314e04c136dc8cd1a2bd545be

Observation 9ac1dd7b-70c8-4f58-a4b3-041e34200782 · outbound

This paper cites URL https://openai.com/index/ learning-to-reason-with-llms/.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ learning-to-reason-with-llms/

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.128959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:462e344e3d5a321491a71945a0255f434f73a35a2a5ebd51cfdb9d6f0ff2aeb9

Observation 22c0fc43-2c26-490e-8a5c-a489cb6e1b09 · outbound

This paper cites URL https://openai.com/index/ introducing-o3-and-o4-mini/.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ introducing-o3-and-o4-mini/

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.131543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:c1073e68f0739dd7106123195ea107f6832b85c537ec4b6cf49504f26f7ce315

Observation 17be9c91-965d-422f-85f3-be94cfbb6859 · outbound

This paper cites Ouyang, J.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Ouyang, J

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.134269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:103c05c275b0f83a939a4af80e0c41ba7d1b09702389e433146474b322ef69ec

Observation 746cce49-6e19-44ed-a3ce-a0c2003f75ea · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.136472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:c38a32ee295fe9a8c651d12946957db86bc0c242190dbf5a96867dad63d90d1a

Observation ec081b8d-129c-43c1-8ab9-c91d41defadf · outbound

This paper cites Paszke, S.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Paszke, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.071955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:256fd4fe631a3382a1540fdb722376c6f56b76bad424d08a54f31b358b2aa431

Observation b34c7e4e-1a27-448a-892a-0bbd00897967 · outbound

This paper cites Wiley Series in Probability and Statistics, Wiley (1994).

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Wiley Series in Probability and Statistics, Wiley (1994)

Reference 39

Resolution
verified exact
doi, observed 2026-05-15T14:24:22.012386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:ebb1dcde83a6c40c3196a6a8331bb33377bad2eb490887be0e9f936f034fa580

Observation 4cfce521-ebf0-4a16-bfe9-efb03efa7f1c · outbound

This paper cites Generalized Slow Roll for Tensors.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Generalized Slow Roll for Tensors

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:24:22.016165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:3d27c1c67f261ad46db39c73d2f7688100d90a85c529648f3db76e9db1640402

Observation 75a53cec-ed67-4faf-a111-953e497e4cbf · outbound

This paper cites Schulman, P.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Schulman, P

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.074665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0b546df09d74275e7071774bc6a08eb6e272dbd53d3dbea2c6fbd4c7a02babff

Observation 1df6b9ad-1039-4be0-b252-ddcca0f8182d · outbound

This paper cites Proximal Policy Optimization Algorithms.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Proximal Policy Optimization Algorithms

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.038379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:736d03913df343f03b6408e769b65f371b310190d843acec2afc2e65a631fac2

Observation 80570ad7-d6ed-406a-8833-791c6d4c419a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.042354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:6d1e0da1e006c3d3c81dfc8b002daf61365f2426fe52dc8b50ba592fe526b670

Observation 55dbd829-c259-4871-b28d-53e033abd762 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hybridflow: A flexible and efficient rlhf framework

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.003970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:20ba526a5bb4c99a415a1cf426258a7c9d18e8f9b6a1161277f240e1871be2b3

Observation 9b490396-40b7-48ba-8703-c5eb8ac69ba5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.046981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:18e0e9ef43473a7a82dd7a2ab3fb8a6505b6636b180ad05bdfb197bcde3ed39f

Observation bf02d8bd-d30e-42e4-80c0-b9cd86e0f88e · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.077208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a12f3a4895e8307f4937ce4e94666d1376c8101ddff7793b01a00e7532712b18

Observation a9bb1a91-530b-4c64-91dc-330177927b5a · outbound

This paper cites INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.057134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0681d6c50b951401f3d90b9284d0ecba4156e18fe07153d0281ef89057bc3159

Observation ac46e50e-a781-4a67-b569-1d9158477b2e · outbound

This paper cites Vaswani, N.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Vaswani, N

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.079530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:7f50dd832e0a99fc55947992ad5cd4caee5c1e543c791e959c9b22f8ff775d76

Observation bb8f25bb-19ec-4261-b3a5-29ee5836815c · outbound

This paper cites Liar” ends the game, then both players reveal their dice. If the last bid is not satisfied, then the player who called “Liar.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liar” ends the game, then both players reveal their dice. If the last bid is not satisfied, then the player who called “Liar

Reference 55

Resolution
verified exact
doi, observed 2026-05-15T14:24:21.993334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:995a299b5c4fee93fb62ff9884dab9ef1e0a47de6ac626d91bc6dd9e677a3c06

Observation 776aaf06-5ed0-4e76-b9fe-01a51c09984b · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.081758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:104848f08fd123b4ba141e6b1923a9a1ffdd6bba513f74e8480992f70fb4799e

Observation b991c518-178b-426d-a6ca-76576e505494 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.065868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:1d08df12c042cf37748abc76c638ec1aa87a15356785e4aae8eb447b983598be

Observation 3628c109-2250-48f9-95a8-63f07031ee6a · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.083916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:41c7f8b4e89aa6f33cb92934d96d05444e320e0e50650ac874bb420881baee04

Observation ed6d1c0e-4631-41c4-aa10-076c7defaa9f · outbound

This paper cites Job Scheduling Strategies for Parallel Processing pp 44–60.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Job Scheduling Strategies for Parallel Processing pp 44–60

Reference 65

Resolution
metadata mismatch
doi, observed 2026-05-15T14:24:21.978165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:2cc0d13edd04c40579ae25f119a3758ca89b11e40e11b2cf4799b46bfa8ab661

Observation 79e61d0b-a2c8-4729-9c60-20162a21aa35 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
malformed identifier
local_arxiv, observed 2026-05-15T14:24:21.986322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:088c7f5360b5a4f11315509607b78d2aaa165654293a7950adbdb0138ffeb0ce

Observation 4fd6d89d-55ff-46da-b1b7-fae958d9c605 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.020162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:e9e22642f9ff8e0d0de8b0dcb766c146abc33692cd1c94b140fc85b119dc178d

Observation e28f579a-43c0-4b89-bbaf-fd5e5bc8e433 · outbound

This paper cites I am EdgeRunner AI.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning I am EdgeRunner AI

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:21.999394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:e35afc795a496f006c5888bd905c88fa9da5fd51e5517592dc6a9b28fe91ebcd

Observation 52324d64-6c04-4d1f-801f-61360ca221ce · outbound

This paper cites Zheng, L.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Zheng, L

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a6e219ebecec639a019f0ccceb5c518436a795ea4faef36303988b72088bdbd2

Observation 59953177-8abe-4da9-bdfe-30c73e0f5279 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T14:24:22.029099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:1c135674df6a711371459effabd1a44466e13e4e13d7cb02838f226458a01f34

Observation ef40b99f-46ba-4556-bd29-cd07bb5f073d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.088384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:b1d6b7ab0f2d033966d2f1b6455a5cf0feeccd4fbe02a76a95271746ac06d13d

Observation d73722ed-f188-4954-af71-82d1754581e9 · outbound

This paper cites For most of the results, we use SGLang [63] v0.4.6 as generation backend and pytorch FSDP [ 62] as training backend.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning For most of the results, we use SGLang [63] v0.4.6 as generation backend and pytorch FSDP [ 62] as training backend

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.090551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:d3003175c55c3de78d75869d4f8a771849c4260a9fc2b6f8376536ccfdcde6a7

Pith citing papers

Observation 82ad9e0a-fd61-4bad-ad06-681f30b30f8b · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:05c8aad864c635ee56da8afdeebfc8dfd283fee883650612e6b5756e107260ad

Observation 87c5e038-0a6d-4e4b-b4a3-bd42cce82f05 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 134

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.288940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:ae7c1d5b6be0c66635a4d7c794e39166e4b32208389f7cfa8e907217f9d6519b

Observation 40d20f02-e5b1-4689-b35d-b4082c9cecc2 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:19.537525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:19.537525Z digest=sha256:3c4802088cbe4c4f4c10fadcc64b14606c468110dec0ae45dca05eddc84840c3

Observation 91e4a796-2636-4899-b357-0aaf145a915b · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.648438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.648438Z digest=sha256:d7491e04486763882837a48ad04c70224301801015fae1b32f62e1ff6e7d5e73

Observation 25fd9d1c-33a4-4ccf-a6da-17df1b42fefa · inbound

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training cites this paper.

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:50:23.494922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:50:23.494922Z digest=sha256:7f0b34e294353c3e170cc29b03305346b91f51a68cd70ff50d36912a38c08c38

Observation f0f1af5d-8ee7-4d5c-b13e-b8847f51641e · inbound

BlueLM-2.5-3B Technical Report cites this paper.

BlueLM-2.5-3B Technical Report AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:53.303978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:20:53.303978Z digest=sha256:c009a1ae56cd8e6af529b2716ab39dcc7923d89c80f15ebaa7da02fb53414213

Observation 89ad4549-b54c-490c-b514-75a11640561b · inbound

MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster cites this paper.

MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:07:34.964925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:07:34.964925Z digest=sha256:48e9fd23182884c26c6cb5c13bb1376100c89ab2865047121ee3157b894d0805

Observation c27dc36d-5ebe-46fa-ab14-4e34213b8488 · inbound

Agent Lightning: Train ANY AI Agents with Reinforcement Learning cites this paper.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.142833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.142833Z digest=sha256:7d2b426e16b5cf8bcefd3168b93913915dcb85e52e9a3c1650965393181ee472

Observation 00e6d183-f73c-4242-bcad-ac5f48f1fa04 · inbound

Reinforcement Learning with Rubric Anchors cites this paper.

Reinforcement Learning with Rubric Anchors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.370681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.370681Z digest=sha256:88365e37dcb69dcf1891b8638e5b8f333a82e13ca4a302d4428c1402f68b226f

Observation 6f41e1d7-81c2-45b0-86d8-93f1976f96a8 · inbound

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL cites this paper.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.391862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.391862Z digest=sha256:9e4848acc211c1e9e6db6f99d6074aa8c56c9526e94d0594c8465d307d4b20b9

Observation 0525c134-88c7-4d1a-a5e2-4c201a03f947 · inbound

AWorld: Orchestrating the Training Recipe for Agentic AI cites this paper.

AWorld: Orchestrating the Training Recipe for Agentic AI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:17.779979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:17.779979Z digest=sha256:f969862ab77cc7c3e29e753fcf09112edb28154be34d0b416ed82b5d600dfdfe

Observation e6c60293-83df-4036-8035-987740f8026c · inbound

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing cites this paper.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.042144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.042144Z digest=sha256:63de1065a30403c1ef2e021fe49a9adbf4a4ba19c792328ad25d9ed0b39a8b84

Observation 8011cfc6-c1fd-46e9-a2f6-a47513364950 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 147

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.401254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a789d8c681234b50d09a1bbff57aac426f3a740727dde7d37f03bdb9824a9719

Observation d8d6583d-a411-4896-ac62-6ef44ab2e356 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.199481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.199481Z digest=sha256:38efe4ad5e54aaee1b9a301014b4b15566fa74297d08017201ca0b8bdb17def1

Observation 413d2afd-76e0-4a29-bcf6-2e2cd3ca6ef8 · inbound

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression cites this paper.

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:50:09.254400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:50:09.254400Z digest=sha256:7e4859c260d2f4e584c71811d941231ce8142846d502fc49ee8da0f69b3dd4a5

Observation 8f02a984-a9be-4d4e-a564-8ec48c2572e0 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:26.538300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:26.538300Z digest=sha256:0c49e4458d8416d1a627282b1a44e316837e7df99d6576b2f5fa578e527168b0

Observation dfc8bfee-2973-4ef5-ac6b-aeaf4c5df871 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.176765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:cc3c3626a5e144cf457032d88ab6251f6b0d09bfc1ed2c4e0f23098612faedd8

Observation 4bee1411-9875-4874-9e6f-8841b42ac7a5 · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:40:14.640754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:4e61e07a103e9bee16c32e9a545ee4eba506326b2669d6ee9824d0dfa4050d42

Observation aacc2d48-dda8-4a96-a6ab-3f900befb151 · inbound

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments cites this paper.

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:23:36.646232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T22:21:26.271796Z digest=sha256:982635d575f79cc7a73669b311040976ccbe26546ab8607ebafefc7f6a006b7d

Observation 70d7ed3e-a225-46ab-9638-255e8247c743 · inbound

OpenTinker: Separating Concerns in Agentic Reinforcement Learning cites this paper.

OpenTinker: Separating Concerns in Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:09.185820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:09.185820Z digest=sha256:7d2f03e1b993757856ede768a54c791a3e86218ada0befed1bf10c3a9c2be9e3

Observation 7d2d5ef5-381b-41af-b18b-44f0ae5fc23f · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.347729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.347729Z digest=sha256:4d9c39998c9cd764913c2a6f256c9415db8e95da416e4ab732eae85543578dbb

Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · inbound

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning cites this paper.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.820681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.820681Z digest=sha256:71eccb57f317d0dc8987f15795ca8ba176f491f04a3f98e8cc7af700925f5085

Observation 26e97cb3-bbe2-4ffc-aa2d-49d925b892c4 · inbound

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning cites this paper.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.092261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:11cd6a73dabf475d344420a3d064105ff006a78f643dfc189db3e2aee218407b

Observation eee55e6b-5056-48fa-a879-7ed6a6f470ef · inbound

Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs cites this paper.

Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T07:01:12.356383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:01:12.356383Z digest=sha256:265da9b3945b067713f47c695af5cc14621553fd3f3bbef7b156943339c8f080

Observation 7c0d5b6e-5517-449d-bb89-e4a7f943ed09 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:40:49.102585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T09:37:57.120779Z digest=sha256:64f8d92032396a75a87d6d2b77a858291c5476d15a54fc2e76542b5ef411c817

Observation 4c092ba1-561c-46a3-a115-ae396192aaa4 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.456057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:e6c512834a1b04b93462b5fdbdfe502730df2ecd57eccd9c60af1e2234e4e0b8

Observation af19c22a-1aa9-4c5b-87fb-42df5799b123 · inbound

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning cites this paper.

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:32:40.837833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:32:40.837833Z digest=sha256:927e44e27d21b52a289b23d2207864291f481c850b12fafe9b5bef2d7c6617af

Observation 084fafff-b067-49f9-a8bc-5ba0b9385e35 · inbound

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training cites this paper.

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:28.930758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:01:54.696390Z digest=sha256:051106599d6da2618262e1973af6feee80a38da7876adbd22d100236503d9f4a

Observation 996aeed6-6acc-4490-8760-039426f88fec · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:37:28.921689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:5ecdf394eaedc5d17d48f87b505433f849be320f4421cd8028827dfe647f2d58

Observation 082f7f1c-3122-49c3-9626-0c5d3de20b55 · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:11.710514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:11.710514Z digest=sha256:691ac76420eeb3cdb4b811674c5d15b74e2926c2e6dd19e2c117790e6022312a

Observation 0167d83b-c497-4164-8fce-0954ae9341c9 · inbound

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces cites this paper.

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:36:17.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T16:33:58.746944Z digest=sha256:5e2a6e1a38d9b6525534fa42bc7fa6b528eaf77ae138a0c295a79b9b18aafb00

Observation c0f8730c-77fb-48c0-ac05-e358dc8b2efb · inbound

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments cites this paper.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.779259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.779259Z digest=sha256:2a91c875755e632517209737a5021649c1755838792676acb05af03e71acf724

Observation 9e196520-b2b4-49e4-9d2b-1e01e7157541 · inbound

OpenClaw-RL: Train Any Agent Simply by Talking cites this paper.

OpenClaw-RL: Train Any Agent Simply by Talking AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T12:56:30.316823Z digest=sha256:971e92f038a032ad0eba2a0eccaeaa9e493fc9936cd41faa3437ad06814e9c22

Observation 9799de75-3d03-40f4-8300-dda427b53027 · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:e896f72970eb4318c30de9f6e8a195bab6fc66da311b2592ad9c8a4dd1ce1fd5

Observation a80d0625-6c83-4771-ae17-5c4bf08115af · inbound

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale cites this paper.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:79a4d965ee4379467d2e5658f42c83cd89165a2e2c07363c9166de6b09f0f675

Observation ec567ed1-d217-4fe4-b075-28002f1fae0f · inbound

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning cites this paper.

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:27:15.270996Z digest=sha256:7d5c728cedeec98ff08ba0ef83328e45a4078899705245bf0bf4ca4bdcef0f92

Observation 15ec06f5-f9ea-4654-b1b6-b027fb62a2e5 · inbound

CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities cites this paper.

CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T04:26:59.781416Z digest=sha256:2c8049508ae495b09669e9304c75989ec0f274204cc761ab897e712214c47bd1

Observation 28d5ebe2-0bb1-4463-abb6-4feedf122d17 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:6f830124e296a09fd897f74746bed1e910d41ff2177bc6913c1c4d1ac5867b93

Observation 3348572b-979b-420c-814a-95fb44087120 · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:73ef29cd4cdccaa44087c22eb494c6f05cd1279d8ee965b1054d6f31eb8cf604

Observation 8795d9fb-7462-4608-9133-2903e6a01b50 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:07710b490488fe4901a58b261226269124b7730c66d9f972016ba276c763d169

Observation 6de70064-9924-4c14-83f4-7eebe1b6b2aa · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:4b29c33a14af686a96698c02e92861671bcc2f09e4e2bde496d498ae01053692

Observation bb3aa539-f7b4-4b52-b501-bf34ab776043 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:0dd93a8d68aeb00561b921265de151a5a16a338f7711ea236714d79e4965b85d

Observation 0172fdfd-c0ec-48e9-b8f6-0b3c487c0e46 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:427422139fd5a0acded5566b7a2fd7e434e73a6c9b6ec306596912e37e343f03

Observation 60736719-82c0-4c9d-8720-03d335fe16cd · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:20da9f0792ec159df17125fe8afcf10299b3bdb122c99c32ec7a47de5a7533ef

Observation 09e56093-512c-44af-9d7f-df841e883179 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:d2f84fdad0a6705c6eef9ac7a14647b1ec5f08f6c489d1d089ee415c2837d4ba

Observation 25b28170-02ac-4885-9097-078af6c18838 · inbound

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration cites this paper.

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:46:59.400834Z digest=sha256:e95627d0d67e61dc0172f3a5667fca9fdaba960be3e43af683a4e4b8d99d30af

Observation 8d917e05-0452-4a23-8e23-8177d57a79de · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:d4787716cdb0aaa8b6d4a81a74a24024c7f64136fcf382eda78cbd7bbf12400f

Observation b8030e71-5ff6-48c6-a20b-9ad790b99e42 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:818d6057e40d407c8b5414a5517d7fd106ef2410fdc8953a00eef745f38bac2c

Observation a2b86f5d-4b1c-425b-aac8-ef90ab2e10b3 · inbound

Position: Agentic AI System Is a Foreseeable Pathway to AGI cites this paper.

Position: Agentic AI System Is a Foreseeable Pathway to AGI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T20:10:36.101426Z digest=sha256:0e12763830b19212065ff1cdd7935441cbd84fd502e82a7fa28c91ef73e53c01

Observation 61174cbe-8304-48f5-95bc-82e7afc8816d · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:a02cc583b5d6c7500f07df7cb38e58829c2037701ae149b721a561793abe5696

Observation f2ca8497-d292-4e85-b552-d7a97237461b · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:6c2616d33244de7cc9748644cd20e7a2e06fbd6efe7d94f6f4d4a8f338a0558a

Observation da7d6f53-7e64-490d-be84-ca520310e3bc · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:02:42.448329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:1d114e4d1578a6ca810a0af1a24a7987ec4384c9f97e749c643ad26dbbdfbfd9

Observation 83b5152c-cd27-45c1-90b1-7541046ec6ff · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.758614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:dd90ecd3215e6aa09212e595d1bc490ea7aa404500715a19676ba8866576cd0a

Observation 4f0699d7-5473-4da1-8a3b-44716b2409f2 · inbound

DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training cites this paper.

DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:09:12.222518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T23:07:18.427357Z digest=sha256:37feaf83ae709b2be76bd17a005a9b2268d7d3a169ba091ec81c3ea0225dcd0b

Observation fe810c69-4791-45df-a5f5-e6c151d4e47a · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:34:02.628850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:0c114e53fc3028c49c3aa8641fdb34502acb759e20901e475e66a0e66e047ef6

Observation 15a3adf2-8211-4f33-84b2-6124d127428b · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.399223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:3b94573474a67cad76e2515b57683ca34731097ad61792ac90fa03cd9db53066

Observation dd84ebee-5311-410b-a8cd-666e49a72249 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.046588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T07:59:43.755196Z digest=sha256:22da92bbdcc2aab57fca504595830ccbc2a0e28959c5459d7e2cb198496e66fb

Observation f6faf070-eb05-4ad1-8e18-b5aed790f443 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:50:23.726282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-25T05:49:08.484663Z digest=sha256:a2016b542b8fde9aecdf9cfa473b58dff3f2e4700c63bd098b3404d35dc95f17

Observation 63abdda3-635c-4f93-b652-925b21489d17 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:57.966303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T18:01:38.509794Z digest=sha256:ab18100a0eebe9a61d8fbb41d4314e4b2e28c79ffdd9c5efd2151874763fb250

Observation 21bb12d9-3fbd-4aec-a1a0-8f5b899e1174 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 237

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.542325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:ad4d11d3b587fcdfd95ba69ab4db135b8a61f8500c071b394db50ae9e42800e0

Observation 4c3d9624-528c-4dba-ba31-706b18038b2a · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:13.397428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:17bfff7afb8202f0f81749336c0ef37af1ebe7eea706107ddc51fac08b164828

Observation 912f1d9e-9ba2-49c2-a818-18da27152940 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.919524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:26:35.100859Z digest=sha256:a1bb13de9d10be1a2d3514422d91f5d73b89e3935b67f718e5c25624b4146614

Observation 00baaa07-cffa-4126-be70-b01ac06ab3a3 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T12:27:08.279842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:27:08.279842Z digest=sha256:030cb1ae2d73b6337a8db72e24581d3ef50d7abae1d4e6d3a196cecf43bbf5cc

Observation 475c3c1b-5bec-4d66-b206-66aed008a67a · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:41.093107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:a19c54171698bc228790e11263a1bd7cc776154608fa9e746bd25b07f97555bc

Observation 43f1bb11-1658-4ba3-8889-5204cecca945 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 131

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:26.415608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:8806a804fc2bd5375917f81bf698c9dfefb8fa7bda3992eb1c5c53ba8d2603bf

Observation bf90f729-bc4c-4830-bad0-4c32d4ac6787 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.826024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:519284d69d420b5e89a793f8ad21fe296cac8032fad6aba04c44c334cbaefee2

Observation e5c4347e-bbdf-42ec-b2c4-c64d284cef51 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:08:07.825297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:0e85e9ab6af9cafd7712452e1abd84517f04f75b8f0c409eaf4464814d8c6fe6

Observation aa490fe4-6ae8-453c-b334-d1b6aa5aa84f · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:17:25.513987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:eda963fdbc19f9305c4e46c569105f1d44cd2817df97b04c55d7a75607c19141

Observation 2dbde3c1-555d-477d-a89b-deda0b49132d · inbound

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training cites this paper.

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:24.703409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T19:06:16.100762Z digest=sha256:ab92f0456d0a79f55bafb428257b162482f5375fd702b67f6bcea9a4a016e7eb

Observation 7b89e743-e38f-407b-b510-32c2ae4d4f9c · inbound

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning cites this paper.

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:39:58.209992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T02:56:47.102678Z digest=sha256:6c71441d9a56dec2cfed24a6ec28ff3dbb164f5010ad879c752281add176fd34

Observation 5149cbf7-dd78-4fc9-8aef-3cc9b5dbac56 · inbound

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning cites this paper.

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:52:35.159984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:52:35.159984Z digest=sha256:0491c236ef70724d26d0f75761ce06b291db67bdc66c26b55111133447a48306

Observation c1d0c60b-fe75-4399-b3b0-a1d287f2b334 · inbound

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF cites this paper.

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T05:43:07.654879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T01:23:07.441218Z digest=sha256:01840d66b3e1e7de8595d0b9d5f365da652a1593bf6b664d369f7960dfe6c8cd

Observation 32bec269-cf7f-47ad-a24a-e0b2e9c40ea8 · inbound

Trees from Marginals: Autoregressive drafting with factorized priors cites this paper.

Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T21:57:39.209960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T21:49:34.512370Z digest=sha256:e384f9469606e70add875460c76366275fd39edf1f43e0919c256f6f2019f794

Observation d5e0de3a-0c46-403c-b27a-60b45309e067 · inbound

Trees from Marginals: Autoregressive drafting with factorized priors cites this paper.

Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:58:05.356765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:58:05.356765Z digest=sha256:595623cdbcbfa214bfd68dc97c5b5b3c33384eaf0574b1846a106f11b430ddff

Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · inbound

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning cites this paper.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.428422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:e3d54bb861e1c3afef555406e19a0b07c814bd951bb20fcac58c724cb1dc2b37

Observation 55ea7ddb-1643-47ed-84ce-597ab67dd347 · inbound

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training cites this paper.

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-07-13T04:42:35.589143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:42:35.589143Z digest=sha256:0983dacebc64d93267889c82b60c03bcf0be70a9e2d149b1fb4457f9c74825bd

Observation f0bd608c-2077-4856-a6e6-5c3f70d9e353 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:37.177652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:37.177652Z digest=sha256:cbd2f9114839e76c8b3603d13100adf000ce3d4c83e73e04e24c27cb51fc41b4

Observation 14b69ad0-a0c3-473e-83bb-8a8b6deaf7cf · inbound

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning cites this paper.

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:29:31.579850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:29:31.579850Z digest=sha256:20eaa5af21a1ea3dbf27c64b816d9cded03e9b206af714ddd7542f3d730ca00a

Observation 7f2db113-7f86-4a24-8845-7519a8acf1ce · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.853734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.853734Z digest=sha256:9d6611a3170e23363521d110918b2f0f6e7e885714bd5f94a005bd6c9b2c3f5f

Observation 4cd5dede-04e4-47e8-a00a-115009c85765 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.311342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.311342Z digest=sha256:e3c167fd3dfed5a39d865cdb5c24bd477f1889626672ff53ea78ee3771bc4258

Observation 4a1c46ae-723a-478a-990e-bf9e9890cca4 · inbound

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training cites this paper.

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T11:13:59.560465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:13:59.560465Z digest=sha256:018a9c9107c6ab3501cb3851c19e2c01020547b6c9bbc90173a64f77607ef70c

Observation 71f91b49-f0e4-484c-a3be-bae211de684a · inbound

RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning cites this paper.

RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:21:54.676497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:21:54.676497Z digest=sha256:5ae24b98c84922a106887e0cfd0f449d76ef81acc9f0e306118c357866715549