Pith. sign in

Paper Citation Record · LEDGER

WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2411.02337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.02337 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:19:15.589837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:20:00.041093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3daaa13-a142-4469-86bb-cb0838180c28 · inbound

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection cites this paper.

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:33:21.204452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:33:21.204452Z digest=sha256:d757bd923232eb875b2564ec1534a100f0be4e3a62b38c69ede981e99025e03b

Observation 3d05642c-2028-4729-99d8-165576046229 · inbound

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning cites this paper.

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T11:25:49.165415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:25:49.165415Z digest=sha256:5d14869714a26588639baec6baf263010a16ca01fdd1731c3240327420f74de2

Observation 1a439e27-a856-4426-b5ad-49f735016f73 · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:32:18.628495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:95809c6e5e037e706a94e253e6c005f5e4057e54ad47a6d6c79a77e50ed14a06

Observation 8a08d509-760c-4940-bbdb-8350116a3a5f · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.751909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:39740b63977a3c8588781d1fa0a6babd64bd61fc5b61f639ac0ad91cd7944359

Observation 4efd62e6-4159-4d0d-b049-b694982913da · inbound

ProgRM: Build Better GUI Agents with Progress Rewards cites this paper.

ProgRM: Build Better GUI Agents with Progress Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.329762Z digest=sha256:1eee1f152769947cb3c71d16c5d295aa68443a8cfc9518267a42f60e55966c9f

Observation c754b48f-8e54-47d5-923c-ac41bbb93d9f · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:02.794899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:02.794899Z digest=sha256:a64776358f4d43d841ba24786aad70ed2e8fdaf5b2b8c48e723fbf53563819c2

Observation 84a7a678-435d-4764-9a4c-23ce173b09bc · inbound

Agent-Environment Alignment via Automated Interface Generation cites this paper.

Agent-Environment Alignment via Automated Interface Generation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:32.880275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:32.880275Z digest=sha256:e9bfe770c730ef5ef7e079cb594d76d1903506a18f6df7e64092e3bfc7913586

Observation f94e7756-8544-49e9-941d-b4d6cd891db6 · inbound

AgentDNS: A Root Domain Naming System for LLM Agents cites this paper.

AgentDNS: A Root Domain Naming System for LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:11:08.422497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:11:08.422497Z digest=sha256:e5e145da0f17f03cdaa49bc4eda943cd8cf17fea7daf3cd534b766d351a7b673

Observation 767db96d-c03f-4295-8692-071b489e8168 · inbound

ZeroGUI: Automating Online GUI Learning at Zero Human Cost cites this paper.

ZeroGUI: Automating Online GUI Learning at Zero Human Cost WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:24.310946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:24.310946Z digest=sha256:356d445e648bded60dfc86a33960edaf340fc5a52105e265ac8af2dc84879f09

Observation 14a271ba-3cbd-410d-abf3-89156620648a · inbound

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation cites this paper.

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:33.721752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:33.721752Z digest=sha256:afe39993f6b8315ea4046381a6e774c292e4b546ece4c13161903384273f7b9a

Observation dcba1132-d734-494e-be0a-fc5a787f1ed5 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060344Z digest=sha256:5366ed60dfc42b958c1700499de8f4de72c8eca979d5181de715a8fe1fbea50d

Observation 59ccd5d8-9b02-48c3-8bd2-ba65dad8f4ba · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.716794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:42428e4336dc1d879048530353854a5a7303f3bec2e0170e6e76a1e2527e252f

Observation 7498d1e4-28c7-4c4e-bd30-e41f5108b9a1 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:21.810260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:21.810260Z digest=sha256:e6d59e943bee2bff9e1d733534c41577743086d9d40f89d7532d8daa953e231d

Observation 2e782046-7782-462f-9b1e-06e866525791 · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.978039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.978039Z digest=sha256:c3a6c0c88f83562c24a5e24184846d03503c26d4e96ce6cd368c5c4cea5efc4b

Observation 23b686ae-f595-477c-a155-9ea064784556 · inbound

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System cites this paper.

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:04.613379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:04.613379Z digest=sha256:acc07ba063da7abc368baf69be4e8633fd74f8949ae63f103490af57cd6702dc

Observation 584b1166-6f2e-4fa4-b163-63d343223acf · inbound

Build the web for agents, not agents for the web cites this paper.

Build the web for agents, not agents for the web WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:00.763378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:17:00.763378Z digest=sha256:6951f3f1c60cf73be6d35f335e470bbc871285ae80efe1846813e2ed46f86f04

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · inbound

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards cites this paper.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:a2afd917cd78fbcaa352bd343a5b37bdfc493b41665333e5560cb96f44f6858f

Observation eea4fae6-f494-4e05-9c8d-05c761871989 · inbound

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents cites this paper.

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:27:37.434374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T00:27:37.360221Z digest=sha256:f6f95757c944cde7e890886cc2de39317e1426223868fac82402abe03f508423

Observation 4a1f1581-f669-4224-9f39-26e04b903bbc · inbound

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis cites this paper.

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:01.164687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:01.164687Z digest=sha256:8f4655a486221cc8ae92b8ee577af1ee6c561d3c6e841cfd5ccef6beefba692a

Observation d5f4170c-1008-4f8d-881f-50d4d395223b · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:cee8da6085bc9dde50e1fa84171052d7e1745e3732d2ce5a7acd403de83bc3f9

Observation 56fc05ce-ba75-47ee-88b3-2906e94b33b0 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:50.893647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:50.893647Z digest=sha256:6e0ab6f6d367ddbbb18fa2d0538c15a742b5df3c89a27bc9faaba91ab6d69ff3

Observation 3420f575-8354-49a0-a0ac-002bd0281248 · inbound

Cognitive Duality for Adaptive Web Agents cites this paper.

Cognitive Duality for Adaptive Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:36:21.986350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:36:21.986350Z digest=sha256:29cffb3a1cceed82156b90e456fec899cd34cfaf7401f0badf44b7cdfda18d30

Observation 8ebe02db-d3f0-423e-8c0f-2cf95809e21e · inbound

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making cites this paper.

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:24.747926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:24.747926Z digest=sha256:de4c9dad5ceb51caaa53d89a576c9081cb05e276130b8a7559845d9ff5b25637

Observation a32608b8-1845-430f-b8b0-6d33fa8819dd · inbound

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward cites this paper.

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:21:50.831536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:21:50.831536Z digest=sha256:f31e400af6687fdad41cfa0421ccd83d7f8608f7eb4522c3d5f990121a6e3e2c

Observation ceb3ec28-6bb5-463b-97f4-03a7fcc6ba6e · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.759166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.759166Z digest=sha256:17a0274ff30085403303cbd47fa3226276368c245e25843ac7a4daca71cf90e1

Observation 6841f209-0436-4463-aefc-e54ffc9774f9 · inbound

Symbolic Graphics Programming with Large Language Models cites this paper.

Symbolic Graphics Programming with Large Language Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:10.896680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:34:10.896680Z digest=sha256:6a2edc37f275cc0b46f7657cce8729e6e796d5e757eb4fde01e2ccd4d7ebd20d

Observation 404606f8-a388-4e91-bd77-0ac751fe4792 · inbound

A global log for medical AI cites this paper.

A global log for medical AI WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:15.570579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:15.570579Z digest=sha256:978bfee6c31e3b12e9d836ec08dc8816d6a14fca25cebb4022c41e4d6327089b

Observation 456f2f7e-ae95-44e5-9e90-486f0853b895 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.456515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.456515Z digest=sha256:1a73655b273386c3fe5f83f8ef402205a163f63cd4345b241128f33e40d6acaa

Observation 3dd41f15-24e5-4f99-a39f-ff294db464cd · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.018864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:7f384c1a76f453ccab7d93f3f12d91d56f7e2b4acf29d220debc268cb46fb501

Observation d6916505-eaec-47c5-8552-e8d1b831704b · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.011939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T20:55:10.819298Z digest=sha256:90b9a4809ff88aa923a54eebf1e8d1a3b95202fe3afa47b917b6ed8bb0b2f19a

Observation 52a3c39b-62df-47b7-9026-a1735bc6fb0e · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T21:29:34.747408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:29:34.747408Z digest=sha256:c2cd3917c74974289344739783ec4708069e2ce3568ed6bde9d52e7655243da0

Observation 59479678-5c43-4a1e-b726-d3dfcd61f296 · inbound

DynaWeb: Model-Based Reinforcement Learning of Web Agents cites this paper.

DynaWeb: Model-Based Reinforcement Learning of Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:37:41.942922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T09:33:30.444057Z digest=sha256:179d85cd5fdeb3d90ccbec9d90b7b6defc02928acd91d9426f96602e1ce95bfe

Observation 577a560a-e36c-4b7c-91b0-9fe6805dd2df · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:43.534953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:43.534953Z digest=sha256:553acf5884e8e97ce01f0afe8eb89c2867a5ad1b9fcf2888dde8adf333adad6f

Observation 8b693a41-6a3b-4de6-9082-f74d0f1123f6 · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.237069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:c9bce7ac32bb7bb6378cf7ef8072907efc5c82b185edd47a2fe38ece7463ed99

Observation 1b0603ab-9dfe-4f64-bc5d-bb4f5a41243e · inbound

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning cites this paper.

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:56.367208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:45:08.769264Z digest=sha256:44ff3f5f83825cdf4355633f6db8b99baf63900a7d2738d0516b2ed8507f04c5

Observation 79d01310-4cf2-4dc6-a202-fc8fec2f88f1 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.877886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:4dcd58083b06c3c9c87bfcfa06cc91f9ef37f1bd6579a27f1c4dc3f83d4ecc1a

Observation 666a29e3-17d5-4c68-ad9e-ed837f4bdb64 · inbound

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning cites this paper.

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:24.706366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T19:21:30.956682Z digest=sha256:e77b14df5a087b211b755676d0fd73c7d7f6a175521127189d3219cd13b47c72

Observation f2208569-4677-4a3c-bb48-212c0babdade · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.349693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:8720cc2802ad5e7146aa8f8f4ee5db99d4955eea67f03cef5f12b54ad07de580

Observation e0674fe6-1e47-49fd-85ed-65a97b37eda1 · inbound

Milestone-Guided Policy Learning for Long-Horizon Language Agents cites this paper.

Milestone-Guided Policy Learning for Long-Horizon Language Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:51:11.409606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T10:50:40.760938Z digest=sha256:4d069886f32d88ab1bb3af149c13dd69028734e170e881cec6769e96da3f9e13

Observation 058e9db2-225c-470b-902c-fa57a3760c99 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.888142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:54316eeb89dc153e18b17c3f229c7f45c2298a0d20e8c67957c3b851e9eeeba3

Observation 11ce60e9-4099-45ee-8ddb-43cd23bb6da5 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.542352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:0d16f80f562fe2550fcfaf5dfc65cd59bbdff8e2cd8b39e3569142db4f4bc0ed

Observation 61a65632-bb1d-4a68-a58d-f995578abf69 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.149642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.149642Z digest=sha256:a3d517b0059ca3412c08b341db999bee54894d785296ba58fb1b6b099e439aef

Observation 0ed81619-69fc-4fd7-9a72-49183809b3a4 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:26.500402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:61ceb74b8799282709fc4778419ec172dddaaa3338f00ba15731650e5500992a

Observation bd028f99-e4e3-48ac-b386-d6b57c581225 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.620192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:34980e620a77e82a6f50f2de75472355ddfd36407ce2b6560e295e603eb73864

Observation b069b3c8-cb86-418f-94b9-e3117a4981e0 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.549697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:a7b2a4dcbbf81630cd0bb2bdd993eed72171a3c4940daa26206509bc8b72c32a

Observation ed399df6-0832-4121-919d-a59193c3a5bc · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:09:51.308730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-21T08:07:59.093519Z digest=sha256:33c43cea237e28c18d31d0694e6f8d6492ccda7f133f693abbfbe1146defd82d

Observation c62bc5f8-cedb-4ae3-a1b4-4bcd212971a4 · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:15:01.478092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T18:39:12.161639Z digest=sha256:5fb84e176d51a79eefbc1ece29ceea94620a673497e5b85f8a41cd813c2ad3d7

Observation 6e5bcc9a-5671-4435-ad9f-e72c017169f0 · inbound

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate cites this paper.

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:29:34.490134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T04:27:25.041652Z digest=sha256:6d8d75afb397bbf99ef3ec1dff191376de2539b4651023969c4fc575f71d89a5

Observation 6557edbf-7468-470d-a8f4-45441a6db87a · inbound

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning cites this paper.

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:36.639556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T09:01:22.412685Z digest=sha256:35a52d9b852e3765724f90b3138665a92d2565671be1ac2e38a09de44e74027d

Observation 3d93c6be-e952-4002-9d0c-81f60346f689 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.605741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:2928dd0c767ae2768ac30b87e1548616c5bbbc0b06f58e0379e12a3a813d3209

Observation 9821d9a6-fad3-4048-a98a-743e8d7f257d · inbound

Deep Research as Rubric for Reinforcement Learning cites this paper.

Deep Research as Rubric for Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.262998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T17:16:16.243967Z digest=sha256:0974622137cf3367c062ba3ee2afcc8c30537b7f84d8edb30c7b61df1e093797

Observation abd50648-c9f2-4285-9a0d-9ff79f69f7f6 · inbound

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning cites this paper.

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:31.170792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T16:22:05.424293Z digest=sha256:646e939c299c9209db0f2eaf9205b31a04ab10ac731ac5a6de8b6b25e35acfb1

Observation 9bcde929-0217-494a-ba96-38d72605a508 · inbound

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation cites this paper.

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.241228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T10:53:27.839408Z digest=sha256:e8c4ae30d7503ccf7cc90d88d097bbb7bc2ad5c629d704016c47d16bd8257743

Observation 0ce9644d-fea6-4905-a03a-468dc6aaa1ec · inbound

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents cites this paper.

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:59:38.183304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T13:56:51.914966Z digest=sha256:8f923de5104d0a31129ce2f23527b743ae2c88ff6931c6176a4010fe3506914a

Observation 8abca0e9-e8ab-472b-b5fc-d427f2488334 · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.042865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:8f5aa1e5d0e8d8d9dde5f26c1ba25ff7de37100a1b82c18b2d7088b26ffac831

Observation 18440b04-b9a5-4dea-b32d-f6f4e07bcc55 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:7023e5e4a57c9d655f78c38844880c808e126162569255068814a1b06aecb316

Observation 3ea66aef-dc4e-4898-8121-a7bf36a89eb6 · inbound

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL cites this paper.

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T06:25:23.264527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:25:23.264527Z digest=sha256:b54482c415931ade719b691ba9c2e92476f0c37b2b9b2a6410e68c8883150d7f

Observation 258eb65e-660f-489a-ac9d-68331475c65b · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.764481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.764481Z digest=sha256:8f8948fa38fb8a244198f024f742207b52445c73cc284e5f71d582642fe1df24

Observation 2ea62fda-50e5-4f86-b47c-1711cdedaf18 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.962904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.962904Z digest=sha256:d3b40c52ee85eba13f83e50c57b9b7468043b3ce7651350d445d851cf4607da1

Observation 919ded30-bacf-49f0-9ce3-940a28ee6b4a · inbound

Software Engineering for and with GUI Agent cites this paper.

Software Engineering for and with GUI Agent WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-11T20:19:15.589837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:19:15.589837Z digest=sha256:8e9fa6ee8db5f7d0cc4f92cd1e8d6a86ba55fe6981cd8d43bb7194f4857617e5