Pith. sign in

Paper Citation Record · LEDGER

WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2411.02337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.02337 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:36.060344Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:20:00.041093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a439e27-a856-4426-b5ad-49f735016f73 · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:32:18.628495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:08976c2014afedabd97b4d7c3144af486eac921c876c6ccaae1778f27d077f8e

Observation 8a08d509-760c-4940-bbdb-8350116a3a5f · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.751909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:1f7470dff2ed69fc52d9c5e9aa94254a105924954f88bb70f767c9efae1d7dec

Observation dcba1132-d734-494e-be0a-fc5a787f1ed5 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060344Z digest=sha256:510a5d3fdbf283920adf6a84e95644640bb14bfc7b60418b7b7c1a1d5bc90a6a

Observation 59ccd5d8-9b02-48c3-8bd2-ba65dad8f4ba · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.716794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:936f59fd588ea241aead49a32f223fcfa876ea919680f43a285bbe599eef4df5

Observation 7498d1e4-28c7-4c4e-bd30-e41f5108b9a1 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:21.810260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:21.810260Z digest=sha256:f4d2f07e1672fe31d887b05e8770319e4772761b26856d8fa9902d79abd20480

Observation 2e782046-7782-462f-9b1e-06e866525791 · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.978039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.978039Z digest=sha256:604ff0f3c0552b479c1848d5110c8baa9eec6cbf35365bc17b09476b19d1fe4e

Observation 23b686ae-f595-477c-a155-9ea064784556 · inbound

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System cites this paper.

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:04.613379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:04.613379Z digest=sha256:baf489870333dd0625ec38867d99945b0b1728b39286d87fdb4e7bf8895e8348

Observation 584b1166-6f2e-4fa4-b163-63d343223acf · inbound

Build the web for agents, not agents for the web cites this paper.

Build the web for agents, not agents for the web WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:00.763378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:17:00.763378Z digest=sha256:854d20933ad665b2fd1462c4f38b017cada71b3f3b29c3eb95d2129eb507a91c

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · inbound

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards cites this paper.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:a10ce0d43e276622a5a16b34d7ae64ab8b6d0b5db7397df1e244700d7be7f633

Observation eea4fae6-f494-4e05-9c8d-05c761871989 · inbound

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents cites this paper.

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:27:37.434374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:27:37.360221Z digest=sha256:65865799ac0f10ab095f083a23c318f04815fd3ac460c18e3e34cd2e434d9b70

Observation 4a1f1581-f669-4224-9f39-26e04b903bbc · inbound

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis cites this paper.

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:01.164687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:01.164687Z digest=sha256:434bc07691467aacc7e755d9be3d5a1e4e7d66909ab92f52c5fa55a9b934f4da

Observation d5f4170c-1008-4f8d-881f-50d4d395223b · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:286f3f818106d58a13e8d020547d343e8bf46986f8a218380d07ef02c0db3dd5

Observation 56fc05ce-ba75-47ee-88b3-2906e94b33b0 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:50.893647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:50.893647Z digest=sha256:a75002c1338cd97a9062223b73a0c01060b934189ee2bbeba0f8a399d664de71

Observation 3420f575-8354-49a0-a0ac-002bd0281248 · inbound

Cognitive Duality for Adaptive Web Agents cites this paper.

Cognitive Duality for Adaptive Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:36:21.986350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:36:21.986350Z digest=sha256:e85be1e15dd2a7a5b9dedf6fe367876f87f8da688a931f8dc0260e80dd769e4c

Observation 8ebe02db-d3f0-423e-8c0f-2cf95809e21e · inbound

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making cites this paper.

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:24.747926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:24.747926Z digest=sha256:606544516c798dcda7f45117c6ac5fd02ea83e4331ee4ea83a96436337c3eda0

Observation a32608b8-1845-430f-b8b0-6d33fa8819dd · inbound

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward cites this paper.

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:21:50.831536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:21:50.831536Z digest=sha256:3eab9f2e2d6f8aaf97cdda886247f1423e2be21b5efeb1612cfd77ff3c293724

Observation ceb3ec28-6bb5-463b-97f4-03a7fcc6ba6e · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.759166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.759166Z digest=sha256:a22fb16c42118046477f463274a212778a31bf4befe53abd36f252926c87fb65

Observation 6841f209-0436-4463-aefc-e54ffc9774f9 · inbound

Symbolic Graphics Programming with Large Language Models cites this paper.

Symbolic Graphics Programming with Large Language Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:10.896680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:34:10.896680Z digest=sha256:7772fb1fe0e80320f202a2b55b2bc6d8316308d959854c546ba7732fc4168f7a

Observation 404606f8-a388-4e91-bd77-0ac751fe4792 · inbound

A global log for medical AI cites this paper.

A global log for medical AI WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:15.570579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:15.570579Z digest=sha256:78613bb966e32fa49aba71437a7776750db47c0470926b2fd4e2457ce9e07076

Observation 456f2f7e-ae95-44e5-9e90-486f0853b895 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.456515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.456515Z digest=sha256:bf0fb8e81b2afdd33f11d102b50019eb5f0eec11c9ac11b16512393392552f5f

Observation 3dd41f15-24e5-4f99-a39f-ff294db464cd · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.018864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:41505ca6f6b29df23b3da9e13aa65d7ac51d97798ec322b7668907170b03a081

Observation d6916505-eaec-47c5-8552-e8d1b831704b · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.011939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:55:10.819298Z digest=sha256:73bbc7946c80c172de217af591371d3941f5579bdda99036a3be299e6d9f7081

Observation 52a3c39b-62df-47b7-9026-a1735bc6fb0e · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T21:29:34.747408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:29:34.747408Z digest=sha256:b0674aee1cfb25e51e6a8bbca21db02014237dc152b376d211143f23d544eb14

Observation 59479678-5c43-4a1e-b726-d3dfcd61f296 · inbound

DynaWeb: Model-Based Reinforcement Learning of Web Agents cites this paper.

DynaWeb: Model-Based Reinforcement Learning of Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:37:41.942922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:33:30.444057Z digest=sha256:4444bc6655e09fb44a967c5d2e9db47067c50700e0dd213812b9e8d1a62add76

Observation 577a560a-e36c-4b7c-91b0-9fe6805dd2df · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:43.534953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:43.534953Z digest=sha256:db533635a3d0d7dcf9a73f14a3911d64990ec4504bef7f6960c558af926e6308

Observation 8b693a41-6a3b-4de6-9082-f74d0f1123f6 · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.237069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:335c9a271092b3ab88a10569d0eaf1e6a41844f6b7388858e74950c04eea26c1

Observation 1b0603ab-9dfe-4f64-bc5d-bb4f5a41243e · inbound

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning cites this paper.

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:56.367208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:45:08.769264Z digest=sha256:2834da36b4951e69f6cb6f5024e91017144358beff70d5fad5e0d49ca47dd258

Observation 79d01310-4cf2-4dc6-a202-fc8fec2f88f1 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.877886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:70661b8e05e371d641b7dd5dac7e7b75590ba0a5301dfe0126b6893a82b4cb1d

Observation 666a29e3-17d5-4c68-ad9e-ed837f4bdb64 · inbound

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning cites this paper.

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:24.706366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T19:21:30.956682Z digest=sha256:4fa4f35c7ea4cdc121aaf35aa252b3b21972409871bec4154a3909584c74608e

Observation f2208569-4677-4a3c-bb48-212c0babdade · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.349693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:c5052218723d68a280d84e1b562ec69b87f5c195b79d037b3ff87d890d97fe03

Observation e0674fe6-1e47-49fd-85ed-65a97b37eda1 · inbound

Milestone-Guided Policy Learning for Long-Horizon Language Agents cites this paper.

Milestone-Guided Policy Learning for Long-Horizon Language Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:51:11.409606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T10:50:40.760938Z digest=sha256:4f4115e5d8f818b11880902b8096a364030000d883db28e745701be0e0698809

Observation 058e9db2-225c-470b-902c-fa57a3760c99 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.888142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:31ca0f64084ccbecc16f27406d857c40290771908f5497895502155ed9857964

Observation 11ce60e9-4099-45ee-8ddb-43cd23bb6da5 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.542352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:bc6558c7c39a4520daeded0da2b04b39c590f8ae75af10c38b390dd8bd069251

Observation 61a65632-bb1d-4a68-a58d-f995578abf69 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.149642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.149642Z digest=sha256:f389edb81fca69329b1fba0ba126a1751625343b3ad8a20e42cfadc580f6e786

Observation 0ed81619-69fc-4fd7-9a72-49183809b3a4 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:26.500402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:04d0becf65a1264e80b63a3f216303e03a17adbe2cf1fa584655424c1886ee68

Observation bd028f99-e4e3-48ac-b386-d6b57c581225 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.620192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:691ee3a7f179c0777dc53848c5f3cc7eaf2333c4a8b1288c1fc097bdbd1356b7

Observation b069b3c8-cb86-418f-94b9-e3117a4981e0 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.549697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:4c48c4f8f852aeeca344fa1a60b25dd138348fe2e786cc18cfa11e15bcd35cc8

Observation ed399df6-0832-4121-919d-a59193c3a5bc · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:09:51.308730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T08:07:59.093519Z digest=sha256:5ef8c52dea5db7150ba27a9480184e5284f9066892f4c1e51652b000b905b413

Observation c62bc5f8-cedb-4ae3-a1b4-4bcd212971a4 · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:15:01.478092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:39:12.161639Z digest=sha256:0eea14bfa993ee5db7c7af241fd826590818a255c4d0ab235c6d0b829e206219

Observation 6e5bcc9a-5671-4435-ad9f-e72c017169f0 · inbound

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate cites this paper.

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:29:34.490134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:27:25.041652Z digest=sha256:ba3e6bf5f1759082de90364922002e9243aa7187a9b59d4c2ece4dfa98ea5327

Observation 6557edbf-7468-470d-a8f4-45441a6db87a · inbound

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning cites this paper.

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:36.639556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T09:01:22.412685Z digest=sha256:10e1e4594008d6ac94fba3b68ddeaa08ed3263ce4705e9d9d764dbc28e880ba0

Observation 3d93c6be-e952-4002-9d0c-81f60346f689 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.605741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:627de6eccdd0b2864488181c2aa55d9746c1b9e75c1161b67640dd60b0282f8c

Observation 9821d9a6-fad3-4048-a98a-743e8d7f257d · inbound

Deep Research as Rubric for Reinforcement Learning cites this paper.

Deep Research as Rubric for Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.262998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:16:16.243967Z digest=sha256:f1af757d4acb12429fb6b296acd4317affcbce39e5bc66661cfaada6d97da353

Observation abd50648-c9f2-4285-9a0d-9ff79f69f7f6 · inbound

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning cites this paper.

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:31.170792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:22:05.424293Z digest=sha256:6c9997d913fd7a47c15d7602a736c1e3beec414a2cd6b55fa2470df22d52d46a

Observation 9bcde929-0217-494a-ba96-38d72605a508 · inbound

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation cites this paper.

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.241228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:53:27.839408Z digest=sha256:b4e4e89209bdbbbb19e1c27a93133ed22823fb0a9a874b7d55c8c0978d18d67f

Observation 0ce9644d-fea6-4905-a03a-468dc6aaa1ec · inbound

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents cites this paper.

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:59:38.183304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T13:56:51.914966Z digest=sha256:b6394fdd9274ac55ba737ebffad0d65ef4457ae481cf237a2a8577a0bbf94cd6

Observation 8abca0e9-e8ab-472b-b5fc-d427f2488334 · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.042865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:6542ad49fe9bcc7d11c7b1f28b9861558aee97ef0eb185653b49cdfadea8a9c1

Observation 18440b04-b9a5-4dea-b32d-f6f4e07bcc55 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:1e0b853884a9832c9c055e50b538e6b8b87fc4ab37c1734ad689e8c3d234a57f

Observation 3ea66aef-dc4e-4898-8121-a7bf36a89eb6 · inbound

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL cites this paper.

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T06:25:23.264527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:25:23.264527Z digest=sha256:dd422eeeb6ca6b71f75da8262eca73f968d3f70d837417451e9be4a0f9cd4428

Observation 258eb65e-660f-489a-ac9d-68331475c65b · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.764481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.764481Z digest=sha256:7d2872b141ffe3ba6572dd073b73f9984db0a9c63bd956be9b4d4a85b560a6b8

Observation 2ea62fda-50e5-4f86-b47c-1711cdedaf18 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.962904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.962904Z digest=sha256:0d156de7db4bfd7d9db6494586445ea0eb1e70c1a70573ee189993db41ddc3c3