Pith. sign in

Paper Citation Record · LEDGER

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2606.17680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17680 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:26:41.376368Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:39.533735Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2da6deae-2830-43ff-8545-85e5e7a432bc · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:29f9d5389123f5331bd27c4568d0b376e7b2156b46c5d8976da062c8fac2bb60

Observation 711195b7-a11c-44c1-bb2c-845039f4a1ea · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:b48b7886c01fb4522ec3700035a92210bb8cc6625f01676b521319d4f6f905d8

Observation ceeae14e-c6c3-4cc3-a04a-732d383a5ed7 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a59388902fc1fe7db37277b6f6660e5bd7a4022bff3e239d9f1e15d10c351ba0

Observation 4db8a43b-9664-4e21-8ae1-8aff5f08d178 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.189097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:30399856cea8268526f3ffd98c9dfdf78dd66630c4146e9cbc04939f513f1133

Observation 8f3db74c-dcc0-4699-a7f3-80e80c2dd5c4 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.192629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:6c3ad7ca728932fcbd56ecf40aec45a9c86da6d8b52d3d9ffcf8bdeab9e6b183

Observation cd3fc0bc-4930-4247-9293-f421fa75a367 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:1f1655846874fd71a51888a696c9491476d5b3dd831ae3c1b681b34b6969c501

Observation 9d53c325-dac9-4724-a77c-ae1b4817b3a6 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:4d54be4462cfda1478e974158d3a8c60e42adf761da215d559c783393b71458d

Observation b54486c1-4832-45ad-b041-210e05db7849 · outbound

This paper cites Mastering Diverse Domains through World Models.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Mastering Diverse Domains through World Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.172700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:b21ff69008930e341e68e744183c3fa4d03c34a739864c5daccdd2e58b152dea

Observation f5bbdc4f-a611-4753-a00b-63bf3205bc8f · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:1329e3a6e204e674c3ed2f6f8f72eac278149f2f49d17937aa3401bb17394693

Observation a4afef43-0771-4834-92b4-227433559d24 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:f5f7653a3b5a2eda9305800fc7736a87d0e0d4df72baa675db55cb596c4b5531

Observation 90130480-67b6-461d-b980-9fc37bc8b88e · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.194544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:8ae8ec574fb8eeea4440ce32044d9597e0a71178b89c2ec340df2d86ffef201e

Observation db926258-f0e9-4868-a4cd-ecc180ae1e06 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:648fe7a02125b4bda1af0f6ea05d0dfadd06298f0dd57ed4a81772c5e9efcdf4

Observation a1ec3c90-97bc-4be4-b40d-bcf0f832ea4e · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.197623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:ef5b877682d95e2045105f4a00db1a5638edc16e57290c17c5f8f28459cb29dd

Observation cbcb43c0-16b7-4365-9548-dcf06d51e6fd · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:58d54680e7e8d1ee39e527dd673004d7dd5c4d9d5562955e6d25f6630d19bc00

Observation 595fd607-6f87-4cd5-a646-60dce89d8a72 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:fcd5c6bc0fc314b9b0590fb2a7daa12078db1d21a828e8582a41105b21ff7b49

Observation a8c97d08-13e7-4678-a77e-d1a5951717dc · outbound

This paper cites arXiv preprint arXiv:2510.20022 , year=.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning arXiv preprint arXiv:2510.20022 , year=

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.171712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a0e8ccdc5c00ec15f2a3f33eff02cc4c1ec52596f2bfc34de3cf9949aa829de6

Observation 83160f63-1a68-4e6e-b3d0-cce42018b71b · outbound

This paper cites Continuous control with deep reinforcement learning.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Continuous control with deep reinforcement learning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:57.167368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a512dab774d0e2748d1b91b364dcb48d8bb3753cad9e2dca9fb03994015ac5c1

Observation 191ff0ac-4af4-44fa-8b1c-8d5f768ee3ac · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:f77a7e3b66c6e917d51a35f1e98bce45ee505bd29a98aa8cdcd0dd9e35466aab

Observation 887d65b6-dfcd-4be6-8423-ff7a7b97d16d · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a0364cc8162171215d8d2d0cdbbaef48220caf988b8a534f3272b401d29ca8a6

Observation 385ab561-0bf8-48ca-a1d1-e64d452494d5 · outbound

This paper cites Agentic reinforcement learning with implicit step rewards.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Agentic reinforcement learning with implicit step rewards

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.185413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a543b957c4fa94223f8c51cfe1a3321e576bdbb15f08073eef081c763a72a5aa

Observation 14d13fc7-2e9b-413f-916c-b255d52a9689 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.160536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:cc112f6c06d98f9206da5720250511281a757c33f9d79240cf73a9e11cd60ffa

Observation f283446d-52e7-4bd4-91ed-b9d299eb7822 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:084051698d49f6a43b5aededc5fd7504324a543b52d96abc8bd160a9271311f8

Observation 24cfe1d1-9f61-467f-863f-3b75235acf59 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:6b2c36fc8595b4c50b19fd5495c8956b484f9d7c5975b920fbcf59e58799ce40

Observation 542bc27a-a086-4963-a80a-3b32caa1ebf0 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.163791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:c9f3e6ee1f722cc41097bb30eb7972f1fe7b29582f79491fcfde4fcdde2035a0

Observation 83621a44-098d-4162-bbfe-63102ddac360 · outbound

This paper cites GPT-4o System Card.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning GPT-4o System Card

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.175228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:8836dafeb0ff9c344a8ca0ac275e22f2bc0abd3f9c4bd67cbb82c6133315b692

Observation 775deffa-c14b-4605-a73c-18cf6bbac1b8 · outbound

This paper cites Efros, and Trevor Darrell.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Efros, and Trevor Darrell

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:27b73ab64ed3dde40631f62cb7cb3f168e55d113a2d4d14d266a6ed43de0032e

Observation 34491557-e07a-4181-af45-8d9285cd6d57 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:3e20c92252c9c90dfbc42b5326724268cc4acc0dd6554453935c74b42c252a41

Observation a05e988d-b170-423f-b555-0c057fdecf97 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.178975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:6671c5a220030baf802844668f615cfb6d9e8d6374018b3efcabbbb4cb55c0bb

Observation bd8aef76-be5d-4d49-b438-7c9de81322a2 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:5542cd2b37fa3f2d422e1bda614b9c5edfb9dd0d5111eb22d350f63ccbea7f58

Observation 4c868402-7811-45cf-a333-0a0be3b46d78 · outbound

This paper cites Proximal Policy Optimization Algorithms.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.159901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:0f20334c6a4e35b7ae6317a96b48b8c7201d68053d58721241b373bde7e81804

Observation 41932e52-96f7-481d-b91d-c7e385eca11c · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:f79f1814bcae63226a092756c5b7f12260c634f490915d28e8740b9a686759b8

Observation 4ffdb9aa-08bb-4d10-8a99-ac3330400095 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:17e112dcf5b51acb02ead9927e2b2701cd7ed0b8ccd3666e3dede0cd3f69fb0f

Observation 6da4d6bf-1f95-46cd-a66b-ecb7935d6eff · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:8bc52b953b91afc6e23f6dc077f3fcd860b6e9524425bf9ac84329ded17c3cbe

Observation 22de5ad0-c187-4b2b-aaca-c2c180df2dfc · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:57.195470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:e217d684480df114c405ff86b29d25befde8be415a192f60472d4c0b3c41970c

Observation 3f465248-216c-4a15-b5cc-74de13c56f9b · outbound

This paper cites OpenAI GPT-5 System Card.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning OpenAI GPT-5 System Card

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.139137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:34afc32b24a1ad248e9cf9971b528e2ea421e4bbec705ea1f10d4a2be22a1b84

Observation 981d3489-4f5f-4eff-b75d-7360eaecfa08 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.133814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:f4fe6c81a7bd29d8d0da74eb4f43069a91ff1eeddfbbda88780d1fdecf0e5321

Observation 62b62c0c-0c8c-46ff-8159-57f9155dfd16 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:7204dd9d2bbc2793736cf0f93781c04e224130cdd0faebaf3d3e2c90075f3413

Observation 150e4003-1bf5-4821-b923-a0af688f916d · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:0f9a0d85801dc2b4f67948ec8d23b17dd7ac421d1349a1ca9ab464b9caa5d326

Observation 2283b5a6-f4b1-4ea5-9d9f-51198fbb4209 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:061b7fabcf54e37260568fd1c1ffd824b443adec20122a6ed3d8be973a78bbde

Observation 0b739ea2-8382-458b-a786-fadaf911452e · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.137468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:406d153d000cdda3e97214ca5c407a104ea679c9cd9579588db09fdbfa20fcf0

Observation 9fba41a5-172a-477f-9198-c5c2054a91e0 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.201947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:64a6a34e07ea159b760e18318ea5ebaae4c7cbbcc09d27e81416f0953a1e3194

Observation 96b797f5-5b92-40a8-855e-4cc321373a95 · outbound

This paper cites Qwen2.5 Technical Report.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Qwen2.5 Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.184283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:eb89e4fb30ecf95b4492eaf9dd3410f3bbc211e90f0c70081dc8642fe07205fd

Observation b3f51b9f-355f-4937-99c8-115368872d69 · outbound

This paper cites Cohen, Ruslan Salakhutdi- nov, and Christopher D.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Cohen, Ruslan Salakhutdi- nov, and Christopher D

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:74d06c606cd8b7488c3cf5438535c32226ab7d379eef0d96443acac637b0665b

Observation d4640620-2a25-4de7-9308-020bea5490f4 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:a07c187e87494ba616d689cb1b67151f54af3b5736e052d28fd55c3b904a77f5

Observation 0fdf2081-60a4-45a1-a430-c37636a34c9a · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:5ffa61c5515a6cc38d324caad0422dcb4c2a2fb02a4aa64c26b6c0d2507e4b84

Observation d078592e-50f5-4632-80e3-02b3fe68a138 · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:46b3d7952429dda878f30ad9a520928d124440de8ed1eeb0d39d1050c347d735

Observation 769e2d20-5573-43b1-bd61-128d919e409b · outbound

This paper cites an unresolved cited work.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T01:26:41.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:0567773ff2077c8397522d9b71b956892a0806f501eeeb29366824bc7da6729f

Observation 21a5a225-3d36-443e-9f07-583c5bb9ae1d · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.198518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:e49c590d00d2ec02f2c0b49319c99cc7244c9e05b6e37756244d8ec6928c98c5

Pith citing papers

Observation c431ebe0-6e2d-44c0-80b9-6bc9fc282a77 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:13:39.900693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:13:39.900693Z digest=sha256:63d9da0324e9829e3c99dff9aa3fb7fdaee9a98e71369e870aaebcc76ec53eb6

Observation cffaf9cb-68a8-49bc-8848-b2a9d3007c82 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:39.533735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:39.533735Z digest=sha256:52655fb9a107b53eefecc8c28c7c42b5d706f7dd2f6f7ce2c1b4216249e602f9