Pith. sign in

Paper Citation Record · LEDGER

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

As of 9 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 4 inbound Pith citation observations for arXiv:2506.04287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04287 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:03:30.907605Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:37:55.374135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.584868Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ab9985b-b96a-4df7-9e36-d89cac920609 · outbound

This paper cites Variational Option Discovery Algorithms.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Variational Option Discovery Algorithms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:25.822971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:25.822971Z digest=sha256:477e357446b9ed8cd02d6ddf38ac4c70d321da458fdbdf1c305c658ba84aad26

Observation 1ac41ca7-a31c-4ac2-885e-2f056241fcbe · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:25.898272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:25.898272Z digest=sha256:29ec3d28cc7bee8fbfe3cb4c2f1601c0983f722548d0cebe8eec119b38f21b74

Observation 6e17bb1d-f86c-4a73-ba9a-116714fb9ae1 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Mind2web: Towards a generalist agent for the web

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:35.056483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:25.999870Z digest=sha256:8061fc3912c3151364d998ac9f4bdb73974943425f54d69282ca695d7e7c2197

Observation 6a38fe2d-1f29-4a32-828a-4b9469d7eae5 · outbound

This paper cites Emergent complexity and zero-shot transfer via unsupervised environment design.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Emergent complexity and zero-shot transfer via unsupervised environment design

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.896987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:26.123499Z digest=sha256:4a0352a494cfe004cf8a307778ddaecdc1833fcb269c14d0a972c506835dbc21

Observation 18c5c05c-2996-4ed9-8338-803627b57746 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Guiding pretraining in reinforcement learning with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.718439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:26.199768Z digest=sha256:8187a9424141f3389e8e6d10f079194f8f50b148b06817ebddd8f49d70f59f1c

Observation 0b8943a3-2a20-48fb-b35c-63242f273e51 · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Diversity is All You Need: Learning Skills without a Reward Function

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.274166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.274166Z digest=sha256:de6b929c3a4a4d66752cb91bb3dfa951972e86c76b4acaccd63b4c2b2cfe2d85

Observation b6583d0c-6308-46fd-9fb3-6eeb0cda2849 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.369118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.369118Z digest=sha256:956084c6764af6c22f7dfefe121a4a68ef12314281ed94c171a7c7371b23ae15

Observation 42f0b07b-99a7-4878-840c-4c883bf3c048 · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Automatic goal generation for reinforcement learning agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.513776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:26.431916Z digest=sha256:d9bf0e91209ad7e389bec90cbc45a6ade49fd7dd33898f51733e65826ec14dc6

Observation c2e78acc-a2dc-4968-a635-029dd4acd4aa · outbound

This paper cites The Llama 3 Herd of Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.535427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.535427Z digest=sha256:431008e237d8d76a37478f2fd82eb62f57dc209a93aa13a94b2b66995d8c4070

Observation ea233952-099f-4a2e-b0b9-190f77d862e9 · outbound

This paper cites Benchmarking the spectrum of agent capabilities.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Benchmarking the spectrum of agent capabilities

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.272353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:26.599008Z digest=sha256:f2c306715ccfbd63676cd168a1ba83ebc094a2d2ae7b9cc2ed14e04019d87868

Observation 4f578e32-4a3f-40cb-b6de-fbc3bf95af79 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.700075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.700075Z digest=sha256:bede6b540dcec6e0a0e226abde0ef91e65f23a64be4cc0bc5c99ce189ba6592f

Observation 66d50625-b891-4aed-ace9-db0823aae1c2 · outbound

This paper cites OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.768525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.768525Z digest=sha256:a4a852c65e47d2554fb7c28e84a9ed23c7db7f89f6e41ffd74fcd0db3f41799e

Observation a82b0163-68eb-4292-87b9-5136b0fe8989 · outbound

This paper cites PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.888965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.888965Z digest=sha256:5c77aa53c54a2b97c9da44abb19683473c17b4de6e8f0e2c394824ad2b8eb800

Observation 67839074-5a44-45b2-a04e-f08c049c142c · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.976890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.976890Z digest=sha256:07301f551c67e4ee3db40f2c52b987ed1a6a248033f2f1c2bb102900c93a9c05

Observation 3b905b5b-0efa-48ac-88aa-190aec4450e2 · outbound

This paper cites GPT-4o System Card.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.078439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.078439Z digest=sha256:b3963bb2f493233efbf54ed9714500b289e1bb812838f75c3127c8c18525315c

Observation 9c6c7823-92cd-41e4-a006-516e73cff31b · outbound

This paper cites Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.155990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.155990Z digest=sha256:0023d683dd2b5d6051c1421f766e47983d26f0d9872d29b75df6ed0ddd5a61c9

Observation 337970c5-63e6-4fb6-89ed-9a2f5323c6db · outbound

This paper cites DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.229732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.229732Z digest=sha256:d85528bead4db9ebf861c8a71d3cbbd780a560195f079cbf33d4cadcaab235b1

Observation 4ceb284d-727c-4bb7-b356-c074b76ae856 · outbound

This paper cites Tree search for language model agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Tree search for language model agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.308410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.308410Z digest=sha256:543db7c9294bafda2e977164f0ba151dd7b865db0cb7fe6b9796883ae6f03f8a

Observation 79a56932-f8fc-4c29-8c8f-c2dd52e78ab1 · outbound

This paper cites Autowebglm: A large language model- based web navigating agent.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Autowebglm: A large language model- based web navigating agent

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.068000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:27.379252Z digest=sha256:75fdff5b1f8283bedaf32c40aa05bd35f8c57e2758a44c9adce9220d062b5485

Observation ecec396b-6a7e-4d14-850f-c1ebb1a1f658 · outbound

This paper cites Unsupervised reinforcement learning with contrastive intrinsic control.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Unsupervised reinforcement learning with contrastive intrinsic control

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.906592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:27.500008Z digest=sha256:e608a32f56662d08a39783e49b5400fe35cd1e5e5e69ccc809acb7935b6504f1

Observation a9aab2d3-e448-45a5-866e-ada7183d5c40 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.589628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.589628Z digest=sha256:9a99203073b9422b142da96ccf900ed0e99713720957f863fcdc30c23d6372c9

Observation 8522c990-18b0-4077-a5cb-feac09909bac · outbound

This paper cites Competitive Experience Replay.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Competitive Experience Replay

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.665012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.665012Z digest=sha256:286e85f05f5c25627a3d5f262124565754ba489af53a5aa22d085e881702098a

Observation e0c742d5-f210-472c-baa9-ba357b71a29d · outbound

This paper cites Weblinx: Real-world website navigation with multi-turn dialogue.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Weblinx: Real-world website navigation with multi-turn dialogue

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.742181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.742181Z digest=sha256:d4618c6e3849393fbf4ed24c4a7cc80b5991a65a350676fbc640a10059b6d7f6

Observation 05674ba6-ef51-4d28-a03c-bbc3aa619128 · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.809974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.809974Z digest=sha256:1eb0bd843a7eb5a2d665ce592dc1f58ee31f6bf43fe4d22e75956ab6097f3a93

Observation e799eca2-72cd-497e-8185-e216d779606e · outbound

This paper cites Bagel: Bootstrapping agents by guiding exploration with language.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Bagel: Bootstrapping agents by guiding exploration with language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.732780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:27.860436Z digest=sha256:5e02b50bc197e957716da5c424271e8cd8c4a651fe0c27eb33db595f59835b37

Observation 488e211f-d605-477c-a681-2f58508d4a4b · outbound

This paper cites LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.944964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.944964Z digest=sha256:6d744a90f676cb16cf9c56006bfa457771ce6893abbdcf7ce957fe830b4f1766

Observation 634f9460-6bfc-403f-8bfb-fe6e1991ade3 · outbound

This paper cites Asymmetric self-play for automatic goal discovery in robotic manipulation.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Asymmetric self-play for automatic goal discovery in robotic manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.037262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.037262Z digest=sha256:e5a4eff7a44360045d29937208616c6d518b5086957edf811dd1306504797b63

Observation c36a30e3-5bda-43a7-b93a-6d06b1404b6e · outbound

This paper cites Balrog: Bench- marking agentic llm and vlm reasoning on games.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Balrog: Bench- marking agentic llm and vlm reasoning on games

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.576120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.126974Z digest=sha256:d0a2488acd7b4d09d65857bd80715237019d0106d3f044e0bccfa3e547c02ba9

Observation ebb98b00-a759-4575-8cd2-94a9444b4612 · outbound

This paper cites Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.263675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.263675Z digest=sha256:acc21422fc9a9034f104ba6cbcd5a36bcef82cab8ebf8ef9aa251463968087d5

Observation 8c74bece-053e-48a5-a5d5-fb27551fcbab · outbound

This paper cites Lipschitz- constrained unsupervised skill discovery.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Lipschitz- constrained unsupervised skill discovery

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.394607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.364199Z digest=sha256:9ba52c78a5a7a4dcf9387a3cef7d3d4706b0d47a5f0a212dd42699aed05e381e

Observation 704a9429-999e-460c-8b6e-f3e02914ac6d · outbound

This paper cites Accelerating reinforcement learning with learned skill priors.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Accelerating reinforcement learning with learned skill priors

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.216401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.477573Z digest=sha256:9d12d16a625d26beb0bd97b0fdb29fd5a7ef7f5366c0b5145f08812d6384d86e

Observation 017bcc25-aebd-4488-a748-ab051c631b87 · outbound

This paper cites Skew-Fit: State-Covering Self-Supervised Reinforcement Learning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Skew-Fit: State-Covering Self-Supervised Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.595829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.595829Z digest=sha256:046e3f2a7544cfb2c8f5fb1963868d4c533b673c0150fe37d4b41b52bddbdefb

Observation c5baef8c-f312-41be-855f-8bc431a1338f · outbound

This paper cites Androidworld: A dynamic benchmarking environment for autonomous agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Androidworld: A dynamic benchmarking environment for autonomous agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.020514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.696270Z digest=sha256:51d1d5dbc08065a287739a4ff8fcead3b0dce3d409d5734081c607e4377a0ec3

Observation 4eeba095-8a8a-41d2-9c4b-4e30f98584c3 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 2023.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.863729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.750473Z digest=sha256:1cad52ec312d3487bfe363d71187a70a286c6033c5c5194e00089d1255ad3d18

Observation 86ac7cf5-2243-411a-8a22-63e118e12f77 · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.855289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.855289Z digest=sha256:e89f0c49c767fc6ba97d4002ed70172bb769f32ee7cc71dbfe32adfd183280a1

Observation 78bac3d6-a90f-4b47-a82c-e9a1c09a807b · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.758345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:28.978049Z digest=sha256:5c5f95453ebc3ba129970e89f2e1b4f7bda08f67e7963dab2237c2d1707413be

Observation abd0e088-0109-4ed7-82e1-872d4053e928 · outbound

This paper cites Open-Ended Learning Leads to Generally Capable Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Open-Ended Learning Leads to Generally Capable Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.057904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.057904Z digest=sha256:cea8a31adae981a7ffb3492f1d410270cc4362d0f11a7ae7303984e41b04b40f

Observation 74b40594-2c33-493c-a6a8-b17576a88022 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.122001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.122001Z digest=sha256:056163736bf81fa1f4888d10f8e4327a4cd142034cde8ed7de79ab894a0cae89

Observation 322b08e6-0357-46b8-a7a1-4d2a4cac3a6a · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.229207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.229207Z digest=sha256:dd3ebbaec8d74ad366f2f744c41a672a2fd4b18fd44fba3e66dfdf33ff8dd36c

Observation 9db23e1b-adfa-4b28-a065-d4a68766d11d · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.338089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.338089Z digest=sha256:420611061e2ab95addfeb6475579fec184746cdfebfc15792011c0b0b836caa3

Observation 4062c452-38b2-4f84-85d2-4f95a5eb7cbf · outbound

This paper cites Qwen2.5 Technical Report.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Qwen2.5 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.407074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.407074Z digest=sha256:0b1636e9700548202c0552517597a97afb9432d79740bdde0462f343b93210e0

Observation 072215cf-aab4-4c81-a214-27e70e67ab8f · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.580264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:29.545300Z digest=sha256:c5649322805369b1df8602939d4e580db28e438d4500df08ba4089a92b0c630f

Observation ddf49a7f-7c9d-4cea-bcc5-3efc2af5cef5 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback React: Synergizing reasoning and acting in language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.626585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.626585Z digest=sha256:a8c07e90e629de70abddb1329a762a4f518815235e600f88a12b00e09b9b9aea

Observation 32b89ad3-01d7-4694-a7af-8cc8b92086b1 · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.691987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.691987Z digest=sha256:d15c7f9804e24173313649756f48d2c162421867595882737d8f51281555e0e9

Observation 5945f212-43e6-4604-aa8f-faf09b6d96d3 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.806310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.806310Z digest=sha256:7753402040b7d39abc83e43b918554fbb8930ec967a74f4deee773a4ce483896

Observation 9a23aa97-d5ea-426a-8e65-38556fbb4eef · outbound

This paper cites OMNI: Open-endedness via Models of human Notions of Interestingness.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OMNI: Open-endedness via Models of human Notions of Interestingness

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.887117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.887117Z digest=sha256:7c5541dec934524236ab3f4740a5b8cf47f79613af3c54b74ef01a3ed4b48bde

Observation 69085d9d-f981-42b0-af6e-6b402b3c108c · outbound

This paper cites Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.961117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.961117Z digest=sha256:85eb29e6a236777227ca49cf69fc5f0e77650964b80c3d49a14320338bbd0c9c

Observation 4b0e4d92-bf31-42da-9d46-796bfc215f4a · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.102392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.102392Z digest=sha256:705fd03120ca6d804f79d70df3764a7aa731818334e0d43fa0473376ef476413

Observation cededf8a-dc85-4318-8a83-feb4550efcb3 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.352897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.204007Z digest=sha256:67c0ed96c3668bc76e965b87e757350efc7fbd3e1ad9e91d046db4c73f15d8b2

Observation ea5536c0-bbce-4aa7-a322-e979271ac04c · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.304449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.304449Z digest=sha256:2c090cb03b7230af7fb1de7666088fdd9a45ab650f1b4a158c15c82866689f56

Observation 630fbc79-01bf-4030-8d21-19eaea262789 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.382349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.382349Z digest=sha256:bf8169385cb8f3ba4d28cfd2587bdb734049131e7928697f7886cdf75f7e74e6

Observation c742e5b0-bc06-45f7-b46a-4acf2b8aabc2 · outbound

This paper cites HTML Element.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback HTML Element

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.932407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.530802Z digest=sha256:3d379d1012f5dee8ceb4f01e1a7f1d6abcf0b4ff6cd84ba22173f8312d10b3a5

Observation 5675d98a-ce4b-4aa2-aaee-7759dd8e7b27 · outbound

This paper cites [button] Search [button_].

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback [button] Search [button_]

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.151987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.617653Z digest=sha256:e84da1baa9febf7b62d1f8a486abde877bee03d744c1141fa3f7b572b0cc29fc

Observation a0e1dcac-a190-44b0-be27-729e07bef74e · outbound

This paper cites HTML Element.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback HTML Element

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.719594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.704096Z digest=sha256:f17c4ed1374b16a6ff53433b9df4fe1ba6aadc3fcb785e4273278fd644a6b14c

Observation 79512170-f70f-40e2-84ed-d87277b7ec76 · outbound

This paper cites Refrain from B during your exploration.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Refrain from B during your exploration

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.511373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.818274Z digest=sha256:0ac7f99f465667ccb6e95a896940d46af9c359a3523bc764b6ad839ef83d9835

Observation ec6dd4e3-552d-4b7c-8235-5a64da8524af · outbound

This paper cites last action.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback last action

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.369594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:03:30.907605Z digest=sha256:016538ed05cf90938468d5ae1e9dbc505838f676ffb7bb65eb8bdc0a2459b243

Pith citing papers

Observation a9521675-1338-44e1-9fd1-7a1fcd88acd6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.668477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:7d0768e73a9df73bae583aa251897cbd3e68e422badc9b341f0adcbba8aec4af

Observation fbe65fd5-d18a-4dea-ab9b-1b4ff2b7432e · inbound

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History cites this paper.

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:47:26.254675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T18:37:55.374135Z digest=sha256:5f027c4c4501993441c90f68591fd0125f62f9369fb43ef764de7173ee53a480

Observation f519e037-47cb-4232-93ec-7bcf3a0b64f0 · inbound

Co-Evolving Skill Generation and Policy Optimization cites this paper.

Co-Evolving Skill Generation and Policy Optimization Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.856666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ac7a411883094e0c86a02a6ad1fb37513b1204fcd7bafa227fbbe1b58469ebea

Observation de972c30-1bde-40f2-b7e7-1512d1d73fd9 · inbound

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills cites this paper.

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.586534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:03:35.457624Z digest=sha256:2be0ff468f729cf947f1d2d67728682c5b43b3b370ad574507d89095cdfc931b