Pith. sign in

Paper Citation Record · LEDGER

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 4 inbound Pith citation observations for arXiv:2506.04287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04287 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:03:30.907605Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:37:55.374135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.584868Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ab9985b-b96a-4df7-9e36-d89cac920609 · outbound

This paper cites Variational Option Discovery Algorithms.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Variational Option Discovery Algorithms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:25.822971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:25.822971Z digest=sha256:1927f6937eb94c5e5a7f003962907571c78b323022b8b39b5c9a0bd2ffb1e3b2

Observation 1ac41ca7-a31c-4ac2-885e-2f056241fcbe · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:25.898272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:25.898272Z digest=sha256:891bb8f128eb88edb7c49755b7843b373bef1c1a68cc2ab68198194a19779797

Observation 6e17bb1d-f86c-4a73-ba9a-116714fb9ae1 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Mind2web: Towards a generalist agent for the web

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:35.056483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:25.999870Z digest=sha256:44ea32c388ec06029f1c69581cca5e20cb49c01e31862a85639085a3e79f1142

Observation 6a38fe2d-1f29-4a32-828a-4b9469d7eae5 · outbound

This paper cites Emergent complexity and zero-shot transfer via unsupervised environment design.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Emergent complexity and zero-shot transfer via unsupervised environment design

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.896987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:26.123499Z digest=sha256:09ed578f22b7b63d63d533132a7dcfdb6b7f5d6456daf1b9d044b9f64af2b83c

Observation 18c5c05c-2996-4ed9-8338-803627b57746 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Guiding pretraining in reinforcement learning with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.718439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:26.199768Z digest=sha256:176e1f175ff91b2adb551c3c91736b30a7caf7e1ec7bb8391f3de97fcd503470

Observation 0b8943a3-2a20-48fb-b35c-63242f273e51 · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Diversity is All You Need: Learning Skills without a Reward Function

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.274166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.274166Z digest=sha256:2116854dfba1d76fc0ee3b47b4aa446d4018e85374f889019aeeebefe3ce6958

Observation b6583d0c-6308-46fd-9fb3-6eeb0cda2849 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.369118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.369118Z digest=sha256:3bce8734bfa47559c9b21f9188293087f3f3cdcd0f92dc71675b09ff7cb270c2

Observation 42f0b07b-99a7-4878-840c-4c883bf3c048 · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Automatic goal generation for reinforcement learning agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.513776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:26.431916Z digest=sha256:037bf7c779911c9bb26dd6fb3b49a7c0f6297179feb697c13fc566bf6e34f0b8

Observation c2e78acc-a2dc-4968-a635-029dd4acd4aa · outbound

This paper cites The Llama 3 Herd of Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.535427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.535427Z digest=sha256:0268bf54d0921400d66725b541c10605eb17fe0e332e90becd30635611f63cee

Observation ea233952-099f-4a2e-b0b9-190f77d862e9 · outbound

This paper cites Benchmarking the spectrum of agent capabilities.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Benchmarking the spectrum of agent capabilities

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.272353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:26.599008Z digest=sha256:9a904c6f7ab86a5821d69a38b9439e2fb6146c366a3027eaa16a831be1311797

Observation 4f578e32-4a3f-40cb-b6de-fbc3bf95af79 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.700075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.700075Z digest=sha256:3f00baf4ad60d81d8693a27e3d4deb097de99121cd0490ea7fba1e862f1d0c77

Observation 66d50625-b891-4aed-ace9-db0823aae1c2 · outbound

This paper cites OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.768525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.768525Z digest=sha256:e63f19e18a25a75039cbf5844c6aadb641343504ce2071c3c885d26a73d6ff0c

Observation a82b0163-68eb-4292-87b9-5136b0fe8989 · outbound

This paper cites PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.888965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.888965Z digest=sha256:2d424c61cceba0fec22bb60ebbc352e835989835482aec8e6ad0464be89caeea

Observation 67839074-5a44-45b2-a04e-f08c049c142c · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:26.976890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:26.976890Z digest=sha256:645d0305b4f443626ab2a80e562d6063bf9f2e733c66de88ecb75961a703a585

Observation 3b905b5b-0efa-48ac-88aa-190aec4450e2 · outbound

This paper cites GPT-4o System Card.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.078439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.078439Z digest=sha256:91bcc5e2ac055f12b50807e15bede0899eb58c3d8cc4a752b621cd6ba4615a84

Observation 9c6c7823-92cd-41e4-a006-516e73cff31b · outbound

This paper cites Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.155990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.155990Z digest=sha256:6256edc799fb6787e0369c89bc083dc26020676423a277a2d1a1c620725d55e0

Observation 337970c5-63e6-4fb6-89ed-9a2f5323c6db · outbound

This paper cites DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.229732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.229732Z digest=sha256:8a78b85b81f7cda36a1bc67220e35478d31c3f98878708b17ecada4b6cab5fa4

Observation 4ceb284d-727c-4bb7-b356-c074b76ae856 · outbound

This paper cites Tree search for language model agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Tree search for language model agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.308410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.308410Z digest=sha256:417ef7f3d10b14906b8c1d7eeecaf5533ac92e8bd5d40b2ab8e5ef4625e504af

Observation 79a56932-f8fc-4c29-8c8f-c2dd52e78ab1 · outbound

This paper cites Autowebglm: A large language model- based web navigating agent.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Autowebglm: A large language model- based web navigating agent

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:34.068000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:27.379252Z digest=sha256:9dbad482c4d7d7feb0163008b01fa1b50f8578a7f8bbd3bbf59da9de3f89a7e2

Observation ecec396b-6a7e-4d14-850f-c1ebb1a1f658 · outbound

This paper cites Unsupervised reinforcement learning with contrastive intrinsic control.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Unsupervised reinforcement learning with contrastive intrinsic control

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.906592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:27.500008Z digest=sha256:8db317de64c1f24783544d0014f88c6d6381a1fc8a6da2c5313aa84166d5c511

Observation a9aab2d3-e448-45a5-866e-ada7183d5c40 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.589628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.589628Z digest=sha256:178235645ad3e0345e06e02fe42f0ce09294f5a4724ba7f7d4ccc43bdd1efd74

Observation 8522c990-18b0-4077-a5cb-feac09909bac · outbound

This paper cites Competitive Experience Replay.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Competitive Experience Replay

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.665012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.665012Z digest=sha256:c2a0e11be0fa08ce7397ac4b4e2ed0228e4b4256040c6679e71b2f480e4b49e5

Observation e0c742d5-f210-472c-baa9-ba357b71a29d · outbound

This paper cites Weblinx: Real-world website navigation with multi-turn dialogue.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Weblinx: Real-world website navigation with multi-turn dialogue

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.742181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.742181Z digest=sha256:0487c276a8ef3fb7dcfb95a72c453e25d354f0da5e35415774183d9de87e607b

Observation 05674ba6-ef51-4d28-a03c-bbc3aa619128 · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.809974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.809974Z digest=sha256:7da93e1562c3b676813f66d01645738c5bc59161b4af43f0c9daacbc22b94739

Observation e799eca2-72cd-497e-8185-e216d779606e · outbound

This paper cites Bagel: Bootstrapping agents by guiding exploration with language.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Bagel: Bootstrapping agents by guiding exploration with language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.732780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:27.860436Z digest=sha256:efb4cb0c1cbdb50ceeadf2db178ed34f4e515eb7b2586a437da61bde7203c6b3

Observation 488e211f-d605-477c-a681-2f58508d4a4b · outbound

This paper cites LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:27.944964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:27.944964Z digest=sha256:99caca5650de04c9dd491abca06e0922d53ff4a1f61b318791b11dbe7cc18bd4

Observation 634f9460-6bfc-403f-8bfb-fe6e1991ade3 · outbound

This paper cites Asymmetric self-play for automatic goal discovery in robotic manipulation.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Asymmetric self-play for automatic goal discovery in robotic manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.037262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.037262Z digest=sha256:eeeb9fa7152bf9abc798fdfef01b6c234edab5cb4e3eb99d69efd51735a4ee4c

Observation c36a30e3-5bda-43a7-b93a-6d06b1404b6e · outbound

This paper cites Balrog: Bench- marking agentic llm and vlm reasoning on games.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Balrog: Bench- marking agentic llm and vlm reasoning on games

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.576120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.126974Z digest=sha256:233fc9f6657ff2603550406d97a2d5e8ae6f80090392e4040776f9c4d4fdaa38

Observation ebb98b00-a759-4575-8cd2-94a9444b4612 · outbound

This paper cites Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.263675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.263675Z digest=sha256:f277066bce2fcc1ea30fc555d7c4536d242b68a456dacdea4bc45d8a8700b00e

Observation 8c74bece-053e-48a5-a5d5-fb27551fcbab · outbound

This paper cites Lipschitz- constrained unsupervised skill discovery.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Lipschitz- constrained unsupervised skill discovery

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.394607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.364199Z digest=sha256:ddfc0bfff7e72d6ad627a995d6c23b637dbdf27c3da34bc1fc4c87d0a511843f

Observation 704a9429-999e-460c-8b6e-f3e02914ac6d · outbound

This paper cites Accelerating reinforcement learning with learned skill priors.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Accelerating reinforcement learning with learned skill priors

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.216401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.477573Z digest=sha256:8a02b465327efbecad2854537ce5cd83e0b1c563537317978383599b7971b09e

Observation 017bcc25-aebd-4488-a748-ab051c631b87 · outbound

This paper cites Skew-Fit: State-Covering Self-Supervised Reinforcement Learning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Skew-Fit: State-Covering Self-Supervised Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.595829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.595829Z digest=sha256:a4973a5379a17b3ce9636d3c6d655d824e745b250bc569a130528771e851fc04

Observation c5baef8c-f312-41be-855f-8bc431a1338f · outbound

This paper cites Androidworld: A dynamic benchmarking environment for autonomous agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Androidworld: A dynamic benchmarking environment for autonomous agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:33.020514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.696270Z digest=sha256:6c05198d2b1c1a728c62df5167844d08901129c51eaf78cdc8bbaefd50201b2c

Observation 4eeba095-8a8a-41d2-9c4b-4e30f98584c3 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 2023.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.863729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.750473Z digest=sha256:7316ec2baf26fc4df853b15f4377075bba443cbe043ba5b0c467685e572d951c

Observation 86ac7cf5-2243-411a-8a22-63e118e12f77 · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:28.855289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:28.855289Z digest=sha256:ecec6ed626f4e59e8b58e64b1a6b86c0ad6f28cacc0bc98cf556fe3acd7d2541

Observation 78bac3d6-a90f-4b47-a82c-e9a1c09a807b · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.758345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:28.978049Z digest=sha256:0b6799d13524bc8978b77f6a984bab1cfa129a49df0d90d8095264b75c16e55a

Observation abd0e088-0109-4ed7-82e1-872d4053e928 · outbound

This paper cites Open-Ended Learning Leads to Generally Capable Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Open-Ended Learning Leads to Generally Capable Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.057904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.057904Z digest=sha256:63adff02b1431e0ddf116281be943047fa95f38f3a2ee36c01d350462e36111b

Observation 74b40594-2c33-493c-a6a8-b17576a88022 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.122001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.122001Z digest=sha256:4fd23414bf1c9eb085cce80b123142519018b7ba5f33548b0ccdba08aafae586

Observation 322b08e6-0357-46b8-a7a1-4d2a4cac3a6a · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.229207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.229207Z digest=sha256:78f11a11343ec09162ab109cad8e1b208a231a5f8e6eecd8216ff31fb9268197

Observation 9db23e1b-adfa-4b28-a065-d4a68766d11d · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.338089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.338089Z digest=sha256:ddf2c0791f3c171ad1cbbadca31d9007ec2dc246529a506c71a78e5570f62393

Observation 4062c452-38b2-4f84-85d2-4f95a5eb7cbf · outbound

This paper cites Qwen2.5 Technical Report.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Qwen2.5 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.407074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.407074Z digest=sha256:7031089e47d66f1997f7d7874e5bbd37f8bcffffa230b34251856b4bc3355d47

Observation 072215cf-aab4-4c81-a214-27e70e67ab8f · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.580264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:29.545300Z digest=sha256:33d3ee30e7b918cc1ea6ef71b83a712379d6454d0a985b9f72b4e4da87142b1b

Observation ddf49a7f-7c9d-4cea-bcc5-3efc2af5cef5 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback React: Synergizing reasoning and acting in language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.626585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.626585Z digest=sha256:72d1845fc7826b5d5afeccb7db3e1ecbea97946ae2c3ecd86aba5be0b211982a

Observation 32b89ad3-01d7-4694-a7af-8cc8b92086b1 · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.691987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.691987Z digest=sha256:bad15ca43a924241b8332acb39e4290867dc9d5fdf60b9071b1b7ba5a854c86e

Observation 5945f212-43e6-4604-aa8f-faf09b6d96d3 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.806310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.806310Z digest=sha256:cfe5f74589d7836fb1a4c8890b23d43635688f854cce32bfa4a413d4b5dafdb0

Observation 9a23aa97-d5ea-426a-8e65-38556fbb4eef · outbound

This paper cites OMNI: Open-endedness via Models of human Notions of Interestingness.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback OMNI: Open-endedness via Models of human Notions of Interestingness

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.887117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.887117Z digest=sha256:caea2296c97f7c80c1985935fccb5af0be781506a8d7c10640e9811207b121ba

Observation 69085d9d-f981-42b0-af6e-6b402b3c108c · outbound

This paper cites Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:29.961117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:29.961117Z digest=sha256:98045c5dddb1f5fa34fa123c0260fcf79d1093a01a86b57129a892d19d288c7f

Observation 4b0e4d92-bf31-42da-9d46-796bfc215f4a · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.102392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.102392Z digest=sha256:35fa64c1e23446c8e0d5b1b4857ed902b87fe5506fe3ca6b3d39b650014d46fe

Observation cededf8a-dc85-4318-8a83-feb4550efcb3 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.352897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.204007Z digest=sha256:623b9b5bae05c6ebfc9c6353dce74fa44cb5f2e4036ac1900cb8d9f675f7fd46

Observation ea5536c0-bbce-4aa7-a322-e979271ac04c · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.304449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.304449Z digest=sha256:4951c08a096d32e14f8bbcfd3c5b234a5c85aa777b056c17bba7e75a97572493

Observation 630fbc79-01bf-4030-8d21-19eaea262789 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.382349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.382349Z digest=sha256:254b1f660fee2d52830e2c9dfafae94515984c8fa4addbbadde1cf355a85cf47

Observation c742e5b0-bc06-45f7-b46a-4acf2b8aabc2 · outbound

This paper cites HTML Element.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback HTML Element

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.932407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.530802Z digest=sha256:f1a9879b56d412166d2b3854b811e9329e08e88f44149a62234c30293cf44c43

Observation 5675d98a-ce4b-4aa2-aaee-7759dd8e7b27 · outbound

This paper cites [button] Search [button_].

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback [button] Search [button_]

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:32.151987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.617653Z digest=sha256:5c53b46e14feaaaa7321bb2069902e48d68392d64e5e10e3f182dae98eb25c4b

Observation a0e1dcac-a190-44b0-be27-729e07bef74e · outbound

This paper cites HTML Element.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback HTML Element

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.719594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.704096Z digest=sha256:e05ac1f40898a8cd5b64adfbcdf28cb2a63f044f4d51e05cd2dca181ef5a15ca

Observation 79512170-f70f-40e2-84ed-d87277b7ec76 · outbound

This paper cites Refrain from B during your exploration.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback Refrain from B during your exploration

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.511373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.818274Z digest=sha256:b372583ccd6807652ea02a84c78b81cae08bc192628b4307d1575a0d314cceb6

Observation ec6dd4e3-552d-4b7c-8235-5a64da8524af · outbound

This paper cites last action.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback last action

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:31.369594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:03:30.907605Z digest=sha256:db76a5bfad7c1923dccccf6c19446ef97307f8af4727b16a42a592ff8609a88a

Pith citing papers

Observation a9521675-1338-44e1-9fd1-7a1fcd88acd6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.668477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:10282518983d4946ca7f7a51e50b0811c2e1ef3bedd2e44b1463994bae67985f

Observation fbe65fd5-d18a-4dea-ab9b-1b4ff2b7432e · inbound

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History cites this paper.

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:47:26.254675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:37:55.374135Z digest=sha256:ba031fd14016f4fad5e3eb6272b50af2f3f8c80ef9a5c04d9fdeb8f587c14d2e

Observation f519e037-47cb-4232-93ec-7bcf3a0b64f0 · inbound

Co-Evolving Skill Generation and Policy Optimization cites this paper.

Co-Evolving Skill Generation and Policy Optimization Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.856666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:b38525c195704fdf71c97133303f9c3f6c1bf8bccd667a8cb2caa24b7b8c499a

Observation de972c30-1bde-40f2-b7e7-1512d1d73fd9 · inbound

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills cites this paper.

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.586534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T05:03:35.457624Z digest=sha256:bb9398bd1b548c8b11ef99cc2958c07e7e3d7ada70f5c63afde806a1331ee7aa