Pith. sign in

Paper Citation Record · LEDGER

MaestroMotif: Skill Design from Artificial Intelligence Feedback

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 5 inbound Pith citation observations for arXiv:2412.08542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08542 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:50:06.496418Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:36:11.391633Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:00:21.648976Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58cfd567-bace-4ebf-800e-0855d8b672b6 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.201993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.201993Z digest=sha256:24b20176a825e6d7743777fb02ed90173fa6cd2e5b023bfcd6e0c4ba4eb0b31d

Observation 137bb165-34c6-4539-9282-818027a12a80 · outbound

This paper cites Modular multitask reinforcement learning with policy sketches.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Modular multitask reinforcement learning with policy sketches

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.440718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.206987Z digest=sha256:b84cf7bfb7f55cd235b89bd0a9fa043b511d71c2a93d1ab3e43e24c40faf6e2c

Observation d396c276-d027-4a34-ab81-f2e5ff80b9db · outbound

This paper cites Senthil, and George Dimitri Konidaris.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Senthil, and George Dimitri Konidaris

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.427822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.211451Z digest=sha256:6365bbfbf18d2462f4f07cb3f696d60e45aad86a82e866906da070dfc9e58bda

Observation a58f98c9-f3a4-4a0a-b3b6-ccbe83e485ec · outbound

This paper cites Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and R \'e mi Munos.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and R \'e mi Munos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.415930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.215953Z digest=sha256:a16c5bbb2719cdd7fedad3c82d5bf1659b37f99cffc84af40b26885e05651b59

Observation c024aa11-85af-4a27-98bd-fd06892f0b7a · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.220070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.220070Z digest=sha256:0de2beaa9cb0d6507d4d5d7db33813cf29681fbc8541cbd7a8c84775867fac28

Observation 4cda2ca2-63e1-4678-acf1-6892541d3ccc · outbound

This paper cites Deep reinforcement learning from human preferences.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.224437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.224437Z digest=sha256:532765ad488f1c5c328d06626784ff04f471fcfe219a081f0b96c0f53536d14a

Observation f8bf0ad2-3eae-4966-874c-7b3826d98809 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.229138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.229138Z digest=sha256:9751d77a7fcf081ccea4c4fb65e8843d212a29fa8aa759e9c8a072349f0e55ce

Observation a3e71b35-88ec-4be5-bad2-cfbfbc9bf390 · outbound

This paper cites Active reward learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Active reward learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.396589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.232964Z digest=sha256:49fd522daf67c4d392d08e22836c291921dabf99dbff8995bc154da6096a434b

Observation 2fe29d7e-3dfc-46a0-8382-6b8c28229f27 · outbound

This paper cites Feudal reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Feudal reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.384682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.236643Z digest=sha256:539a1c8ca49a015a538792fb6b369af721ebb496b33e10c5a1485b2828bd37e5

Observation f9f3dd4d-7826-44fc-a441-4a139b69556b · outbound

This paper cites Drescher.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Drescher

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.372383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.240419Z digest=sha256:ba319bdb1d95ef9f297399ed813d12c12a1c1f8fd93c41c0b27332a89f61a56d

Observation e87da5df-3086-4d12-8a65-f5f9e5ca3bf0 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

MaestroMotif: Skill Design from Artificial Intelligence Feedback PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.244955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.244955Z digest=sha256:5c68024782ed605de3b97a98ca6d1179f95f149106da48cab9701ce6ccb7b926

Observation 8381b071-cda7-4d3a-848c-a0b31aa62227 · outbound

This paper cites The Llama 3 Herd of Models.

MaestroMotif: Skill Design from Artificial Intelligence Feedback The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.250015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.250015Z digest=sha256:d6f0c7f31e069b29e06461133025212f52071c1edfc1ec7013edc42105f5b572

Observation 736d0f78-8f4f-4f21-b084-caa495e58e4c · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.254663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.254663Z digest=sha256:24e52c345c9abdde1a502691733c5911e400e991cca35665032c6e6d124ce66d

Observation 9a433a8e-88eb-4a4f-839b-38d789c1320d · outbound

This paper cites Hart, and Nils J.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Hart, and Nils J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.346142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.258412Z digest=sha256:0d4983837e1aacb01adeb6773401f4fa892218b1dc441d2a76cde54682c9be07

Observation c967e085-3cf0-45cc-8d01-c892a1bd063d · outbound

This paper cites Variational intrinsic control.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Variational intrinsic control

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.333552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.262519Z digest=sha256:6ed505644b9cffbc9dd91435665775a481c3e80a5036af648710c41b6cd58f3c

Observation 6258b150-4377-4a54-b181-21a646b85509 · outbound

This paper cites u ttler, Taehwon Kwon, Donghoon Lee, Vegard Mella, Nantas Nardelli, Ivan Nazarov, Nikita Ovsov, Jack Holder, Roberta Raileanu, Karolis Ramanauskas, Tim Rockt \.

MaestroMotif: Skill Design from Artificial Intelligence Feedback u ttler, Taehwon Kwon, Donghoon Lee, Vegard Mella, Nantas Nardelli, Ivan Nazarov, Nikita Ovsov, Jack Holder, Roberta Raileanu, Karolis Ramanauskas, Tim Rockt \

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.322523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.266672Z digest=sha256:03ea5a2497aa0e7318f0b1aa9e4cc06a8a29d21574a85e61a18a3e50b4eef1ec

Observation 825399fc-fc59-4d19-adbb-6043b2211222 · outbound

This paper cites Dungeons and data: A large-scale nethack dataset.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Dungeons and data: A large-scale nethack dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.310372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.270432Z digest=sha256:90777b9739b9d69a50651be65510ffcd96e9ccf99aef63713ac5c3038e847ca4

Observation bc0b6704-6e62-4635-bd98-45973b0765ad · outbound

This paper cites When Waiting is not an Option : Learning Options with a Deliberation Cost.

MaestroMotif: Skill Design from Artificial Intelligence Feedback When Waiting is not an Option : Learning Options with a Deliberation Cost

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.274759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.274759Z digest=sha256:c66a8cb5f63ddf6c5b101e2c1570795a3b6bdd2366159e0d32219d6e7b5cd876

Observation fb84a863-430b-4439-952e-2295f123ac18 · outbound

This paper cites Exploration via elliptical episodic bonuses.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Exploration via elliptical episodic bonuses

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.298339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.279401Z digest=sha256:9cc84d6c6cae4e62da68dd63e0fe63f58e79174a245ab401fd14eee5993f8744

Observation a18b4ca7-2838-4a8b-a077-ba32edaf0af1 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:50:07.285986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.283884Z digest=sha256:8289df5ac9f3bc005d26f5cb4e8154abb8cc8c4eef90cb85b6c53f330141f161

Observation 3bee4fe8-bb36-4a22-988f-2115b4ed6d84 · outbound

This paper cites Open-endedness is essential for artificial superhuman intelligence, 06 2024.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Open-endedness is essential for artificial superhuman intelligence, 06 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.273395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.287785Z digest=sha256:a5b0d6fb84e0a9eb7f001e50d4981fce47349f29a55857fba57ebab16bc2f2b4

Observation e990c377-7225-4081-af6f-e8912b474127 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:50:07.261438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.291930Z digest=sha256:e063cb7e62ea0c6530e18217f0e1356600c0f25f1f381d3c0c885647bb839e2a

Observation 6711e68a-137c-44ff-96ac-5c8b362deb8f · outbound

This paper cites Options of interest: Temporal abstraction with interest functions.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Options of interest: Temporal abstraction with interest functions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.248119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.295915Z digest=sha256:d0a6336e4fd82c7fd35075e9c6a252433b23d310b488b47a14b5f6654554293b

Observation 64eb8769-ca71-4adc-9880-3ba01a4cf1c5 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:50:07.235751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.300665Z digest=sha256:8ecf28a444dfa48c2db579c4f0f2b8e90d55ba2d82a753c16f29b5c2ac370a45

Observation 1fc87d47-db2f-4926-b6c2-cdbb31bdb224 · outbound

This paper cites Motif: Intrinsic motivation from artificial intelligence feedback.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Motif: Intrinsic motivation from artificial intelligence feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.304422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.304422Z digest=sha256:2c9ac7ab217f81f2804d5241dd8aa41e725134f4624dac1ff26641788fd58c7f

Observation 914032c9-ad51-4f15-808c-50011e2a8f9b · outbound

This paper cites Keep your options open: An information-based driving principle for sensorimotor systems.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Keep your options open: An information-based driving principle for sensorimotor systems

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.215052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.308405Z digest=sha256:be5f3696d019cdd80254bb170f770cb8fe2931cd63651280409b20ea31df92f4

Observation 7f331168-b01b-445c-8738-fff879ea7452 · outbound

This paper cites Interactively shaping agents via human reinforcement: The tamer framework.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Interactively shaping agents via human reinforcement: The tamer framework

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.312568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.312568Z digest=sha256:b54aa4c9f7bdb73a2bb890b60009894c576f86a9d36198ccda20cfcca1d322cf

Observation 060cb00f-3593-4a44-b224-55f1cabc66f4 · outbound

This paper cites From skills to symbols: Learning symbolic representations for abstract high-level planning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback From skills to symbols: Learning symbolic representations for abstract high-level planning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.195485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.322091Z digest=sha256:46b73a5e66b06f0c8c4fb8018d82796cfff00f6c12958b98ee469a9bca893553

Observation 67cfa847-3ac3-4374-957d-56e0c425eabd · outbound

This paper cites Practice makes perfect: Planning to learn skill parameter policies.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Practice makes perfect: Planning to learn skill parameter policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.326391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.326391Z digest=sha256:3ce1b50fd814a37f502c75f84697954fcdc47a03304b8d2ebec64855e5068c62

Observation 66c0ba1c-9fe4-427f-a50e-9aa05329953f · outbound

This paper cites u ttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rockt \.

MaestroMotif: Skill Design from Artificial Intelligence Feedback u ttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rockt \

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.330268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.330268Z digest=sha256:7582545aae37dc1d2a0823eaa68b1a08119f868a1b9f80ad716c87dda99ec0a1

Observation 6ddcd9fd-3060-4456-8f03-c66ec95e385f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Gonzalez, Hao Zhang, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.334577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.334577Z digest=sha256:ca0fc3a5b7d0129f6cb3ad9a66770b2a0fd46e7948c38378b727ef83100c93f9

Observation 4cbe3cbe-b9b4-4914-8abd-a9e93af5e7e2 · outbound

This paper cites Code as policies: Language model programs for embodied control.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Code as policies: Language model programs for embodied control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.338517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.338517Z digest=sha256:73fd0359a93a02ca57522b81688a959c41d3df1f45d2cd944970b06036c69bb5

Observation 51bdf37e-afd1-4fc6-932a-1412f4cb66e2 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

MaestroMotif: Skill Design from Artificial Intelligence Feedback LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.342410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.342410Z digest=sha256:51f8d7f5d80047494b1c2c82265bcd33424d46bfdeefb51623680ba35985b74b

Observation a153dfa8-2807-45ae-a59f-19a3b5c41397 · outbound

This paper cites Goal-conditioned reinforcement learning: Problems and solutions.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Goal-conditioned reinforcement learning: Problems and solutions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.161693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.347887Z digest=sha256:2799bbad885e072d0c223de8f934670f8ea56bbe61630af9522f0a5fa8c0cc13

Observation 159ac966-efa9-41af-a8b4-dcf577d5b637 · outbound

This paper cites A laplacian framework for option discovery in reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback A laplacian framework for option discovery in reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.148018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.351838Z digest=sha256:68880a21d529564dea2eb6188d8712499be3e192dfd3227282c9657065a561bf

Observation f2290405-1187-4902-9ebd-e631078d6ec7 · outbound

This paper cites Machado, Andr \'e Barreto, and Doina Precup.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Machado, Andr \'e Barreto, and Doina Precup

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.135628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.355951Z digest=sha256:804b4bcc4b55d1fe2d45bc9f8b46a6f7003043e24eabdec7bc4abab6ee8e7f0c

Observation d1883cf2-22b1-458e-83c1-da7ca0e2bb9a · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Self-Refine: Iterative Refinement with Self-Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.359714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.359714Z digest=sha256:e9bc508bb92acc801d87d4053c27f73e347d0048047ef2184288b29dbe9bf22c

Observation cb0f5c7b-b459-4184-8d78-219be8b5d5aa · outbound

This paper cites Concurrent hierarchical reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Concurrent hierarchical reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.124086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.363855Z digest=sha256:a3bab289295662dcf80be0c3fa7bdd71d980a2ace4422c5f791a36a24a19ebe4

Observation cb4782cf-d8fe-4b23-84b5-5ddf4b847f6c · outbound

This paper cites Howe, Craig A.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Howe, Craig A

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.368361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.368361Z digest=sha256:9fafc977064ce13b8dc88a0cfbda981e20a6c3e6faa1da30d48d85845199ac18

Observation 02efb005-5ddf-4603-b9c6-5dc1be3884b3 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:50:07.104896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.371824Z digest=sha256:4680473f055b283b51bb2ebf4cc621e10b6e36bd4e8694d944bc668d51859ecd

Observation 7ab2bc44-2ad8-4b6d-9475-c8d6bcf5bf6c · outbound

This paper cites Q-cut - dynamic discovery of sub-goals in reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Q-cut - dynamic discovery of sub-goals in reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.091341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.375187Z digest=sha256:464a6c950fa82db2c088e72b9c6568bf65e900c0813d41453812516e6608c32f

Observation 1e7fcc68-6323-4854-9c59-08bac42ce7e0 · outbound

This paper cites Embodied lifelong learning for task and motion planning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Embodied lifelong learning for task and motion planning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.076869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.378876Z digest=sha256:f821e5415c091528dd9b791421871ec6fcbc67dc0048e4824fedab7c64d7b7b5

Observation 6d52731a-f09e-4e5c-8c26-cb6ca188c23d · outbound

This paper cites nle-sample-factory-baseline, 2022.

MaestroMotif: Skill Design from Artificial Intelligence Feedback nle-sample-factory-baseline, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.064439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.383274Z digest=sha256:d4dfa9eb7a51794a5548b1a22ae9785de1c1418fcf09ae2543af6d166fc68420

Observation 7aa0d318-7163-4f1f-96aa-de87c7503870 · outbound

This paper cites Nethack: an illustrated guide to the mazes of menace, dec 2022.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Nethack: an illustrated guide to the mazes of menace, dec 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.051581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.387934Z digest=sha256:6b01aa3b2dcfdec43dc67dc720163548d461ae5d6e850f5df29fe2786bc9bde9

Observation 7f4b56c5-a0ac-4ffe-96d6-d64cfe14125e · outbound

This paper cites Courville.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Courville

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.036570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.392445Z digest=sha256:a3651305ca5140c63c100cdb73264c1392c7fe1b1491af4f715f97ee3f2c56b6

Observation d0726826-1054-4226-bbb3-f5782a1bfe97 · outbound

This paper cites Sukhatme, and Vladlen Koltun.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Sukhatme, and Vladlen Koltun

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.397227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.397227Z digest=sha256:39519350f351acf7b406ba6182cae9abaa78c7425d50d945ac4ff18a9e3ed196

Observation 290c1b5b-f223-4a35-9d6d-07323d1b1f60 · outbound

This paper cites diff history for neural language agents, 2023 a.

MaestroMotif: Skill Design from Artificial Intelligence Feedback diff history for neural language agents, 2023 a

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:07.013920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.401323Z digest=sha256:52d85da7d2a48b0a2b6806dd30a259dfe43a6daee76c50fc16cead4b1dfc660d

Observation d8b61072-7d39-4baa-8e71-0860a909afed · outbound

This paper cites Nethack is hard to hack, 2023 b.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Nethack is hard to hack, 2023 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.999286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.405725Z digest=sha256:adce76084638487aa7f42221831f9b262046030a5e490daf95db455c312f3c97

Observation 2aa3c6f0-8ea0-4125-bf91-4afd1a9dd0f6 · outbound

This paper cites Temporal Abstraction in Reinforcement Learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Temporal Abstraction in Reinforcement Learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.986082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.409903Z digest=sha256:669f342ce1337a4afc843d0e4167ec84f52d599d2aab711c4779d9a03777fd7b

Observation b2b4dd38-44df-41b9-bbc3-70f94f6d17e8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.414644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.414644Z digest=sha256:89d95049c9853b9725d09828d1bb343a5470ff93f2807b9002943ec529cf2ec7

Observation 62fe808b-6720-489a-b511-6c819663d1d6 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Reflexion: language agents with verbal reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.973678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.419067Z digest=sha256:9df62498f77098cc427a0becd006a562d5c594acb1f5ab2c12f0be77032aa829

Observation 2e33b19f-8649-43dd-97ad-46cdcfbe71fd · outbound

This paper cites Tenenbaum, Leslie Pack Kaelbling, and Michael Katz.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Tenenbaum, Leslie Pack Kaelbling, and Michael Katz

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.961693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.423946Z digest=sha256:8094409d63887bd24557e7d59840b2fbbf6b72aadf34f54116d58be236e99598

Observation 821e2848-6d8b-4513-a18c-dbf647b41f48 · outbound

This paper cites Learning options in reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Learning options in reinforcement learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.950439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.428479Z digest=sha256:108bf0afc8d7f624ace626a0299c87644f79217d4ae64fe85589e1acbeebd21c

Observation 20731ef9-84cd-4a4f-94c1-294cc7b334f6 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:50:06.936818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.433101Z digest=sha256:99d250d439e492f4e0b48e88ef1afb164149d3f18afc911dfc6fa8df55972754

Observation 4ec9dc09-481a-45bd-b0bc-791c239cbf39 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Sutton, Doina Precup, and Satinder Singh

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.925459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.437170Z digest=sha256:6744b19a7e7e170c715166f18ba18f1f203a938df8f4c4eea23374570cdf1bc0

Observation b391c326-2e30-4cf5-b5fd-ba74910b3128 · outbound

This paper cites Autoascend -- 1st place nethack agent for the nethack challenge at neurips 2021.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Autoascend -- 1st place nethack agent for the nethack challenge at neurips 2021

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.913623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.442037Z digest=sha256:e3b284cab6ab09d3ffb8568df2a021cc8b740f46816b54e40cfc8938193e894d

Observation a3faedc5-bc65-45eb-8ffb-1fe2904d2a42 · outbound

This paper cites Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.900456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.446680Z digest=sha256:b8653a12db839a3105196a4f914045d22cd786d5ec862fdb5e2abcea0612bfe1

Observation 9ca69d07-8aa2-4091-8106-a53243740010 · outbound

This paper cites Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.888832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.450860Z digest=sha256:8c51cc20aa4ee25da6783363147c0ff9499b64fe93791e9ef776d4fc60946b29

Observation ac99ff11-84ca-489e-9fa4-8d1beaff95d8 · outbound

This paper cites Python tutorial, volume 620.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Python tutorial, volume 620

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.877013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.455099Z digest=sha256:fe8aad1921f1fdb946da084356653aed2dcd944e91b6a026e953b1ea7ebd3e54

Observation 4389f62e-b837-4b8f-a3b7-09800ce198c9 · outbound

This paper cites Feudal networks for hierarchical reinforcement learning.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Feudal networks for hierarchical reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.864933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.459789Z digest=sha256:54fdb8d392c55de540bff8b7bbeb441af452c4f72279e99963b5c8ffbbcfaa0a

Observation 0785c02a-8e1d-46f2-9cd7-6990b9d3d309 · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Voyager: An open-ended embodied agent with large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.853347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.463901Z digest=sha256:2dc44c404d15e9976e41742bd4537b0e3512c017b903514dbc931d5c4f47e305

Observation 0a074345-0ba5-40c0-a3ba-e6ab1fcd5cc5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.467363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.467363Z digest=sha256:520623cb329016fdc9963dcbe3525e69266e0be3203405b92f23c2ff27e52eb3

Observation e09f9c08-f7f2-49ef-8014-a744ecfef580 · outbound

This paper cites Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.470977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.470977Z digest=sha256:e09926effc755675000b17473b5627d2446672ff14bebfabbd0aa79b150190c6

Observation 15162f35-fc65-4bc2-9b97-ed6e66e2a274 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

MaestroMotif: Skill Design from Artificial Intelligence Feedback ReAct: Synergizing Reasoning and Acting in Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.475121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.475121Z digest=sha256:954a0c2e365215565398557ea7ba093e96beb49589c5e20c7dad9c7caf1e404a

Observation 4c8cd7bd-24a2-44a5-be15-edc10e0377d3 · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Noveld: A simple yet effective exploration criterion

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:50:06.833613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:50:06.479218Z digest=sha256:5faeee4855662e2025a9d6c023351ed7c6818806d983bcc05473f6098b7558a5

Observation 5e230a5c-e34d-4a87-be1a-5a88680b59fd · outbound

This paper cites write newline.

MaestroMotif: Skill Design from Artificial Intelligence Feedback write newline

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.482822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.482822Z digest=sha256:9605b0993b8b3792c1243651d7e3d4ce9cd7f1e1ecaab3d697b7845a03fbb934

Observation 5720ff46-3fcf-4189-9450-16e4606d7d07 · outbound

This paper cites @esa (Ref.

MaestroMotif: Skill Design from Artificial Intelligence Feedback @esa (Ref

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.487128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.487128Z digest=sha256:d8993768ae31284e36b361d318c1f601d7b5db1ee95bd463db551f57ac807876

Observation fc778430-d7a7-4baa-a838-0dc3ce3fd0d8 · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.492214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.492214Z digest=sha256:733e260025053da67efda1a20a837be086531bad0582908e47527beae6ea4908

Observation 5f034dd9-0daf-408a-8d16-9a277611284a · outbound

This paper cites an unresolved cited work.

MaestroMotif: Skill Design from Artificial Intelligence Feedback Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:06.496418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:06.496418Z digest=sha256:db618eca962fdeb8972e37cbfc28520315da4d75c740f744dbce04546ea60f3e

Pith citing papers

Observation ffe85de8-775e-47e8-96cb-5b6c8bb0ff93 · inbound

A Self-Improving Coding Agent cites this paper.

A Self-Improving Coding Agent MaestroMotif: Skill Design from Artificial Intelligence Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:36:11.391633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:36:11.391633Z digest=sha256:077bf7ecbe402d0de8282508f3f2275108725ef9c97cab7ac96c4706f27e3fdb

Observation 95e1046d-2733-43be-832c-05820eaa8411 · inbound

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities cites this paper.

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities MaestroMotif: Skill Design from Artificial Intelligence Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:29.684409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:29.684409Z digest=sha256:c14831456cc23e168ff1e77e5267a87302e9da9c73e090d773952b038aaa855c

Observation 7cf223ad-0ca7-4803-bddd-95b9d7d0ca99 · inbound

Hierarchical Behaviour Spaces cites this paper.

Hierarchical Behaviour Spaces MaestroMotif: Skill Design from Artificial Intelligence Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.774139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:33:48.527621Z digest=sha256:4acd858276e6250665018f53206c8aaf8829a53f669fb9bf75417a2a5e86f3f8

Observation 0150f5a1-ed7a-4c53-be4b-88a5ce04855e · inbound

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents cites this paper.

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents MaestroMotif: Skill Design from Artificial Intelligence Feedback

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:58.723344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:22:32.713175Z digest=sha256:8c60fec3b43bfc7207082528892a6e3ef451d5f04d8e0e88739d0d75fdfb1e8e

Observation 868116f4-dc41-43b7-a744-ee1ee21c8bbc · inbound

Goal-Conditioned Agents that Learn Everything All at Once cites this paper.

Goal-Conditioned Agents that Learn Everything All at Once MaestroMotif: Skill Design from Artificial Intelligence Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:21.652397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-25T04:59:48.867927Z digest=sha256:220e0d493ed9f2180edf3b179ca867de27bca22d8f4fb32aef1672190983b0ce