Pith. sign in

Paper Citation Record · LEDGER

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

As of 20 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2506.20061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20061 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:33.026420Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:46:28.503923Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee5ef227-57f1-48ca-94f1-05a33cc845a6 · outbound

This paper cites Universal value function approximators.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Universal value function approximators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.816841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.373096Z digest=sha256:373b337e96a62706ff4e56051d96dd87396d51206cd5635e6e57587b88120abd

Observation a7e33da3-9a98-4b9b-9623-c7aba574a393 · outbound

This paper cites Hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.500447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.500447Z digest=sha256:65c273043d5d45deb33a1e6d53fdf78f60764a43865f16635776c11412b60fb0

Observation f9e5e4fc-9e4b-413a-b71b-5f24f81619dc · outbound

This paper cites Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.615893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.615893Z digest=sha256:4f2418705f30605158ff592832ff082b2335fe7ebb1c00eb6aabb2a3c478ea6e

Observation ce835828-b91f-43ec-9f59-4485abcca790 · outbound

This paper cites Grounding language for transfer in deep reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Grounding language for transfer in deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.565489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.777682Z digest=sha256:369d50f321431bc44f3bc62f9134c4f4facf3bc791d5b38e8203c524ba226aaf

Observation 37532636-43af-4b52-8434-53dd47ee7afa · outbound

This paper cites Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.905488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.905488Z digest=sha256:05b2a05a75ea931d6aa47af722ee8fb4703e0ba995347e4b182f8d61a25a3fd4

Observation 7974a871-e96a-47a5-9f66-4fe6ab15f670 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.107695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.107695Z digest=sha256:6b4d13bea2842d0d8f093dcc30f4c8130ca68243c1f0a68a1a3a762d04ff3c61

Observation b101b800-f6c9-4d69-9320-e99fe1209ddf · outbound

This paper cites Maximum entropy gain exploration for long horizon multi-goal reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Maximum entropy gain exploration for long horizon multi-goal reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.295432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:30.241905Z digest=sha256:f1352a374dceae1cd515fedaf085de2dd6e6942a4e64e9c9ee2b432732e93dca

Observation ef5a2712-587e-4d1b-b38e-843eac71b0c4 · outbound

This paper cites Curriculum-guided hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Curriculum-guided hindsight experience replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.378167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.378167Z digest=sha256:2c4256fffa884c68b17655805127039be9e62ed47069d74ce702c744b289e15f

Observation 38706882-630a-49cd-93d4-47fb0f010657 · outbound

This paper cites Exploration via hindsight goal generation.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Exploration via hindsight goal generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.546500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.546500Z digest=sha256:93ec2151d07cedf168c5d84d19275eacd7d5da3f78c4914b7f27383bdc3bbba5

Observation cb3ba943-5387-44ad-bab6-55dd54832e3f · outbound

This paper cites Visual reinforcement learning with imagined goals.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Visual reinforcement learning with imagined goals

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.719316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.719316Z digest=sha256:2d784433fad70dc9b5205ddcf625b4408fdb0427756d2dc0a9534df7548c7693

Observation 0676bdcf-9d3c-4064-a8a6-e31cf92c2201 · outbound

This paper cites Unsupervised Control Through Non-Parametric Discriminative Rewards.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Unsupervised Control Through Non-Parametric Discriminative Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.883300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.883300Z digest=sha256:3a9eb9ddc353f95ec0c11811287e6a7bdaac8d2068c6560fa802603fa1358617

Observation 9d332c08-ac5d-4255-8d10-a5b3c3b91e42 · outbound

This paper cites A Survey of Reinforcement Learning Informed by Natural Language.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models A Survey of Reinforcement Learning Informed by Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.066702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.066702Z digest=sha256:34ae0e0e218c5dd56a5b449a971a6a9badd97cfad01ad09fd5afd05ec172411f

Observation cfa8ee4d-b0b4-4fb2-b154-95b72e688eac · outbound

This paper cites The wisdom of hindsight makes language models better instruction followers.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models The wisdom of hindsight makes language models better instruction followers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.061171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.203145Z digest=sha256:56e136eb35e2dbfa8ab0ce07b7e985611058de3a4ded3b32f7fe70de9a773345

Observation 6d3a108f-8727-45c9-92fd-7970ba17c401 · outbound

This paper cites Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:03:33.440232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.331751Z digest=sha256:1d1ed6236bcdb9a70456e86cfca72a84e368f62c108aaa38c946d8fe8750f697

Observation 154d05e9-42a3-44d3-a35f-73f261b2fe91 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.451743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.451743Z digest=sha256:9e8b30fb5af946778fa2c4cfe3f9a2ebc97e0f3e77f2d1aaf62cfcd268be3643

Observation ce92ba1f-c190-465b-bf2f-deb45477acf8 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.726003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.573213Z digest=sha256:7d8c343d3f0ce8478e82a3f4112f37d2c8de5b350eac650c2d78dc27fe03878e

Observation 87ef1c50-298c-4ceb-91aa-0af950e38492 · outbound

This paper cites Inferring Rewards from Language in Context.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Inferring Rewards from Language in Context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.747052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.747052Z digest=sha256:a3a4cfeb238feb114f37865df5aae0d2a2b72c734463f4c96367571f8d7b9bbe

Observation 472cdb42-28bf-4cde-9f5a-e3c525469997 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Guiding pretraining in reinforcement learning with large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.435119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.894792Z digest=sha256:06ca1fe8daa78d9a1ec727cbfd8b3e6b46df04b43ea6bb91c4a535ddcb79fb69

Observation d4cb4566-d3f4-470e-87b5-8ecfd58751f8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models React: Synergizing reasoning and acting in language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.011774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.011774Z digest=sha256:60629c99fd93cde7464608f9b781c40448995458ba1b0ba97efd569269228038

Observation 125b9fc7-cde6-42f3-bf25-12de208fc593 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.159440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.159440Z digest=sha256:4c7c82dc9102eed2c27f06afe3fdebf132849ce3cb18fe5455739cda3b397167

Observation 4899dffe-6d59-40f4-b0e2-6883148a6f09 · outbound

This paper cites M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.302643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.302643Z digest=sha256:7eda7db496930e5596e361bda4de86025a77a2ec92964ca047e58d4ccdac3233

Observation 68263a5e-58f2-4328-8d97-e32ffec8b761 · outbound

This paper cites Puterman.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Puterman

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:32.436169Z digest=sha256:dbf2f8dec8b9a028dce8a3965e5ecc8b3a4baebe37b21dceede2f0db33eef93d

Observation 5d3772f7-f1c5-4ac3-82f1-a90506281224 · outbound

This paper cites Sutton and Andrew G.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sutton and Andrew G

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.619630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.619630Z digest=sha256:d21998d247ba176f563072c5577c1df277a37729aab1ae72d959a019001ef635

Observation c80792ea-c059-4d23-9320-953d70e826d4 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.776946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.776946Z digest=sha256:8df21e2578ee815dff8616ff4266f20f2615e811e1777a27c704e683ad79eb47

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · outbound

This paper cites Simplifying Deep Temporal Difference Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:64b54385634984661aa7c59ef559c3c339f16fc70d3c2cfac04a8eb2b89be45c

Observation 191974e1-80f5-4535-ae57-ebc654313a33 · outbound

This paper cites Prioritized level replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Prioritized level replay

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:33.846448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:03:33.026420Z digest=sha256:0411d223577d418483ba636332a4b294f1c96d6230f3123699d7bc6811409fbf

Pith citing papers

Observation 0e39e0b9-01c2-4be0-9674-7f933d8fded8 · inbound

Learning More from Less: Reinforcement Learning from Hindsight cites this paper.

Learning More from Less: Reinforcement Learning from Hindsight Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T00:46:28.503923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:46:28.503923Z digest=sha256:96a3dae8a9c2772a77d2fbc246344e4723b95a439461304c5001673af627139a