Pith. sign in

Paper Citation Record · LEDGER

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2506.20061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20061 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:03:33.026420Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:46:28.503923Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee5ef227-57f1-48ca-94f1-05a33cc845a6 · outbound

This paper cites Universal value function approximators.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Universal value function approximators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.816841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.373096Z digest=sha256:58b707152ccd9c3a4b58157453ce3037deee894b8f986d501ef0d5540c42f7db

Observation a7e33da3-9a98-4b9b-9623-c7aba574a393 · outbound

This paper cites Hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.500447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.500447Z digest=sha256:6ede131b7366a9fc94893db5351330a1cc00cd4e8bad5adac6fbeeafb555456b

Observation f9e5e4fc-9e4b-413a-b71b-5f24f81619dc · outbound

This paper cites Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.615893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.615893Z digest=sha256:9871be88ac021b17b456fd711ac8d9bd8e6622ee375ae19421e0b82b2b16d540

Observation ce835828-b91f-43ec-9f59-4485abcca790 · outbound

This paper cites Grounding language for transfer in deep reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Grounding language for transfer in deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.565489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:29.777682Z digest=sha256:165e036e8d776f3809c9bdd795908b8dde195f4f2acbc736a346648afc184b48

Observation 37532636-43af-4b52-8434-53dd47ee7afa · outbound

This paper cites Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:29.905488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:29.905488Z digest=sha256:41ea412e623d7c166739da73235d9dede345a9a56702df2f5cf59534bfa6affd

Observation 7974a871-e96a-47a5-9f66-4fe6ab15f670 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.107695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.107695Z digest=sha256:eb75a24ff04f1d8b0b156f2036e70cb2bccbfa961a744f40576f7de811ca1996

Observation b101b800-f6c9-4d69-9320-e99fe1209ddf · outbound

This paper cites Maximum entropy gain exploration for long horizon multi-goal reinforcement learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Maximum entropy gain exploration for long horizon multi-goal reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.295432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:30.241905Z digest=sha256:99f6e692082a7ac7472a7f59f24653f1c70e8e54a4c635ff3be5d1522421bdfa

Observation ef5a2712-587e-4d1b-b38e-843eac71b0c4 · outbound

This paper cites Curriculum-guided hindsight experience replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Curriculum-guided hindsight experience replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.378167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.378167Z digest=sha256:b015ee693d3b4e2a637648258e275d389631c95a974a947edb0542bba0e9c80d

Observation 38706882-630a-49cd-93d4-47fb0f010657 · outbound

This paper cites Exploration via hindsight goal generation.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Exploration via hindsight goal generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.546500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.546500Z digest=sha256:40b8fea4a53f5f5209d9df478468b93f7959b55f8006ac440c82b177db97eb5e

Observation cb3ba943-5387-44ad-bab6-55dd54832e3f · outbound

This paper cites Visual reinforcement learning with imagined goals.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Visual reinforcement learning with imagined goals

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.719316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.719316Z digest=sha256:0d8eb8d99dea482deca0d09f79c5f1a5c31daa450becdb676c112aa88ea01013

Observation 0676bdcf-9d3c-4064-a8a6-e31cf92c2201 · outbound

This paper cites Unsupervised Control Through Non-Parametric Discriminative Rewards.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Unsupervised Control Through Non-Parametric Discriminative Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.883300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.883300Z digest=sha256:ec57b8db7f7e0c281d0e1135862dfc2d8b980517b08f41510b4889f00ff4b1fa

Observation 9d332c08-ac5d-4255-8d10-a5b3c3b91e42 · outbound

This paper cites A Survey of Reinforcement Learning Informed by Natural Language.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models A Survey of Reinforcement Learning Informed by Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.066702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.066702Z digest=sha256:dc411fa5354e487ef9f1e6cd95f4adaf10328d0a46a0ea1f64bc8fb2d591e28e

Observation cfa8ee4d-b0b4-4fb2-b154-95b72e688eac · outbound

This paper cites The wisdom of hindsight makes language models better instruction followers.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models The wisdom of hindsight makes language models better instruction followers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:35.061171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.203145Z digest=sha256:70b779c0ce2e17322ba00e2ee6c825164d3849a842332a02933eaa61b6353b2f

Observation 6d3a108f-8727-45c9-92fd-7970ba17c401 · outbound

This paper cites Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Analyzing the Stance of Facebook Posts on Abortion Considering State-level Health and Social Compositions

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:03:33.440232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.331751Z digest=sha256:5df68993fdf8345f8555bd4883db3add091f49f6357f06134b2f714a337030ae

Observation 154d05e9-42a3-44d3-a35f-73f261b2fe91 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.451743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.451743Z digest=sha256:1382e19d336eef620ccab598602ff1cc1ea90144ea937fac778b07de83f368a2

Observation ce92ba1f-c190-465b-bf2f-deb45477acf8 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.726003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.573213Z digest=sha256:c93014ee03319d32720946092848dea52c92c8e4fc38aa77ec4dea01c8dfd609

Observation 87ef1c50-298c-4ceb-91aa-0af950e38492 · outbound

This paper cites Inferring Rewards from Language in Context.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Inferring Rewards from Language in Context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:31.747052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:31.747052Z digest=sha256:7f8a410588d6859853a923d1ae1591576f99a3953d048152bc67faa19a737a05

Observation 472cdb42-28bf-4cde-9f5a-e3c525469997 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Guiding pretraining in reinforcement learning with large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.435119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:31.894792Z digest=sha256:a92b0d0f255ab33b1b9403ef6148320478aa0f782d95415e198a70dd026c54ad

Observation d4cb4566-d3f4-470e-87b5-8ecfd58751f8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models React: Synergizing reasoning and acting in language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.011774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.011774Z digest=sha256:b559e2856a0b4a87d4cadf1432b6d9147a1908c1ff2c16ff651e532b493155d8

Observation 125b9fc7-cde6-42f3-bf25-12de208fc593 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.159440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.159440Z digest=sha256:c8df4e2a1eeb01ea2aa55c9a7e17b6b57acf0b225fddd97d23345d3eb6ccba5e

Observation 4899dffe-6d59-40f4-b0e2-6883148a6f09 · outbound

This paper cites M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.302643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.302643Z digest=sha256:35c22de96c6dda3f1b340c2133f3ed3e87591f76ff41a31e4da10cb4513ff51d

Observation 68263a5e-58f2-4328-8d97-e32ffec8b761 · outbound

This paper cites Puterman.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Puterman

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:34.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:32.436169Z digest=sha256:1eb80488f82f73beded6d9bcdbab7d0453f10f0938a90ba875bd2f799fb7d070

Observation 5d3772f7-f1c5-4ac3-82f1-a90506281224 · outbound

This paper cites Sutton and Andrew G.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sutton and Andrew G

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.619630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.619630Z digest=sha256:5c5fca7e2a8b91dd14add09fd9052c65f564cf8242995e41a1a92a1348505f5e

Observation c80792ea-c059-4d23-9320-953d70e826d4 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.776946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.776946Z digest=sha256:a63f2548e6a07db282c8c60c570722aadb6b2c5d861d0835c614051df69d5927

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · outbound

This paper cites Simplifying Deep Temporal Difference Learning.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:56ebad87ab02fc0b4c8b444a94b8bce3281f2c72094e09a82b51652a8998ff50

Observation 191974e1-80f5-4535-ae57-ebc654313a33 · outbound

This paper cites Prioritized level replay.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Prioritized level replay

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:03:33.846448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:03:33.026420Z digest=sha256:2cbf34ea9af4314f4043b4520acb7eb9f1b0d4f07905f710471422f9f6a899aa

Pith citing papers

Observation 0e39e0b9-01c2-4be0-9674-7f933d8fded8 · inbound

Learning More from Less: Reinforcement Learning from Hindsight cites this paper.

Learning More from Less: Reinforcement Learning from Hindsight Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T00:46:28.503923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:46:28.503923Z digest=sha256:d0aee218f980f3baa48f50bb03d7e01736d4deb430dee54ecf4629514a4e56ec