Pith. sign in

Paper Citation Record · LEDGER

PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2106.05091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.05091 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:52:27.258667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.257016Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 006cde4c-6c80-49fd-81c9-d12af47de548 · inbound

Active teacher selection for reward learning cites this paper.

Active teacher selection for reward learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:56:01.750823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T05:54:43.873174Z digest=sha256:0e5115d620851608e6d6d1245231b83fae55121375a96275961753ff6dc4554b

Observation 89384354-e872-455e-bf1e-cc4fd718b371 · inbound

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning cites this paper.

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:52:27.258667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:52:27.258667Z digest=sha256:d09ec0596a7a6323014b45b81897fc881c3a0deff03412e2c167dd624e2b611d

Observation 480fa63f-cede-40c0-b88d-6a38f8e293db · inbound

Learning from Active Human Involvement through Proxy Value Propagation cites this paper.

Learning from Active Human Involvement through Proxy Value Propagation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.213302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.213302Z digest=sha256:9a63157287a16577f429859f0a0e54f7ea39972df6bc8f4381be8a7911326dac

Observation f5750ba4-345c-479c-b362-073007f18f0d · inbound

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning cites this paper.

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T01:04:04.085762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:04:04.085762Z digest=sha256:b7dd248d38bd35186d6985cc39c5c60793714916547c8d55e635e18a75320dbb

Observation 67a4b0f7-931d-4a38-82ba-3573b732027a · inbound

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning cites this paper.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.977920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.977920Z digest=sha256:ff22ae3e8a4efbd4fa12200698bb7da000ca92294a3a4490b3f48748ea6da904

Observation 8190da58-22ab-43de-a3f7-162654f53e83 · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.030007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.030007Z digest=sha256:c91dd71bce024b8be8f3cd940bf31e3b967142da2ff80661961f28eb70cd5393

Observation 2a675f9e-6ba7-40da-b27d-cadf3d44b6e5 · inbound

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning cites this paper.

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:31:36.503042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T13:27:48.404574Z digest=sha256:e688f0c4c53f94dd843c0376b077f3f4fdc2b8c0cf2d701548990fab9f8147a8

Observation 9f862af4-11fc-402b-ac64-e871402ee531 · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:36.384982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:36.384982Z digest=sha256:4dab05d8e721b951a73d3332152a8434808af65bdc4a3f5529c4be13e532d18b

Observation ee2d89a7-e9a9-4d9b-ab0d-6fb0937dee98 · inbound

CueLearner: Bootstrapping and local policy adaptation from relative feedback cites this paper.

CueLearner: Bootstrapping and local policy adaptation from relative feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:57.368857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:57.368857Z digest=sha256:e179aac095162207ccd63bede2fa43d7530122197593e1c22404220ec5ac490d

Observation bf990ad4-6fe6-4e04-848f-c4f42f575879 · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.210585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.210585Z digest=sha256:7d544f3978e61a6e7fe6ea417b8194990f794c30bf6634bbb07ae85a2a1ced50

Observation 58fb53c3-1d03-4e44-9b60-4f6e10056144 · inbound

Active Query Selection for Crowd-Based Reinforcement Learning cites this paper.

Active Query Selection for Crowd-Based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:02:25.976213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:02:25.976213Z digest=sha256:b66a079c5d7eaf8980578dd398ab440a29a28b97a76460f4a080d79d566e5512

Observation 10b2cd0d-6808-4b23-a199-26ae8d81cedd · inbound

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning cites this paper.

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:06:56.860428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:06:56.860428Z digest=sha256:5dfcdcaa641e1b585c6cc026babdbbb44d97926ab64ca10988c922b8edd5dbeb

Observation 294511cf-d0ef-4e06-a517-4c9dee980e38 · inbound

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview cites this paper.

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T22:33:25.723393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:33:25.723393Z digest=sha256:33c021e718afaac5e309a546167057e9c685fc40b223224cd1b01e847957062f

Observation 7127ee49-baaf-4be9-99f8-d999492dfd71 · inbound

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback cites this paper.

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.090762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:35:03.577967Z digest=sha256:7c3a12c8ab1a0080e0b7857e70eb54e13e71ca6d3eab03278f6277c9aa713377

Observation 8544e529-e9bf-48bc-baa1-3c64508c4798 · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.258720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:040226726bdb1b6b190211b90fce12ae13dcf86f8e022b406ea54f626b095131