Pith. sign in

Paper Citation Record · LEDGER

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning

As of 24 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2506.13741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13741 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:04.401624Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3ca3b8-39bc-4df9-9bda-4faae0bb9f96 · outbound

This paper cites Forecasting with Multiple Seasonality.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Forecasting with Multiple Seasonality

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:33:04.542813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.096688Z digest=sha256:70135826f8fa40d97229f136ba34565b7b2b50357cc282e4e8ca54feee136e7b

Observation 0ff5fe61-f701-4d0e-98cf-f78aa95601dc · outbound

This paper cites Batch active preference-based learning of reward functions.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Batch active preference-based learning of reward functions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.769359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.165524Z digest=sha256:33c6e75cef74ae6eeac6c38b3156d504294b134daefd2cd0ca2f827b3b533284

Observation 3ca188d0-d1aa-4191-bee9-ae299a861ddf · outbound

This paper cites Landolfi, Dylan P.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Landolfi, Dylan P

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.761093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.194615Z digest=sha256:b7145dca3215aaeaf8b9587326137d54fbae37c2cc062d2919fc8cf7381892a4

Observation 16f14e3a-6d11-45d0-a1a3-9dbcc885d1ff · outbound

This paper cites Rank analysis of incomplete block designs: I.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.271354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.271354Z digest=sha256:d037043375e187ac1b5a43ff95cb5dbbab7b537732cb51901eac775804a8da4b

Observation bb8ca152-c1ee-466f-8552-22153d40bb14 · outbound

This paper cites Christiano, Jan Leike, Tom B.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Christiano, Jan Leike, Tom B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.747117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.342628Z digest=sha256:36f466a3b0b34219ec16271aee90ce62f97a2bde34a89eab83b1685a2189f381

Observation 75ee63ad-1625-4f68-87ca-0306261e0258 · outbound

This paper cites Autonomous skill discovery with quality-diversity and unsupervised descriptors.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Autonomous skill discovery with quality-diversity and unsupervised descriptors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.738394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.370760Z digest=sha256:9e639361eb0b55fa19dc9e7550630cc52e68189e6482d823d3dbcc6067d70000

Observation ee5644ab-4798-4701-ab72-26e93eb93764 · outbound

This paper cites Faster Improvement Rate Population Based Training.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Faster Improvement Rate Population Based Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:33:04.530615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.441449Z digest=sha256:e896e45f2b08518e40cc2082b1631ed15250a57ccb2f8045260dde17f1c5530a

Observation 98854a81-2221-4aef-b8d9-d1f072050231 · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Diversity is all you need: Learning skills without a reward function

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.730037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.502883Z digest=sha256:8e4070bc7c0346ce88d13b0b46afb91f2231894e0909cff25c9b7ec2a25f7b7a

Observation 8b8647fa-d609-449b-9991-3b6ff6253b6e · outbound

This paper cites Jax-lob: A gpu-accelerated limit order book simulator to unlock large scale reinforcement learning for trading.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Jax-lob: A gpu-accelerated limit order book simulator to unlock large scale reinforcement learning for trading

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.721729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.566117Z digest=sha256:ce557dbbf2420159ca614daeb9c649d681f252d3e1cdb455c72d0a4a45140a25

Observation 7018460c-2e08-436c-9e75-6936fccb3a54 · outbound

This paper cites u rnkranz, Eyke H \.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning u rnkranz, Eyke H \

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.622234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.622234Z digest=sha256:b1301cc0b57c08506f9e546f28e5aba68a7e1e2049c29eb0c7ac2b1223d16139

Observation 845db6e5-ee25-43ac-98cb-4ebbacc3c51d · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.685817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.685817Z digest=sha256:356105e45cf5481bc13ce2b1abd39acb779235f8eea5fbd727540b0c82ab23dc

Observation 20f8516b-548f-49ce-b3b0-246b310d433a · outbound

This paper cites Russell, and Anca D.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Russell, and Anca D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.707679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.753777Z digest=sha256:fc93aeb6981beca664fb151dbdf40648f93ea9819430bff2f9cd0d1204c9228c

Observation 2dc743d1-2a3b-46ca-9edf-4d6bdf6d34c9 · outbound

This paper cites Query-policy misalignment in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Query-policy misalignment in preference-based reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.698293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.789081Z digest=sha256:982d1e331e18086294715e9540e67b61444c3cf3066bb9c7be78db2296548cb1

Observation b34a91bc-ede1-4ca6-9fce-25bd053e69b7 · outbound

This paper cites TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.820968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.820968Z digest=sha256:f79ad827093beb05afbee8a8f2ea1b3ec6b3297635c813e90c4d46a369bfe08a

Observation fc53ef44-60ad-4df3-ad40-b688b6a136cd · outbound

This paper cites Population Based Training of Neural Networks.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Population Based Training of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:02.870174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:02.870174Z digest=sha256:478f912dc37c0689df464e6d6b7bc642146377dff24ec906767e321fa89cee7c

Observation 4a483ab6-d007-40f8-8f9b-c6b3a0bf95af · outbound

This paper cites One solution is not all you need: Few-shot extrapolation via structured maxent rl.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning One solution is not all you need: Few-shot extrapolation via structured maxent rl

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.689812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.922990Z digest=sha256:f98ecab9beccfce384569af40eb161b78b6c18a34b23477fcc38114dfdefd099

Observation b0a333ab-5437-4079-b7de-e88058b5fd29 · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.682059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:02.965011Z digest=sha256:bca7bb36d9068678819b6daffae3a73ae73092bbd0fbd6ce111a879cb1677bd5

Observation dea18995-5bd0-4844-b2e9-15ef7185ed2b · outbound

This paper cites Smith, and Pieter Abbeel.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Smith, and Pieter Abbeel

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.674287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.009880Z digest=sha256:3ecb3d3d9d243f57bf5c7704cb9112135e58c9404b13b80545c216fbe7710bbe

Observation 22081c72-68ed-4432-befd-c43cfb4912b9 · outbound

This paper cites Evolution through the search for novelty.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolution through the search for novelty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.666897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.076339Z digest=sha256:f283e2ed6c427ac9ae70bf21aa14ba1c5f79190bd5f21333e25c6e1bb2b32713

Observation 7394ec4d-2201-40c3-adc0-e427cf205823 · outbound

This paper cites Evolving a diversity of virtual creatures through novelty search and local competition.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolving a diversity of virtual creatures through novelty search and local competition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.659279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.161033Z digest=sha256:910572e681ce7ef879b248c11cff193be26e277c1496c8e6e6272f92bd322213

Observation 8f765897-1cd7-4aa4-b3e0-3abdf2ead841 · outbound

This paper cites Reward uncertainty for exploration in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Reward uncertainty for exploration in preference-based reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.651000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.215678Z digest=sha256:5f3e252e99da65141eddc40a057ac80addf410fc2627791f5c9c798157256179

Observation d7e0be92-746a-4f3c-b441-eb77449568ab · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.643003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.280078Z digest=sha256:5f43b42dd250bb6eafa51bf286cdf7880825fe8c199aa1fc3af21795aa0eb92c

Observation 8b9e4ef6-97bd-43a6-9945-99b25d8274ca · outbound

This paper cites Efficient preference-based reinforcement learning using learned dynamics models.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Efficient preference-based reinforcement learning using learned dynamics models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.634836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.339377Z digest=sha256:b5fe6c3d0f645e507f33ce6b181ec6237ca070291bef713b2ef02ae5cc46f678

Observation 6856a5f5-707d-44d1-863e-9b9c6e4b3f86 · outbound

This paper cites Variquery: Vae segment-based active learning for query selection in preference-based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Variquery: Vae segment-based active learning for query selection in preference-based reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.626755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.389931Z digest=sha256:e18455e5958d9ae1bc8de382858700217ab9c1c06469609411bdef1687c213aa

Observation 5e0af0b1-3115-40a0-80bb-045cb1e8d4b5 · outbound

This paper cites Rewards Encoding Environment Dynamics Improves Preference-based Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Rewards Encoding Environment Dynamics Improves Preference-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.458839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.458839Z digest=sha256:7687b194a55170652f6f1bde7cbd2b1f3a62f0cec6ddc58b93efbafe7c74ca95

Observation 10ecdd6d-fdf7-4749-a094-8816e6e2010b · outbound

This paper cites Illuminating search spaces by mapping elites.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Illuminating search spaces by mapping elites

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.520942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.520942Z digest=sha256:338b0e953977a4faeb1a17139a85ac319bf8b5a8b80285fdfb0a18847ae534f7

Observation cf25001d-17db-4534-a18f-7a5581b6f219 · outbound

This paper cites Xland-minigrid: Scalable meta-reinforcement learning environments in jax.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Xland-minigrid: Scalable meta-reinforcement learning environments in jax

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.618180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.600291Z digest=sha256:b4d77ed3f94161ab8c214c093c01d80247887338d315286935030374835ecba0

Observation 6f73c65c-9e1d-49c8-be74-979b7d1236bb · outbound

This paper cites Policy gradient assisted MAP-Elites.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Policy gradient assisted MAP-Elites

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.608896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.650089Z digest=sha256:8669465dc12149bcb6738bd69ff5f817110976444d1dbff0b8657b06a4b6dae9

Observation a166f64c-c237-412e-972a-f1dbb3340dce · outbound

This paper cites Discovering diverse solutions in deep reinforcement learning by maximizing state--action-based mutual information.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Discovering diverse solutions in deep reinforcement learning by maximizing state--action-based mutual information

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.600764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.732003Z digest=sha256:e59ce29eb5fa3773bc86fafd1cac69bef6a1a82924ace2c06bc5e2553d9d8371

Observation c0ed7339-82d8-44be-a7b7-a506a191fd40 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.800435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.800435Z digest=sha256:2012d135ac570fba611e0a4a47fefba9a17cb3be13ca827698cc04536b500734

Observation 47273825-0496-4a5f-87f8-5a271a79c99c · outbound

This paper cites Effective diversity in population based reinforcement learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Effective diversity in population based reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.592571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.883178Z digest=sha256:c705f9d6111fd4d96d8bf85a0fcd62f0e519138abf47acae48d43d23a6fb8790

Observation 4d3691d9-59db-49ee-868b-4f9f7203e226 · outbound

This paper cites Jaxmarl: Multi-agent rl environments in jax.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Jaxmarl: Multi-agent rl environments in jax

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.583969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:03.948814Z digest=sha256:93946dd82d551ab1a5809a76a19f7b3e874204e7ffc00a691b050f20e387a28e

Observation cba07e3f-6afc-47cc-8173-ad5d6961a15d · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.015747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.015747Z digest=sha256:992798b00a4c3d2b1d985d9550e5c0b61b069163a9803e43e8ca54d517551939

Observation e602c915-ec49-4297-8476-52b38f3e57eb · outbound

This paper cites Dynamics-aware unsupervised discovery of skills.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Dynamics-aware unsupervised discovery of skills

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.575052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.098928Z digest=sha256:c0ce45979de7e20d339b46b029f45db6f8b6f8cb1c7868032488c7ccd3a586a6

Observation 100cb00e-0586-444b-b90d-34a2793461b6 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.163238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.163238Z digest=sha256:58b1b22308109cb1511bdcde3bce0d7add93f2aa376c581a5b950a97410d2464

Observation 8ea88ba8-32ca-403b-9456-f0b60699cb2d · outbound

This paper cites DeepMind Control Suite.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning DeepMind Control Suite

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:04.252878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:04.252878Z digest=sha256:eee1ab791c35e041c4d7616753b71d3ef96d53d8fd2d46bf17d79c3994211a8d

Observation c6be78c0-e1dc-45bd-92cf-e897ea2f2946 · outbound

This paper cites Ball, Vu Nguyen, Binxin Ru, and Michael A.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning Ball, Vu Nguyen, Binxin Ru, and Michael A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.560627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.332746Z digest=sha256:bb204250e672ee83a8aabca9f548b557eed28a62ea3b2df21ea1731bbe1639a8

Observation 1591528d-92f0-4345-a103-0f26b8bbe050 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning A survey of preference-based reinforcement learning methods

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:04.552166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T00:33:04.401624Z digest=sha256:d1ab82e0cd23c92b6ce847a3a4f0bbac56951b7688650d9af503db7590372fd0

Pith citing papers

No inbound Pith citation observations are available.