Pith. sign in

Paper Citation Record · LEDGER

Learning from Active Human Involvement through Proxy Value Propagation

As of 10 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2502.03369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03369 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:13.369371Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d390698-79ce-46a3-a28b-e79ecaabfe62 · outbound

This paper cites Agent-Agnostic Human-in-the-Loop Reinforcement Learning.

Learning from Active Human Involvement through Proxy Value Propagation Agent-Agnostic Human-in-the-Loop Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.097061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.097061Z digest=sha256:186c495cd00e6c97aeb4729669baefe932a960fb438f3db0a35d5665015c6cd8

Observation 1f779148-5b50-4f46-ab15-ea9f0849d2da · outbound

This paper cites Constrained policy optimization.

Learning from Active Human Involvement through Proxy Value Propagation Constrained policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:14.042162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.102328Z digest=sha256:1e12b224b5c029dc47143e9f5e663c4a5562a014bf0411f281a0e0f3865334b9

Observation 6551bf87-7e4c-4241-9215-45c03fb950f1 · outbound

This paper cites An interactive framework for learning continuous actions policies based on corrective feedback.

Learning from Active Human Involvement through Proxy Value Propagation An interactive framework for learning continuous actions policies based on corrective feedback

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:14.031627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.106908Z digest=sha256:346bfa26cbe6ee76872091010d11bd61efe3086dc2467b05b35ec3c14bb488c7

Observation 9f113462-bbb9-4ff1-bc2c-4d1475b62433 · outbound

This paper cites Minimalistic gridworld environment for openai gym.

Learning from Active Human Involvement through Proxy Value Propagation Minimalistic gridworld environment for openai gym

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:14.020818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.111768Z digest=sha256:cba831bbb77f3874d667469cdbb4e122f476a8594211f8ef5faa0f889526bc4e

Observation 43d96413-8c44-46c5-a6ed-b9ad7354217f · outbound

This paper cites Christiano, Jan Leike, Tom B.

Learning from Active Human Involvement through Proxy Value Propagation Christiano, Jan Leike, Tom B

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.115607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.115607Z digest=sha256:85412e670ee787dfbc296e786f73a0c0afb2f871da084523007571422b71f6cf

Observation a61c1ff7-38cd-4a67-a958-2eb43de93f67 · outbound

This paper cites Open Problems in Cooperative AI.

Learning from Active Human Involvement through Proxy Value Propagation Open Problems in Cooperative AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.119853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.119853Z digest=sha256:8ee46478946e2a7e8224f1384643a5ab5c51c40ca36a453435d46c5b1ccd4172

Observation 4ef4784d-5dbf-4407-9428-04d38b84f334 · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:14.003269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.125143Z digest=sha256:a250a555922c33b1793a0fdc745af267a0243fc0b86f5127074073cc99e34773

Observation 5cbdc2e8-a9e3-4a6e-92db-05407621c93c · outbound

This paper cites CARLA: An open urban driving simulator.

Learning from Active Human Involvement through Proxy Value Propagation CARLA: An open urban driving simulator

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.129124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.129124Z digest=sha256:6356d39baf0436b5763c971b17ff541beb89fb3c4c10ea09297b437c8ba14016

Observation e5614c5d-7d61-4601-9830-9e1c162b2a6f · outbound

This paper cites Learning robust rewards with adverserial inverse reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Learning robust rewards with adverserial inverse reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.985809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.132885Z digest=sha256:09237f2198658d3768689923d02250a6d98ce5a49f86d3e59841ec709a3fc446

Observation 7b39b107-02c1-44ed-804f-de83d3e30c30 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Learning from Active Human Involvement through Proxy Value Propagation Addressing function approximation error in actor-critic methods

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.975624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.136365Z digest=sha256:e99bec5ebc056479f61964de3806236efcb46229b578effc5149c483a75fd497

Observation 7222c1b5-b6b1-4e9e-a42b-73c4e1fd9997 · outbound

This paper cites Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation.

Learning from Active Human Involvement through Proxy Value Propagation Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.964698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.140422Z digest=sha256:ade513c6b973e56bb119d5682fd3816f64a58705a2f7b66200b7514a0b3c5014

Observation 8aac2172-d86a-40e0-b6c3-1092cba33595 · outbound

This paper cites Learning to walk in the real world with minimal human effort, 2020.

Learning from Active Human Involvement through Proxy Value Propagation Learning to walk in the real world with minimal human effort, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.953813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.143910Z digest=sha256:8b77657081441a2a9e7724735515061dfe3185c27afd8254c20bde1f4dd80285

Observation 5814ac2b-f61b-4abc-a940-19b105a23d71 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Learning from Active Human Involvement through Proxy Value Propagation Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.943700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.147681Z digest=sha256:dfb47db19a57f9c53c3003d27126cf36fda2af86d833ee7ee9554f13f0b172dd

Observation a12310a7-8404-448e-b892-16650f1dd590 · outbound

This paper cites Generative adversarial imitation learning.

Learning from Active Human Involvement through Proxy Value Propagation Generative adversarial imitation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.932348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.150972Z digest=sha256:5b93f2eedcb78b4cefd261ffdb42f7bf3b83c7902c5ead34d2bf2b67e1bf4d0c

Observation 2d9f4182-d0a1-46b2-a062-953718f16410 · outbound

This paper cites Precise Synthetic Image and LiDAR (PreSIL) Dataset for Autonomous Vehicle Perception.

Learning from Active Human Involvement through Proxy Value Propagation Precise Synthetic Image and LiDAR (PreSIL) Dataset for Autonomous Vehicle Perception

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.156609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.156609Z digest=sha256:29c9d11d225c7307291711cac7f539b3c91f05f1996a0c1d3e671904e8b36ba8

Observation 2d88299e-4df6-4c4a-8b8b-cf1c3d882959 · outbound

This paper cites Learning to share autonomy across repeated interaction.

Learning from Active Human Involvement through Proxy Value Propagation Learning to share autonomy across repeated interaction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.921905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.175680Z digest=sha256:7e73707258351d33782374aef608da9308b775bac23a2b8bda69be04ad9dd4c0

Observation bbbdf038-5ea7-40bf-b2db-ddc029307981 · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts.

Learning from Active Human Involvement through Proxy Value Propagation Hg-dagger: Interactive imitation learning with human experts

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.911648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.195567Z digest=sha256:effbcd79ea43812c58bdb767c66c822d05dcaaff16499167db4327e6a22f5e20

Observation 515442c5-2d78-450d-8dab-92b0fa12c494 · outbound

This paper cites Learning to drive in a day.

Learning from Active Human Involvement through Proxy Value Propagation Learning to drive in a day

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.901302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.199268Z digest=sha256:e1a6212f0b7810e677b1209e4b7e9ef64c173a5da650134e4040ebd427e655a9

Observation 12e15a05-11a0-42cc-a893-0fdd3527d49e · outbound

This paper cites Reinforcement learning from human reward: Discounting in episodic tasks.

Learning from Active Human Involvement through Proxy Value Propagation Reinforcement learning from human reward: Discounting in episodic tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.891483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.202906Z digest=sha256:3396bfa6cb08831d0129de53ddaf55dbcb13dab6d3293edff6bd35a6539efb35

Observation 8ec8c7b7-888c-4308-a845-7c7bf919bb68 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

Learning from Active Human Involvement through Proxy Value Propagation Specification gaming: the flip side of ai ingenuity

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.880895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.206101Z digest=sha256:281f8a00421f0c57c72d4199dd387de8e40668da2dcd610afb968b506668555e

Observation eae9d699-d97b-4fa8-8a8f-b935d18a6caf · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Conservative q-learning for offline reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.870747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.209424Z digest=sha256:c472b454cae116686eb7094d6bc3c11e2c6a3261345c4367505d6a0038015ba9

Observation 480fa63f-cede-40c0-b88d-6a38f8e293db · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

Learning from Active Human Involvement through Proxy Value Propagation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.213302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.213302Z digest=sha256:9a63157287a16577f429859f0a0e54f7ea39972df6bc8f4381be8a7911326dac

Observation ab5c7713-fe20-4ef9-a9f2-ad7da4b506d9 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Learning from Active Human Involvement through Proxy Value Propagation Scalable agent alignment via reward modeling: a research direction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.217444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.217444Z digest=sha256:f426f205558abf148e2ead7d6f05ea0626279848dc11353f3a85a561b0117233

Observation 800cc48d-9186-41e3-bec0-9b126f665946 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Learning from Active Human Involvement through Proxy Value Propagation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.221593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.221593Z digest=sha256:c1a04d8b7ea6c7e86abd16b6ee8f6b1701a1748a18124a4ea607caa3c03c6679

Observation 41784e27-ff14-4fe3-a9ec-8d65cb273d50 · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.860131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.225792Z digest=sha256:2eb11a5f9e54adbb97b4e16e708463bf8c94f7c464369c424dfdb621e875f06f

Observation 613635c1-2af8-4d10-9e08-ba39b93faf50 · outbound

This paper cites Efficient learning of safe driving policy via human-ai copilot optimization.

Learning from Active Human Involvement through Proxy Value Propagation Efficient learning of safe driving policy via human-ai copilot optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.849745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.229681Z digest=sha256:ec8ea98e72d32755fd2bf0bd38cdd44d118935a62cfec462c1538426d5cca852

Observation 6b5d3aa1-10ac-46a2-adf8-93d386449437 · outbound

This paper cites Interactive learning from policy-dependent human feedback.

Learning from Active Human Involvement through Proxy Value Propagation Interactive learning from policy-dependent human feedback

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.839140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.233935Z digest=sha256:12ac83b1169929ba4fb4ec53fe0b79e55ce90da8bc2a197119b027021d69fc10

Observation c17b998b-b5a3-4ca8-bef6-efdf4058b2bf · outbound

This paper cites Where to add actions in human-in- the-loop reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Where to add actions in human-in- the-loop reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.829020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.237652Z digest=sha256:02b9430a44a7e2297602188d6cf05e6e750083bb7ccf9f319375ae7ae8254197

Observation 1010a489-38aa-4c5b-b1b4-c5e7c20ef80b · outbound

This paper cites Human-in-the-Loop Imitation Learning using Remote Teleoperation.

Learning from Active Human Involvement through Proxy Value Propagation Human-in-the-Loop Imitation Learning using Remote Teleoperation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.241309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.241309Z digest=sha256:085d7b01fee067f3c629dd0204e1f3d16a0ed8b0b0ee2d28b56fb2bc858fbe39

Observation 60815919-ae78-415b-9ebc-0f7f148902dc · outbound

This paper cites Ensembledagger: A bayesian approach to safe imitation learning.

Learning from Active Human Involvement through Proxy Value Propagation Ensembledagger: A bayesian approach to safe imitation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.818504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.245371Z digest=sha256:6a18c20ef9943b311dec7fb5555b42b29109238bdb48ab866ae38ac6bf6b71fd

Observation ced1ce6b-ef94-4b41-b8bb-c876ef40c9b6 · outbound

This paper cites Human-level control through deep reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Human-level control through deep reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.249898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.249898Z digest=sha256:69d74b40898f1150bf1a82de2fe90f6b3f33264b4005d3300d8674654975b54a

Observation 5e3bba32-8576-4491-8fe4-a9c3e45548cb · outbound

This paper cites Interactively shaping robot behaviour with unlabeled human instructions.

Learning from Active Human Involvement through Proxy Value Propagation Interactively shaping robot behaviour with unlabeled human instructions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.802496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.253953Z digest=sha256:3e1683ff16b5aaaabee06df47a892a5685d984a51e45d468140f68db8e13f634

Observation 485f51ee-91ce-4237-afda-1ca95d0b9bde · outbound

This paper cites Deep exploration via bootstrapped DQN.

Learning from Active Human Involvement through Proxy Value Propagation Deep exploration via bootstrapped DQN

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.792031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.257745Z digest=sha256:73bc4e4a426f638a6d9c74a1f569fd4b5a374ef8517f958c0355829f45b9b42b

Observation ee7e1128-c1dd-45f1-be1b-0f85cc47d448 · outbound

This paper cites Training language models to follow instructions with human feedback.

Learning from Active Human Involvement through Proxy Value Propagation Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.261475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.261475Z digest=sha256:8483047dfc6ae25d25a55d172d82f4392952b601a49c30053d7ca92a6c5e4559

Observation 2f37cc8a-19ae-4da6-b640-437947344033 · outbound

This paper cites Deeptake: Prediction of driver takeover behavior using multimodal data.

Learning from Active Human Involvement through Proxy Value Propagation Deeptake: Prediction of driver takeover behavior using multimodal data

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.782039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.266006Z digest=sha256:f940004692ec7a266eb3c766a19559e7e3d057dda4a04adb8c9ec4493f3cd947

Observation d77a5421-8af5-435e-a7f1-2d5400274dd5 · outbound

This paper cites Learning reward functions by integrating human demonstrations and preferences.

Learning from Active Human Involvement through Proxy Value Propagation Learning reward functions by integrating human demonstrations and preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.771865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.269625Z digest=sha256:9e10dfd0ed5cbd38c2f203ab8570480423f1a4a030c6f8387a87c7aa27584005

Observation 3504d69c-1fdf-4a9a-b3b2-fb81bd3064b2 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Learning from Active Human Involvement through Proxy Value Propagation Stable-baselines3: Reliable reinforcement learning implementations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.761852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.273446Z digest=sha256:1e41fd02300211c6d3eb2ece59ec2c226b8ff0a5ed1c9dd2d535cbe2a4bf746a

Observation a4aa344c-5a33-4a04-b2c4-819b55859e9f · outbound

This paper cites Shared autonomy via deep reinforcement learning.

Learning from Active Human Involvement through Proxy Value Propagation Shared autonomy via deep reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.751260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.277130Z digest=sha256:32237a2a551ed36edc15cf82ba612d161223d076455878eab6e701df8a71342a

Observation 2a7b02cd-c928-4ab8-8a6f-e684a891d10e · outbound

This paper cites Efficient reductions for imitation learning.

Learning from Active Human Involvement through Proxy Value Propagation Efficient reductions for imitation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.741060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.280788Z digest=sha256:c42e6ab257a3784ebde36ff6927ae2fb8cb2815510d775c5fd22aff4212b41e9

Observation 512fd1f9-fd5d-493b-8cf5-ce44b21155d6 · outbound

This paper cites Human compatible: Artificial intelligence and the problem of control.

Learning from Active Human Involvement through Proxy Value Propagation Human compatible: Artificial intelligence and the problem of control

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.730757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.284438Z digest=sha256:4d549e8f819d505178824d14ed8c82bc2aa6b053fdd8d3ba092084044210583c

Observation ef5d8521-5e86-45ba-8c00-014e4e319f1a · outbound

This paper cites Active preference-based learning of reward functions.

Learning from Active Human Involvement through Proxy Value Propagation Active preference-based learning of reward functions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.720405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.287996Z digest=sha256:73c1b04af7f2e9ed64d9de30a6c41a72fe2d278d6c035f04530b585f4e3be0ab

Observation 43b7a1ab-eb22-4dd8-92f6-0e5c6ad147e6 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Learning from Active Human Involvement through Proxy Value Propagation The StarCraft Multi-Agent Challenge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.291676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.291676Z digest=sha256:a04469d51759b2e5a1b8dda5d1e9df150e3ca0ef47409b52ec47e468f57e7335

Observation b959a36c-801b-41d3-becb-2a8a2b0779bf · outbound

This paper cites Trial without error: Towards safe reinforcement learning via human intervention.

Learning from Active Human Involvement through Proxy Value Propagation Trial without error: Towards safe reinforcement learning via human intervention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.710352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.296095Z digest=sha256:d11ca7d6b57ec7f91072c1ca1295e8439503a513a56386034cd222101a137fd9

Observation a52999c7-3485-4783-b7a0-55bf2061e842 · outbound

This paper cites Prioritized Experience Replay.

Learning from Active Human Involvement through Proxy Value Propagation Prioritized Experience Replay

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.299795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.299795Z digest=sha256:eb640bf4a5e436b9adbd30c9142db019d625e41a49011f4dc98f88cb7af8dc16

Observation 94d15502-5157-47ab-8d20-976a467ab046 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Active Human Involvement through Proxy Value Propagation Proximal Policy Optimization Algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.303556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.303556Z digest=sha256:e5954aa62c80754032be2321502b9fede380b8cda80a14cb9d0f666d0c27c558

Observation 0398c6e5-90a9-4f6b-a123-049b82a417d6 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Learning from Active Human Involvement through Proxy Value Propagation Mastering the game of go with deep neural networks and tree search

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.306927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.306927Z digest=sha256:bef1d8884bd3518f127b7a33cb9e0ecb70ecb66024b24a47fd788d96b471aa25

Observation 7fd6d2e5-f549-4dab-a89d-eb87b501cb25 · outbound

This paper cites Learning from interventions.

Learning from Active Human Involvement through Proxy Value Propagation Learning from interventions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.694734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.310183Z digest=sha256:18eef9627c0f678ffe0609e43eabdf23e8d004ff0e0dd90fe06260dc1e9534d1

Observation c3b36b0b-6a65-4bec-bebb-593a2260a381 · outbound

This paper cites Responsive safety in reinforcement learning by PID lagrangian methods.

Learning from Active Human Involvement through Proxy Value Propagation Responsive safety in reinforcement learning by PID lagrangian methods

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.685480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.313523Z digest=sha256:790812704c2449a8b9578ff41f484791f2024c3364c5f5cadc4663b37e94de9a

Observation 16f063e0-5ee7-4fe0-a36f-97e2e770258a · outbound

This paper cites Intervention aided reinforcement learning for safe and practical policy optimization in navigation.

Learning from Active Human Involvement through Proxy Value Propagation Intervention aided reinforcement learning for safe and practical policy optimization in navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.675802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.316900Z digest=sha256:d6275452c00a1aa92ed04102ffaa96c4c85d26e9ef56374924214435b3d76356

Observation 1f9894bd-0bf9-4556-8a26-6636f5946c53 · outbound

This paper cites Appli: Adaptive planner parameter learning from interventions.

Learning from Active Human Involvement through Proxy Value Propagation Appli: Adaptive planner parameter learning from interventions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.665765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.320318Z digest=sha256:90b09bd6865cb8d46f4c15bd749e6279b6d9f9b89a2a82b640cb81a1a5754071

Observation 478fe29a-200e-4244-aaec-9c1429f8ed03 · outbound

This paper cites Apple: Adaptive planner parameter learning from evaluative feedback.

Learning from Active Human Involvement through Proxy Value Propagation Apple: Adaptive planner parameter learning from evaluative feedback

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.655365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.323639Z digest=sha256:d11ad075e79767b96bbfba6bd91e3214d984a60effdae2be8d7e6539131fab31

Observation f6731b1a-369e-46d1-ab34-e27019e4ca7c · outbound

This paper cites Waytowich, Vernon Lawhern, and Peter Stone.

Learning from Active Human Involvement through Proxy Value Propagation Waytowich, Vernon Lawhern, and Peter Stone

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.644940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.326895Z digest=sha256:88569412520bcabd061747bed7b38892fdc302f335e4a0ee9649d8da4cb58356

Observation 9efe55a1-6684-4c74-95cc-3541b6f622e2 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Learning from Active Human Involvement through Proxy Value Propagation A survey of preference-based reinforcement learning methods

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.635146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.330249Z digest=sha256:2938cb3257753bfe5dc34172d6b27382133296f691d56a81992d5af4681aac7c

Observation fff5f2f2-54f5-4012-8fb2-f5dd6ee30d87 · outbound

This paper cites Look before you leap: Safe model-based reinforcement learning with human intervention.

Learning from Active Human Involvement through Proxy Value Propagation Look before you leap: Safe model-based reinforcement learning with human intervention

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.624382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.333912Z digest=sha256:bb20f9bb476a3e06db44fe63531dfc9b16d5ddd1ac858cfa6fb271139039cdd7

Observation 4a6b2247-a7b6-4c53-8a33-62eec05ac09a · outbound

This paper cites How to Leverage Unlabeled Data in Offline Reinforcement Learning.

Learning from Active Human Involvement through Proxy Value Propagation How to Leverage Unlabeled Data in Offline Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.337362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.337362Z digest=sha256:aee51a44651b21a865727d05918dbd8241a2a0b5d4fd3c159271bd9dddb02d99

Observation 1cba8519-29cb-44e7-beb5-2e893ce4d370 · outbound

This paper cites Query-efficient imitation learning for end-to-end simulated driving.

Learning from Active Human Involvement through Proxy Value Propagation Query-efficient imitation learning for end-to-end simulated driving

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.613347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.341319Z digest=sha256:b6c7afa88f9e213508263154c99273b4700d323990b5d0868f41a293c62b8845

Observation e950529c-f767-45f5-9422-d19e5ac28403 · outbound

This paper cites We also compare the behavior of agents learned from PVP and TD3 baseline.

Learning from Active Human Involvement through Proxy Value Propagation We also compare the behavior of agents learned from PVP and TD3 baseline

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.603389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.345404Z digest=sha256:ffb03eb3a8b4705195d91524aa196f3203e9ef382ab798653b5c394ccc8a5d16

Observation 6068c2d7-ed81-48b2-89b1-a6aefef1393f · outbound

This paper cites We present the behavior comparison between PVP and TD3 baseline.

Learning from Active Human Involvement through Proxy Value Propagation We present the behavior comparison between PVP and TD3 baseline

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.592962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.349364Z digest=sha256:419a6642117b29dbd0b7a3c96156353e3a555dc00db2870e316d1d617954a11c

Observation 9bf4e8f1-30c6-4542-af09-7e41c28670bf · outbound

This paper cites PVP performs well in GTA V and can drive smoothly on the highway.

Learning from Active Human Involvement through Proxy Value Propagation PVP performs well in GTA V and can drive smoothly on the highway

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.581248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.353351Z digest=sha256:9e0ad117eedf84cc45c538df45933edde25e135c69da3426de1428789c6845fb

Observation c5bcf62e-6cee-4520-b026-ec406bf02213 · outbound

This paper cites If the agent drives in the wrong way then the displace- ment reward will be multiplied by −1.

Learning from Active Human Involvement through Proxy Value Propagation If the agent drives in the wrong way then the displace- ment reward will be multiplied by −1

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.570786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.357519Z digest=sha256:aa85bbe9d5e7d2fa4ab6bb864520150f600b4e80e652e455abae617ac0b6e8d2

Observation c52ab57e-f03e-4c80-a06c-8884e315b45d · outbound

This paper cites If the agent drives in wrong way then the speed reward will be multiplied by −1.

Learning from Active Human Involvement through Proxy Value Propagation If the agent drives in wrong way then the speed reward will be multiplied by −1

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.559970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.361332Z digest=sha256:354be659ee9afa04b60264b22c865071868c4148b5cab836e8442dc3c7f94301

Observation 17d20be8-c420-46e8-acdc-458c05267c24 · outbound

This paper cites Otherwise, it is 0.

Learning from Active Human Involvement through Proxy Value Propagation Otherwise, it is 0

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:13.549236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.365174Z digest=sha256:798de1d9912f3c511625359c172df7e78faa07eec59402563b3b8ae96c17b698

Observation f0fc79af-f7b3-490a-81a4-4890185e7071 · outbound

This paper cites At that step, we set Rdisp = Rspeed = Rcollision = 0and assign Rterm according to the terminal state.

Learning from Active Human Involvement through Proxy Value Propagation At that step, we set Rdisp = Rspeed = Rcollision = 0and assign Rterm according to the terminal state

Reference 63

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T05:02:13.537874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T05:02:13.369371Z digest=sha256:d53947a06e81c58d430595ad3ceca117d0eb99b69f58b4eb0f194320333626c1

Pith citing papers

No inbound Pith citation observations are available.