Pith. sign in

Paper Citation Record · LEDGER

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving

As of 13 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.03568.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03568 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:47.517274Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b953d159-f33d-4b89-859a-3537b2c5e1ab · outbound

This paper cites Learning to drive in a day,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to drive in a day,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:49.081032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:46.996668Z digest=sha256:33d113eb3276099b4c279de7cf49f09a94d5396ce9536af57dbaa8f9877342c7

Observation ea909278-58b2-43ac-9010-e526a626c62c · outbound

This paper cites End to End Learning for Self-Driving Cars.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End to End Learning for Self-Driving Cars

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.054690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.054690Z digest=sha256:135c09452ba05349062381920a1172d50354445e23442b13540bfcba26b683ab

Observation 26b284c1-6c32-4c92-bd57-e8d5f8d7c391 · outbound

This paper cites Dense reinforcement learning for safety validation of autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Dense reinforcement learning for safety validation of autonomous vehicles,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.062703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.062703Z digest=sha256:5005fe0b3a0032c97977b8a494fc20b3e153530a2e882678cf9b9a19ada9b9da

Observation b9599d1f-fd6d-4551-a9b6-ad054c3bd309 · outbound

This paper cites A survey of deep RL and IL for autonomous driving policy learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey of deep RL and IL for autonomous driving policy learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:49.016724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.074645Z digest=sha256:bb6a616eb70e3ad6c420dc99a86fd3b2343ebbcc797b92a5f54dfd876e5754d9

Observation a10452b6-967f-4279-b900-5e0790312f10 · outbound

This paper cites A survey on autonomous vehicle control in the era of mixed-autonomy: From physics-based to ai-guided driving policy learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on autonomous vehicle control in the era of mixed-autonomy: From physics-based to ai-guided driving policy learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.994054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.080690Z digest=sha256:fde05c09a55dea90fe6984b88ed0f7c6992019a7b081edb641c30ad07edfb10a

Observation d6c01a77-4c2a-4f80-91d0-67b85257cbb5 · outbound

This paper cites Deep learning for safe autonomous driving: Current chal- lenges and future directions,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep learning for safe autonomous driving: Current chal- lenges and future directions,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.087129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.087129Z digest=sha256:f355fd92b568351842946a67fcee48a8cbc65709d7f71610f8b5ee04a7c0e492

Observation 70e94c2b-9b6d-496f-9258-84821eaebf43 · outbound

This paper cites Survey of deep reinforcement learning for motion planning of autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Survey of deep reinforcement learning for motion planning of autonomous vehicles,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.092742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.092742Z digest=sha256:8b33b3bc0a6205932d93c5a41b49aa52f43c6b013a14ec109f1fe7c9788a7610

Observation 33fc5d1f-f2a8-4f43-b19f-8a3c79b377b0 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.099754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.099754Z digest=sha256:44690878faadba8a221a3f547a9afdccf25169db609fe9e465d627901a31a271

Observation fbac7cfb-67d7-4cd9-9adb-34ae7100f6ce · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning for autonomous driving: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.938561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.108862Z digest=sha256:ffb274500ea82c01e8cb56699f3e8262d49ae8b5659da5544b5ece9984dc99c4

Observation 2a382cf7-ee65-43e4-a623-303b4ec4c4fa · outbound

This paper cites End-to-end urban driving by imitating a reinforcement learning coach,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End-to-end urban driving by imitating a reinforcement learning coach,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.913663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.119541Z digest=sha256:3259f04896bd9dbc81e8579f7a24d68eb0095d81b9a40ff2a2eca0a256de64de

Observation 69b0d92e-7f70-4794-a84a-fc54ae364c6a · outbound

This paper cites Reward misdesign for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Reward misdesign for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.887744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.129878Z digest=sha256:05e5291d1c21107e3a26baea8be937a9c80d1f3065bb1bb3d44c2f2f2d30f595

Observation 5d2bb129-4ecb-4225-92d3-85b487ae0178 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.141529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.141529Z digest=sha256:1d56d7c6bf656816676551139e5d0eac927b234c1999319d4df8bc0c49966d4e

Observation c5789285-6345-42fd-8fa0-d698171dc58a · outbound

This paper cites Demonstrating specification gaming in reasoning models.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Demonstrating specification gaming in reasoning models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.152726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.152726Z digest=sha256:0fc08f987445b5c9125a0db28ebead701892e2407add03df4823fd5dcad290d5

Observation 3ecdb816-979a-4115-8595-9a98158cc2fd · outbound

This paper cites Toward human-in-the-loop AI: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Toward human-in-the-loop AI: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.861260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.159984Z digest=sha256:6e8e7973f1023e4c28b0bd2150da9ba6685248051d4cf4b0ffdaf2971d9a3807

Observation e200d637-9a49-4c9b-bcf8-f7d3425a8012 · outbound

This paper cites Hindsight credit assignment,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hindsight credit assignment,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.841270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.176862Z digest=sha256:cc56411109e70157c803780ab8670433e166bf809b772029b41aa5973a5a0dc0

Observation 37afddd2-4f33-41ff-8d6a-2707ab0e72bc · outbound

This paper cites Trial without Error: Towards Safe Reinforcement Learning via Human Intervention.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trial without Error: Towards Safe Reinforcement Learning via Human Intervention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.203825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.203825Z digest=sha256:0b422112a64b009a1b073b42bd243bc891f6fdb2c70a9acb96a108012f77c764

Observation 7435c22e-26b8-487a-b920-9a468755d7a3 · outbound

This paper cites A survey on imitation learning techniques for end-to-end autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on imitation learning techniques for end-to-end autonomous vehicles,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.820149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.211989Z digest=sha256:7ba6ca748bab4fd9b3038f955d12abfabefed29f3128aa08e563b1eb98656c9c

Observation 43b0205f-a0b0-4a5c-84ba-28463da293fb · outbound

This paper cites Conditional predictive behavior planning with inverse reinforcement learning for human-like autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conditional predictive behavior planning with inverse reinforcement learning for human-like autonomous driving,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.790007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.224555Z digest=sha256:03cf757dd3c56ef9352b9172386f961beb020ad88963ff2c327eef98e5a6416e

Observation 52d91390-1107-4988-9158-bdb23e1d15f6 · outbound

This paper cites Pattern recognition and adaptive control,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Pattern recognition and adaptive control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.747424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.230363Z digest=sha256:813b80b5aa36e67ba3be9d886cecf2fd0735cd2faa2eae26592a8069371c6f72

Observation a4696d27-913a-475b-a557-d457f26f83c7 · outbound

This paper cites An Algorithmic Perspective on Imitation Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving An Algorithmic Perspective on Imitation Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.235938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.235938Z digest=sha256:5967683572c6078306c6ccdc150271e63eb342076e151943cec498dd6b0987a2

Observation 2f971b72-1e49-456d-8b9a-ff91feb8d2ef · outbound

This paper cites Learning a decision module by imitating driver’s control behaviors,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning a decision module by imitating driver’s control behaviors,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.726411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.242273Z digest=sha256:de60670fe37087342049833d812a1badbb2584148858379f734290a462855208

Observation 33693b59-16e0-45ff-8b2c-7ff49c725d71 · outbound

This paper cites Conservative Safety Critics for Exploration.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Safety Critics for Exploration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.251327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.251327Z digest=sha256:ca46a104a915d79aece741e89b163d0b2d8af3fd8b8bc58eba84deab9a33748e

Observation 9ab2d0dd-d2b1-4721-b75a-7bfb2264ad49 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavior Regularized Offline Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.259634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.259634Z digest=sha256:0e7b81580bd69ace1f925cbc15a13fca450d4b45310dd319dceb4587eacfdcd2

Observation 7e335325-1715-4ac0-a391-11f2cc86de03 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Off-policy deep reinforcement learning without exploration,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.703985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.267759Z digest=sha256:78505f7a323889f9fa90d450907270a52a6cc7809dbf3d4cfb4541f81e244232

Observation 0098b27e-2e4b-42a2-a36d-324d9d0fd4e7 · outbound

This paper cites Adversarial inverse rein- forcement learning with self-attention dynamics model,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Adversarial inverse rein- forcement learning with self-attention dynamics model,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.675772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.274511Z digest=sha256:f7b2b1e6ee6816f99abd891cff4e49770a5c366cceb38d4c8bdc06007a303286

Observation 9434015b-8421-4144-b37f-18da759de9d6 · outbound

This paper cites Efficient reductions for imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient reductions for imitation learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.648437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.301201Z digest=sha256:4b51bc6045e58efb540802ab11fc19e923c443cd675e33d395d7f95e8c122e18

Observation 4442ef43-f0c1-432b-9059-456c2821fa2c · outbound

This paper cites Exploring the limi- tations of behavior cloning for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Exploring the limi- tations of behavior cloning for autonomous driving,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.625527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.314152Z digest=sha256:ea2b5fc3ce5daa7ef11108374f2eedd35e0be502ba05aada34541ffee0bdef22

Observation 6bf27769-625a-48d1-ab28-945048db412f · outbound

This paper cites Behavioral cloning a correction,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavioral cloning a correction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.590978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.320883Z digest=sha256:dbcd1fd469779d34388333b59fa78522053144af077b3496219c069dc818d366

Observation ec8175b5-24a6-42d4-bdcc-c042ea77d078 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.561472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.328632Z digest=sha256:115a09fca24be121f10d094057c885c80cba5580bb06a4d82e7f12224ee09e6a

Observation 024cbaec-61db-4023-908b-e53548811411 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.533279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.335160Z digest=sha256:8b6d2e6b8c798317cf022c9a5dbefea0d981be1b10049ee0be3204c82706cef1

Observation f1763acf-d4a0-4831-a456-b69e1506e204 · outbound

This paper cites Query-Efficient Imitation Learning for End-to-End Autonomous Driving.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Query-Efficient Imitation Learning for End-to-End Autonomous Driving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.345392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.345392Z digest=sha256:4c441d49da8c6d5872191d41a0d08b6791b14a21a259c44c2d380997364e1512

Observation 70cfac0c-e7e9-4dfc-a259-3032037c6108 · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hg-dagger: Interactive imitation learning with human experts,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.354950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.354950Z digest=sha256:64a69a3379de7570945a17a712be1a3bb21508c5b0dba7ef16e56a25af61d0a9

Observation 8d0258af-51c5-4dba-a66c-2a63b26941c0 · outbound

This paper cites Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.489799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.379002Z digest=sha256:ec2b4f05f4fab1a3628f53193f0565a66ece00ba0fcba012bdbf3d30cabb5199

Observation 9a20d808-d509-43e5-b925-2926ac9b8c0d · outbound

This paper cites Expert intervention learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Expert intervention learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.467741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.384324Z digest=sha256:725f76d78827ad1aa5579f39b3f8f5c50ef0f9a230b9737a65e1a902965ece93

Observation 93b1d758-f991-4c18-92d1-157a51a2a22b · outbound

This paper cites Human-in-the-Loop Imitation Learning using Remote Teleoperation.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human-in-the-Loop Imitation Learning using Remote Teleoperation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.390201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.390201Z digest=sha256:2dcb081e7518761b42b3c7d114cf7148a4432ef3d7944262d1600d1fbb11e5c7

Observation 5cc54a30-2cf5-4e27-bca1-c6b64468a705 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning from human preferences,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.434427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.398761Z digest=sha256:5e6ac02ff42afb4db29990ee585c53beebbf2734fee7729621dd23b099b3ce45

Observation 558d2604-d8a2-45cd-96b4-1858083f53f1 · outbound

This paper cites Batch active preference-based learning of reward functions,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Batch active preference-based learning of reward functions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.407579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.404330Z digest=sha256:cbf93f3dc7267625c6249c43969a0f522b10c81064077468758144670d9d1b50

Observation 64f9f942-7044-4d1b-b85d-3446d11595d6 · outbound

This paper cites Learning Reward Functions by Integrating Human Demonstrations and Preferences.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning Reward Functions by Integrating Human Demonstrations and Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.410249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.410249Z digest=sha256:c0e90d7f3d47d86892c7df89d1618eef7925ad19cbac3579c72d106e15926cd3

Observation 67e09e0c-5b81-46f2-8a1d-49f877c648b0 · outbound

This paper cites Efficient learning of safe driving policy via human-AI copilot optimization,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient learning of safe driving policy via human-AI copilot optimization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.378955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.416144Z digest=sha256:02b556ecc806aa5fa6aae3118e7e3ba236f3573f1f7bec7fab58afe93dc8a8c8

Observation 67936181-90d0-495a-866a-28f330be8d8c · outbound

This paper cites Learning from active human involvement through proxy value propagation,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning from active human involvement through proxy value propagation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.353963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.421979Z digest=sha256:1bb1c55cfc8cb04033a29f7ac8dde2eebf597d83d5bd8d2af1da975f54d0108d

Observation ace12819-7be9-4dbb-8251-53a121e283b2 · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.427864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.427864Z digest=sha256:e6ad6aeb93f1dff66c9d26820c7b7e0affcfffa52d07a2bd6cf75dd4ec2f311a

Observation 355d23c3-b0f1-4007-80e5-e2862f3e8c3b · outbound

This paper cites Socially situated artificial intelligence enables learning from human interaction,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Socially situated artificial intelligence enables learning from human interaction,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.332758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.433857Z digest=sha256:24aa388490a19406e52930df14856953f0693f170144b03beab1d669f21072fc

Observation 268b589d-5665-465a-8013-6846a2e03e11 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.301583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.439626Z digest=sha256:82f9efe5aa12ec4420efae413dc4d6251311be8066d5634bb825fae7fe53e25c

Observation 09d7c262-872f-4a76-b8bb-ab18800494a0 · outbound

This paper cites Human as AI mentor: Enhanced human-in-the-loop reinforcement learning for safe and effi- cient autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human as AI mentor: Enhanced human-in-the-loop reinforcement learning for safe and effi- cient autonomous driving,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.271114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.446612Z digest=sha256:60ba016307010e5d2d972fcf9bf9fc95b135f1b8342ff9523527bf2488dc93e5

Observation dd045378-b146-4e55-9cf1-b5559197a214 · outbound

This paper cites Guarded policy optimization with imperfect online demonstrations,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Guarded policy optimization with imperfect online demonstrations,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.234155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.451753Z digest=sha256:c5e915a2264c1c9390aa2043b13abb118fcfd0ec599eb352c830fcba026775c6

Observation d56724d3-6b63-4773-84a8-f20f074ef99b · outbound

This paper cites Trust region policy optimization,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trust region policy optimization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.209925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.457418Z digest=sha256:ccba4c0366c6198ba332650cab4e8191e897841b55ce5d099c762990148e982e

Observation fbc2e929-024c-40ff-8841-12f2dfc434be · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.176999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.462594Z digest=sha256:60f8706c260d2a50575d9d627077402fbfcf13901c921028c9f0efc64f1ea559

Observation b46c8c8a-366e-447a-bee1-d91116c5d1d4 · outbound

This paper cites Responsive safety in reinforce- ment learning by pid lagrangian methods,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Responsive safety in reinforce- ment learning by pid lagrangian methods,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.150462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.468686Z digest=sha256:d805b35bd62c164149ab42bfeb60b822fe47fb4937cc788bf7ee1c05e3be63a8

Observation 8d55fd2c-4562-4271-8430-f18d391150b2 · outbound

This paper cites Learning to Walk in the Real World with Minimal Human Effort.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to Walk in the Real World with Minimal Human Effort

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.474001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.474001Z digest=sha256:74e66af3cfd367dbb0449ba8cdf0fb3286770ada505b114bc1a3b3713d6f0929

Observation f771110b-c8bf-4dda-8076-11d34b2454a5 · outbound

This paper cites Conservative Q- learning for offline reinforcement learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Q- learning for offline reinforcement learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.125073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.479725Z digest=sha256:acdb79cd3293fe48f4a1c853f81fb468430d95a5e0dd42be560c5bd983a75813

Observation 2b598292-45f2-4e28-9d94-55dfe16f77f7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Proximal Policy Optimization Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.485274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.485274Z digest=sha256:aa5fe42cfec373d8f6cb59abbd6f8e61defa01514e7626360b4a31f183c05f04

Observation 389de93b-200b-4f6b-aa22-ce32ef49726e · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.085564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.497444Z digest=sha256:1f3c788174201d864640060d19f85a7c7e0e4f085ceabc3f126612d8383767b1

Observation 0b2a3f9b-c5c4-4ba3-8bb6-24d538ca433f · outbound

This paper cites A framework for behavioural cloning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A framework for behavioural cloning,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.059247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.504779Z digest=sha256:d9cf5e8b4859f5f0d19e62852f4129d98f4f42c25f11f521104482aaa3de85c8

Observation ee91778d-321a-480a-a9dd-e402f4b66a43 · outbound

This paper cites Generative adversarial imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Generative adversarial imitation learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.037949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.511187Z digest=sha256:f555d908b326885f4c0194549bb8178ebf813b1f5cfe9b7888ef6ebb7ec5cde3

Observation ec87f5b9-e218-4ec9-9066-043b308dffc2 · outbound

This paper cites an unresolved cited work.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:47.999906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:06:47.517274Z digest=sha256:fadfc2f1856668d121ba37c94d3ac40f52505fae8565035113e328de4b91adf6

Pith citing papers

No inbound Pith citation observations are available.