Pith. sign in

Paper Citation Record · LEDGER

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.12095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12095 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:30.592023Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e260941e-9d6f-4510-b2ae-7ccda0e66247 · outbound

This paper cites Learning humanoid locomotion with transformers,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning humanoid locomotion with transformers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.278036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.499106Z digest=sha256:d6796d1328b0cb8ff5ebb6b2963f3277ba3cdee0a045e488e9c929c7b91ff673

Observation d7889185-0ee8-4d79-918b-ee4f71d97b26 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Real-world humanoid locomotion with reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.115658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.537410Z digest=sha256:4872323ad44e9df65b7b0de00dfb7e24f027c68d25f8b6e402b5855cd0684e0d

Observation 52a8d7ac-5b49-41e5-8d74-14e8d27525b2 · outbound

This paper cites Model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model predictive control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.932501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.589024Z digest=sha256:ab4846661e0b3b93ac954ef6dd6974df0d4723c52f7ef79af466dbb16273b32b

Observation f4bab07a-8f8e-4064-9cf6-977f1283d99c · outbound

This paper cites Aleatoric and epistemic uncertainty with random forests,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty with random forests,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.656935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.628274Z digest=sha256:c1d9a29aeec04d44ba97ad71a62bb8746926a73ddcaa326d3946c320f5749ecb

Observation 87224d6d-60e6-4b83-a8dd-19afe8e017c6 · outbound

This paper cites Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.696736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.696736Z digest=sha256:ff3a88f7bd712374fd27a0af0496436f8df0eaa4a57680e97f5f3b63a70ee144

Observation 4d584cc8-db7c-4a6e-81f4-d036127ff71f · outbound

This paper cites Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.960812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.748895Z digest=sha256:305e4ad681a6e15b6d16fa09b64927ebe16ae34aa1eb287029a617f31675c8c9

Observation 8b5c315c-d480-435d-84ac-a50798c6409d · outbound

This paper cites Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.909236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.806038Z digest=sha256:8ac012c01ee3c0235016a25d708f115e74b5160bccc5147272506ab71829ed93

Observation 878a5feb-84b3-4579-beea-fa4067b7500c · outbound

This paper cites Temporal difference learning for model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Temporal difference learning for model predictive control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.890578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.890578Z digest=sha256:a2ebebaf9b49c56a55e5314ce60ec35d734cdc05b66e2357475747d6fae3b4b7

Observation 0e2d9b79-dbe9-4b1f-9065-b6a28d55aaa9 · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Td-mpc2: Scalable, robust world models for continuous control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.541850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:29.976944Z digest=sha256:30d9d97c4bebc01dd29f4cf83ae58cc59341a334fb29f1db9416da716be2417c

Observation 56bbb6f7-6191-4408-94d9-414ebade58b5 · outbound

This paper cites Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.489789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.035313Z digest=sha256:60fb88082fda75267a93960e5bc8c6f2e48f511588eecc31995e142bc91b4085

Observation 70dd384c-1d61-43f9-ac3f-887c72d72b08 · outbound

This paper cites Model-Based Offline Planning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model-Based Offline Planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.087984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.087984Z digest=sha256:cb6abbcaccba7a9411887c16c486bbc03d366aa26abe3121f93e43f4a6bfd137

Observation a772cadb-ca4b-44c8-844c-52e155ff9de9 · outbound

This paper cites TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.170831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.170831Z digest=sha256:8316e64ce15809e6131f97630ec857ce67b35eac89cb1eebf70db1cd1894f827

Observation 986fb49b-5f58-42ba-9610-d3c7d1488490 · outbound

This paper cites V ovk, A.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion V ovk, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.198273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.198273Z digest=sha256:c7999b4652746e97d5e839e926ba7e92fb4a674b76de5d9665e90fa8f70981c9

Observation 6cdcabf3-1b02-4939-94c8-9de1834b17a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.305025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.305025Z digest=sha256:19aa673c46298f33b998227629775fbd2e03383654bc8e6c22744d698b8fee82

Observation 8b14b05d-c8fa-4c20-b22a-d15cb39aa7d0 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.356457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.356457Z digest=sha256:e91904f9c2b08b76b3a5dcec051b86f9830412800a33b9ec3c568ec935ef2291

Observation c337b9cf-f2b7-4f20-9d21-9d6a3f968b11 · outbound

This paper cites Reinforcement learning for humanoid robotics,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Reinforcement learning for humanoid robotics,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.425332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.415385Z digest=sha256:23c24bd3833cda0575494ff677b37d62be5242bf4eb09b30e8d9196c45a63e3e

Observation 11b382bb-3341-4f1c-90fb-c1db1f7a0b97 · outbound

This paper cites Learning off-policy with online planning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning off-policy with online planning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.380410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.423302Z digest=sha256:1d5824b4220ac2c0b50ee087645094da005354294b649884813207de6859d36a

Observation f123cf6a-1eea-4d54-8c5a-9b4320964f46 · outbound

This paper cites Conformal prediction in manifold learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction in manifold learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.357234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.428872Z digest=sha256:76b5d0a411bd334fe2f7bdfe8764d88800104a5c24ee0637e0da1299a19024b0

Observation 614f192f-7e1f-4661-83c0-acde746bf96c · outbound

This paper cites Conformal Prediction with Learned Features.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal Prediction with Learned Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.435680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.435680Z digest=sha256:5bed598146e044eb76c7f7a43eed3434bd038571f2ddfffadc17b968cc92d9c2

Observation 2d667aa7-e41b-4da9-884c-c297c762fdd9 · outbound

This paper cites Confor- mal prediction for semantically-aware autonomous perception in urban environments,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Confor- mal prediction for semantically-aware autonomous perception in urban environments,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.338215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.445140Z digest=sha256:0698d7369ddcf7b739be7663c0f2dd5bbc9fc2154bc02646ff39c47d1430ead5

Observation c3a52c1e-a316-4937-83dc-e277ce094e56 · outbound

This paper cites Conformal prediction for uncertainty-aware planning with diffusion dynamics model,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction for uncertainty-aware planning with diffusion dynamics model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.307658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.452592Z digest=sha256:3af2becf8e0bc1c9da3479e52e8ed4d041241fff5cd856acc3bca9a3ea626174

Observation 07eba7e6-feff-4612-8708-cba7edc3e584 · outbound

This paper cites Adaptive conformal prediction for motion planning among dynamic agents,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Adaptive conformal prediction for motion planning among dynamic agents,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.286345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.459658Z digest=sha256:6c23f1fad795594b43f007724f0ab0c80ae6122945e663bb07dca18f937c5011

Observation 82632f0d-106c-4c04-aac0-bbe697569ced · outbound

This paper cites Safe planning in dynamic environments using conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe planning in dynamic environments using conformal prediction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.267133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.467615Z digest=sha256:929218c6d6ff466ed80036e0ec6a787f998dbe2ab501a5ff81ff17552d96e718

Observation 6f144c20-bd22-4315-af55-7676327a4b1d · outbound

This paper cites Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.246167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.473135Z digest=sha256:c8967df33b78229e9c25daef83770cb9dd7520625c53dd426ce4b41aa5032bf6

Observation 94498e85-f25f-4691-9792-b54b2eee61c2 · outbound

This paper cites Conformal decision theory: Safe autonomous decisions from imperfect predictions,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal decision theory: Safe autonomous decisions from imperfect predictions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.229092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.480932Z digest=sha256:c913a7c18ee07d874c22166a1e3d7723693f267080790e4b6a7955e576b5347c

Observation 668e69c2-056f-45a2-9bbf-c26ed3c50442 · outbound

This paper cites Safe pomdp online planning among dynamic agents via adaptive conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe pomdp online planning among dynamic agents via adaptive conformal prediction,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.206160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.486796Z digest=sha256:befa97168554d40eebd768218fa1eb81023f89685034b56f7d8f1aa2ed267c02

Observation 9982d18c-a52c-4002-948c-34f7c9562c1e · outbound

This paper cites Conformal policy learning for sensorimotor control under distribution shifts,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal policy learning for sensorimotor control under distribution shifts,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.180556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.493848Z digest=sha256:924308c62de6d64bf00bd02be047475f07e9f4941c56a03da14ae01f779055ef

Observation 45659071-8e67-416d-9c9d-f0cb3d6260ba · outbound

This paper cites Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.502164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.502164Z digest=sha256:d0353b708dadda7a7a81c0e129720fd6205eba3bffcfe4292a13372f08e7a10d

Observation 129a9df3-0ca0-4cd3-9f37-ab4523e42d85 · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stabilizing off- policy q-learning via bootstrapping error reduction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.157675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.512999Z digest=sha256:cf123b561f2bc18dcf80e68d7635bcfebd6707b38c6e9aec66059036f2383e62

Observation c8364e28-c826-42d7-b104-9488c5a1d8bd · outbound

This paper cites Conservative q-learning for offline reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conservative q-learning for offline reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.521022Z digest=sha256:0283222d4f1c1d5035640d712e5b9a93722ce3503f33655d81e6fb5a67ee59a8

Observation 03d3fb43-b9bb-442d-b1ed-df58065d65db · outbound

This paper cites A minimalist approach to offline reinforce- ment learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion A minimalist approach to offline reinforce- ment learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.103450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.527463Z digest=sha256:ba7d87bf49aebbc64064f21edaba8d523d0d78275803ff281d06b5c7b33ac11a

Observation f0913dee-c670-42da-9b0f-7160d230e533 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Off-policy deep reinforcement learning without exploration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.079366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.533658Z digest=sha256:140c0b2e7f75f7c88ccc60fa276268e86078f2ae3f02d7ba63a03630420f8f38

Observation f69c47d4-639e-4bb8-a9b2-5ee79a991814 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.542382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.542382Z digest=sha256:f8f0c501a91a8417780625b746424dd917eb4a82815ede7c0389990b19a28195

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:39313c88daec28e0ece5845017e47f3736f9d4f92611570c76376ef39ef866e0

Observation 05760fd6-b9e7-4ee7-9971-fb4e4a86156d · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Offline Reinforcement Learning with Implicit Q-Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.555683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.555683Z digest=sha256:92491d3094a91661221263dbb49207aa18ef957451a349d907cc43dda0187f13

Observation d9f23434-f2ee-4508-8fb0-c5c7e8889472 · outbound

This paper cites Aggressive driving with model predictive path integral control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aggressive driving with model predictive path integral control,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.561578Z digest=sha256:94fa31ec40878e71eacd3223ee609799a6e59f02af07e9d16395cb94528e22d4

Observation 4e8df83b-7ac8-4750-ad88-c1415d4c9e99 · outbound

This paper cites Trust region policy optimization,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Trust region policy optimization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.569861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.569861Z digest=sha256:c09722fcb274eeaef14896ef17b819c664dd27a8d3fc4375c9c07dd217cd6689

Observation dc279f4d-39da-4c85-b56e-c235f4528584 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.019435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:27:30.576382Z digest=sha256:3845253594305390b34e51489e1e139abb07ef82a6bdb448120d37e71fe373fa

Observation f335dd01-54b4-4e0e-88e6-d9ab13ca0f86 · outbound

This paper cites Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.585942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.585942Z digest=sha256:69d9c57004a31290e2240b535b75ef425098ce6357e335ddba4b85a165dbd534

Observation 4a59ac21-e20c-4335-982e-fa8a68de9b1a · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.592023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.592023Z digest=sha256:5e1045be76f1319cb8739443d20950ea9583bb3b6c5af8e460e020e77ff33d3d

Pith citing papers

No inbound Pith citation observations are available.