Pith. sign in

Paper Citation Record · LEDGER

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.12095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12095 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:30.592023Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e260941e-9d6f-4510-b2ae-7ccda0e66247 · outbound

This paper cites Learning humanoid locomotion with transformers,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning humanoid locomotion with transformers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.278036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.499106Z digest=sha256:b68ef6692d95d67871b073bf95103a3ef4899ccfb5ed635fbcce1dda64e93cb1

Observation d7889185-0ee8-4d79-918b-ee4f71d97b26 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Real-world humanoid locomotion with reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.115658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.537410Z digest=sha256:d47a86841d3302b92445bf11a915fd0604ecebe4dafe5d4cc054284995ea4e43

Observation 52a8d7ac-5b49-41e5-8d74-14e8d27525b2 · outbound

This paper cites Model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model predictive control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.932501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.589024Z digest=sha256:2dcdd374eee1c12ea6efedb25ea5ac987e82911931505e8df912fc9481a44c0e

Observation f4bab07a-8f8e-4064-9cf6-977f1283d99c · outbound

This paper cites Aleatoric and epistemic uncertainty with random forests,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty with random forests,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.656935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.628274Z digest=sha256:286cd6894027b9115545a902945833a33c58405ce5191118a36d4c147a5f7d19

Observation 87224d6d-60e6-4b83-a8dd-19afe8e017c6 · outbound

This paper cites Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.696736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.696736Z digest=sha256:60e87beab1c81dd72c9767c2a9d7be438278f35ef6b610e89d68b312a8504412

Observation 4d584cc8-db7c-4a6e-81f4-d036127ff71f · outbound

This paper cites Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.960812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.748895Z digest=sha256:9f21a51d26c290b93fd20feeb75349be0d115964e1346f7e4e2b922290af4b53

Observation 8b5c315c-d480-435d-84ac-a50798c6409d · outbound

This paper cites Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.909236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.806038Z digest=sha256:1375ddb059cfd2f19ea7e376b15151a3ff00e85dbed50a7b1742fc09d12fbcb8

Observation 878a5feb-84b3-4579-beea-fa4067b7500c · outbound

This paper cites Temporal difference learning for model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Temporal difference learning for model predictive control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.890578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.890578Z digest=sha256:009c7ac06b7b4dac143e91c84391aef62aa03dd334bec107d54c1603f59cf303

Observation 0e2d9b79-dbe9-4b1f-9065-b6a28d55aaa9 · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Td-mpc2: Scalable, robust world models for continuous control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.541850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:29.976944Z digest=sha256:eba6373ecc879f947327b49274b222fe5e752632f2b8f1beddf46feef89bebff

Observation 56bbb6f7-6191-4408-94d9-414ebade58b5 · outbound

This paper cites Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.489789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.035313Z digest=sha256:7960be65ccfa80011ffa267b1a8c974d0394888774414a038492f53c08eb8cba

Observation 70dd384c-1d61-43f9-ac3f-887c72d72b08 · outbound

This paper cites Model-Based Offline Planning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model-Based Offline Planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.087984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.087984Z digest=sha256:8d3812dda01e71ae7cbbe9c7edd15114274152a4f07a77c71be7c6a2e012d09a

Observation a772cadb-ca4b-44c8-844c-52e155ff9de9 · outbound

This paper cites TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.170831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.170831Z digest=sha256:4b9c944c9af7e440a7b783ab2f5ad9394c25573870a61f458289fac95a0daf18

Observation 986fb49b-5f58-42ba-9610-d3c7d1488490 · outbound

This paper cites V ovk, A.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion V ovk, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.198273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.198273Z digest=sha256:c7333bdd67a91360c1518244a48118ba99694c6a0bd1ced61a93527adb981b49

Observation 6cdcabf3-1b02-4939-94c8-9de1834b17a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.305025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.305025Z digest=sha256:c48aa7e101264ad0b181f530f3d0ebade1369c4cb9d1b0bf29cd167bced2e96c

Observation 8b14b05d-c8fa-4c20-b22a-d15cb39aa7d0 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.356457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.356457Z digest=sha256:92446f074fa49b3b86cbedd58e34ff0b993ad32d94bc117e4e04453a50f2d7f4

Observation c337b9cf-f2b7-4f20-9d21-9d6a3f968b11 · outbound

This paper cites Reinforcement learning for humanoid robotics,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Reinforcement learning for humanoid robotics,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.425332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.415385Z digest=sha256:04b15e401b42772dccd6d8484769d8a02fd36da1b3ef7a7b6ac153dc9a720dd6

Observation 11b382bb-3341-4f1c-90fb-c1db1f7a0b97 · outbound

This paper cites Learning off-policy with online planning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning off-policy with online planning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.380410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.423302Z digest=sha256:a7a17222b490d25367f764d2c222fdb2b327d755d1faa319f466732366a0460c

Observation f123cf6a-1eea-4d54-8c5a-9b4320964f46 · outbound

This paper cites Conformal prediction in manifold learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction in manifold learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.357234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.428872Z digest=sha256:f96199c4b1b3a7d51da1359bd00d4cfec22f93ae4cab7f1c869c97d8c69ca489

Observation 614f192f-7e1f-4661-83c0-acde746bf96c · outbound

This paper cites Conformal Prediction with Learned Features.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal Prediction with Learned Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.435680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.435680Z digest=sha256:cb12ed07184dc75b54b3f3db4f18cf734a27ec655863930a7ee5661e80905273

Observation 2d667aa7-e41b-4da9-884c-c297c762fdd9 · outbound

This paper cites Confor- mal prediction for semantically-aware autonomous perception in urban environments,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Confor- mal prediction for semantically-aware autonomous perception in urban environments,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.338215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.445140Z digest=sha256:9c68c85f0557a5b01a4189c8ce4cbc50c07603a006f0945b332982e6f1f8ebe1

Observation c3a52c1e-a316-4937-83dc-e277ce094e56 · outbound

This paper cites Conformal prediction for uncertainty-aware planning with diffusion dynamics model,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction for uncertainty-aware planning with diffusion dynamics model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.307658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.452592Z digest=sha256:c26ea5eb2f7ade7ae96276440b79681541789ba0723af3995847b9594367b32f

Observation 07eba7e6-feff-4612-8708-cba7edc3e584 · outbound

This paper cites Adaptive conformal prediction for motion planning among dynamic agents,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Adaptive conformal prediction for motion planning among dynamic agents,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.286345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.459658Z digest=sha256:736095fa2e107cd2b09f796a9a8f3dbd74ca89c73fbfb011177e25f59b5425c4

Observation 82632f0d-106c-4c04-aac0-bbe697569ced · outbound

This paper cites Safe planning in dynamic environments using conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe planning in dynamic environments using conformal prediction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.267133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.467615Z digest=sha256:a597b643226a4a7278a62ebee94cbf5e33d9d1cab1dd045d321321fb0b5a7e3a

Observation 6f144c20-bd22-4315-af55-7676327a4b1d · outbound

This paper cites Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.246167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.473135Z digest=sha256:5d032142dfa5c300e34dc24ca671d2b83c51eab80bc0a73dcab39d01e95ee086

Observation 94498e85-f25f-4691-9792-b54b2eee61c2 · outbound

This paper cites Conformal decision theory: Safe autonomous decisions from imperfect predictions,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal decision theory: Safe autonomous decisions from imperfect predictions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.229092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.480932Z digest=sha256:0ad98288e3662082e6ca7a9bd3fd91240ddc5cdc2dbb9a5f4297bddc7e19163f

Observation 668e69c2-056f-45a2-9bbf-c26ed3c50442 · outbound

This paper cites Safe pomdp online planning among dynamic agents via adaptive conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe pomdp online planning among dynamic agents via adaptive conformal prediction,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.206160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.486796Z digest=sha256:baab478674e5e833ca7fa830c1ca7479b762eee2b8f6ab34c7151bf424163bb7

Observation 9982d18c-a52c-4002-948c-34f7c9562c1e · outbound

This paper cites Conformal policy learning for sensorimotor control under distribution shifts,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal policy learning for sensorimotor control under distribution shifts,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.180556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.493848Z digest=sha256:ead26a4d86bdb29484e2698c80ca3ed4c4191db8fc4a7e64753071351bb2f9c1

Observation 45659071-8e67-416d-9c9d-f0cb3d6260ba · outbound

This paper cites Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.502164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.502164Z digest=sha256:a5ae9c374ac12ae8771c405aaeaf45a3c85826814db88768ceafda6c159d0d76

Observation 129a9df3-0ca0-4cd3-9f37-ab4523e42d85 · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stabilizing off- policy q-learning via bootstrapping error reduction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.157675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.512999Z digest=sha256:81eb1d027aaf8dc72049b2a90b15c6115dcba5d179b1e0153e20f8795b950a1c

Observation c8364e28-c826-42d7-b104-9488c5a1d8bd · outbound

This paper cites Conservative q-learning for offline reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conservative q-learning for offline reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.521022Z digest=sha256:aa192df2cef3b4b00823f324d1685032f9075ce01402129f5d4fadecd6bc80af

Observation 03d3fb43-b9bb-442d-b1ed-df58065d65db · outbound

This paper cites A minimalist approach to offline reinforce- ment learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion A minimalist approach to offline reinforce- ment learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.103450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.527463Z digest=sha256:492897bfecfabbfb050b1622f21e3aa429973f213a44cdb9f60491887cb645e8

Observation f0913dee-c670-42da-9b0f-7160d230e533 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Off-policy deep reinforcement learning without exploration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.079366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.533658Z digest=sha256:12d42e4af6d949f7215d0b9ab07991fc21a3509eab19400705b334b84d9f04df

Observation f69c47d4-639e-4bb8-a9b2-5ee79a991814 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.542382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.542382Z digest=sha256:dc3ffafd5e5b69fb039d5aa6c81e424644b614b2c89f5c3a001c5758ff9c0e2c

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:faa91d5f89d685878da10ced71f24c1c5db102d9e85b4dc8260c4c392c1060fe

Observation 05760fd6-b9e7-4ee7-9971-fb4e4a86156d · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Offline Reinforcement Learning with Implicit Q-Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.555683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.555683Z digest=sha256:12a4ff522779019e64505e4662a76f37d796aca1066e2ca8085765419fe3f72b

Observation d9f23434-f2ee-4508-8fb0-c5c7e8889472 · outbound

This paper cites Aggressive driving with model predictive path integral control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aggressive driving with model predictive path integral control,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.561578Z digest=sha256:4cf8a58f97214447a268adcda184c4e0d5fc2dbf99598d45dec6f3178b29784a

Observation 4e8df83b-7ac8-4750-ad88-c1415d4c9e99 · outbound

This paper cites Trust region policy optimization,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Trust region policy optimization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.569861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.569861Z digest=sha256:ee04579e3e3e2ea51a870f23afa33bccca67bdc78fb3b5dc5b71f93b639e9ad7

Observation dc279f4d-39da-4c85-b56e-c235f4528584 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.019435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:27:30.576382Z digest=sha256:48df5d2694d3b4c5d9c173f78a4a497ade1ad4f1df925255e31aa0adcccdc348

Observation f335dd01-54b4-4e0e-88e6-d9ab13ca0f86 · outbound

This paper cites Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.585942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.585942Z digest=sha256:f46dba39fdb35c75f14f3851fb65326e19c7641155a605e00ff303565cda5555

Observation 4a59ac21-e20c-4335-982e-fa8a68de9b1a · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.592023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.592023Z digest=sha256:6f4459fc86fc1ed04f6ef1dee6bc4935a262ba6fedab3060af8d09571128c78a

Pith citing papers

No inbound Pith citation observations are available.