Pith. sign in

Paper Citation Record · LEDGER

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2502.03550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03550 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:38:30.745532Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:30.170831Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:56:27.801679Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f56e1c6f-7fa0-45ef-98e2-c52492be8718 · outbound

This paper cites Model-Based Offline Planning.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Model-Based Offline Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.124877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.124877Z digest=sha256:1fe0e0624c4a5580ea7e5b4fe80d7747c0e46b03fcf3cc0ffd609ab5f7239f28

Observation c2bbd487-d4d9-40bd-8fdc-7979a37f7ef1 · outbound

This paper cites Bertsekas.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Bertsekas

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.996494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.192382Z digest=sha256:3875d4d664a79e5c29c3c282cd6fb43532f94dbb2ec8e6d0aff65f58118eb317

Observation 519b3369-66e9-4e7d-bf7e-dffe0a053f9c · outbound

This paper cites Bhardwaj, A.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Bhardwaj, A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.980648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.197110Z digest=sha256:31e33785ff9678475f1d0ebccdede1fe0c677b72ccc7d68c9ef0af8a1098e6eb

Observation 08590fe3-6236-479f-a724-0f7540a02bd6 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:38:31.964782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.202166Z digest=sha256:c7a8a832391d5356b9146497dd61215246d05b7a94e7cdadebcc00a301928781

Observation cabca893-2be3-4a28-9cfc-4f077aa2c774 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.207033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.207033Z digest=sha256:81663b5b19d81109c5354c3cd5cc6724009ed1a29c7084479a302f1d25244a3c

Observation 5a428069-57d7-4dc9-b4a3-e2abcd1a6b02 · outbound

This paper cites Fujimoto and S.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Fujimoto and S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.919520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.211877Z digest=sha256:aa443d513894511136a3f10766190e0c029b87e5f1d22a6b10e4aa7d228dba6e

Observation 7442f3a2-d4a2-4579-9e17-3af9a8ab19eb · outbound

This paper cites Fujimoto, H.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Fujimoto, H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.829642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.218712Z digest=sha256:ff4c342d9edb54a7952550dfd270fc31f3981e73c16f209b1c7f82fd3bc8ebe4

Observation 8feb9364-cea6-4e24-bf8e-c535ff65f48e · outbound

This paper cites Fujimoto, D.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Fujimoto, D

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.747103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.223435Z digest=sha256:3b57b7cdd77e969010bc000ee3453b0f1f69079bc728c2043677cab6973bb816

Observation 8b590690-0025-41e2-9905-07b1a5960926 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Extreme Q-Learning: MaxEnt RL without Entropy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.228255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.228255Z digest=sha256:efa761ff150e9d7a1f3c34918fcf735c294e1084a7a5fc3b77d7b9873dba8a0f

Observation e5977cab-fc46-48e4-b042-c3811fc215c3 · outbound

This paper cites Haarnoja, A.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Haarnoja, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.675738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.232877Z digest=sha256:c71d10cbd4790aa40a916e438c43488ccd808f9ef0789090b44359d395591492

Observation 70b6a54d-f06b-48c5-ad77-27857d432087 · outbound

This paper cites Hafner, T.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Hafner, T

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.237424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.237424Z digest=sha256:32fd98c19d20a613558158c8bc49c6494283cb1f4bd2760c139f3fec78405d82

Observation a660f189-2834-4629-9f96-021408a568d7 · outbound

This paper cites Mastering Atari with Discrete World Models.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Mastering Atari with Discrete World Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.242136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.242136Z digest=sha256:43379c9e5ef561111efc0f9f2f0b2098842b059302cd6dd7e427b4543d9ff912

Observation 401fa314-6b9f-47b7-a504-5b3bfd1767f3 · outbound

This paper cites Mastering Diverse Domains through World Models.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Mastering Diverse Domains through World Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.257078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.257078Z digest=sha256:eecdae7023119f9e7af1af75f0d7059a19c817c6e68a8e242e5c4ba3807e6edd

Observation f4b7eb98-7dfd-4f39-9e84-d717639c59f9 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.266687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.266687Z digest=sha256:71126b47990a051ffaae12efece0bf92371fba4d750e99c0641766b01cfa965e

Observation 4249080f-d4b3-4d86-a741-110d2eac392b · outbound

This paper cites Temporal Difference Learning for Model Predictive Control.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Temporal Difference Learning for Model Predictive Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.271605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.271605Z digest=sha256:1066c8837c179959d66e426d8d35dba0c2d909c7e00118c766e0cd4160201aec

Observation eeaeff53-9369-48a2-8fdb-d34285396c27 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.276596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.276596Z digest=sha256:d935c2d3f71b076a5e62fa2e2014ada8dc22fa9c6630abe2f765d8ad48691dfb

Observation ae0332d3-7a15-42c5-a263-eb211c34bf9c · outbound

This paper cites Janner, J.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Janner, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.642661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.281062Z digest=sha256:a662149ead191608acc38934e9f4fc179cc84ede64e67680eb52a52493a299d9

Observation 6141572d-b3b7-4136-bcbb-60c913ca303b · outbound

This paper cites Kabzan, L.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Kabzan, L

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.626186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.285679Z digest=sha256:6568c5f1179e7f92581b10d7c8ab1f583be0cab1efbcaf700416132519a48ebf

Observation 42cae131-7b0f-4643-bdf9-76a84831b285 · outbound

This paper cites Katsigiannis and N.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Katsigiannis and N

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.609507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.289823Z digest=sha256:f09eaf06e7ebcf047212cf3ab18143c6be966a604217d11e1ca41581147a130d

Observation 9c18f673-7e58-471c-aa6e-cb5990781dc0 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Offline Reinforcement Learning with Implicit Q-Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.294338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.294338Z digest=sha256:0bdb7da3019807edae2b5df194314647029a1cd3d56090ee3074784cc268d8ca

Observation 7a85f7ab-7872-4eb4-9af3-31ab4ef8bcb1 · outbound

This paper cites Kumar, J.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Kumar, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.594121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.298917Z digest=sha256:6ebd1a534f7be54f2266aeb9b39bc878bde18d115fe67e2cbd2f9381ccecc6fd

Observation 83656bc0-ea89-49bf-9cc9-47ae1768b611 · outbound

This paper cites Kumar, A.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Kumar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.579052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.303175Z digest=sha256:ea9d611555f1a1abc1a8b0244e02f39d76381179e36d06295c9a4053a7402f28

Observation 2e7546d1-e634-4b71-b34c-7210e38c3061 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.339026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.339026Z digest=sha256:422a104d736b5bb5b2ff2f471798bd9f93e779399c7ef32191481a2836725bb1

Observation 029326ba-1ce2-4930-958f-0a7ecdd9392c · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.371033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.371033Z digest=sha256:45568c3fc24fcc7c6f92c4f812ab8b558b68d38f05f09ba4e71e212cd8eaa9de

Observation e0318fbd-c343-4565-ab6c-4468e4f59fd5 · outbound

This paper cites Littman and A.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Littman and A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.563079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.410665Z digest=sha256:8b9c9cc8b7a2f1d1aed09f79a1b24dcd9fa41486cd9413f85765758169cd5e40

Observation 7348b843-5152-447b-b0ab-b73dd3d03181 · outbound

This paper cites Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.455729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.455729Z digest=sha256:80004a4da009a17a851c8ce6c3428ae85febc9fe52e6c928d6f5a6c3bd487bbb

Observation 1e4ef373-e834-469e-920c-4bac2418e548 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:38:31.546528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.487926Z digest=sha256:379850828a5b34a87515a92d2b2e2f4d28fd53c1fbee3449c849d512bfaafc5f

Observation 5f696695-2beb-46e4-b953-857243b74732 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:38:31.530900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.524226Z digest=sha256:940d674970e224be3533962378bc8ae71ae0d5e1ad6d36244461eabb19fbe71e

Observation f4b860dc-19a2-4a85-a1c0-4331ef5d6a08 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.572672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.572672Z digest=sha256:2de7d11f56b2a5bdc678f77c9b1ee0046d7a76e2a38de34894361168d5ee9a6e

Observation 79e80661-4606-431d-bf29-13365e22ffd8 · outbound

This paper cites Nakamoto, S.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Nakamoto, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.515069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.609097Z digest=sha256:a4d4efc2e11fa4cc14386c21cfdbf6d6842e10f17eb0d205705033f3e0342a28

Observation d7bf83f9-7ea1-4056-8267-c873f50e8fe9 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.689021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.689021Z digest=sha256:e94c298622dbfd23f9097535cacbd554baeb3d234c867fb455df850743393c33

Observation 630c9d3d-4412-45d3-b40b-f0db30e070d4 · outbound

This paper cites Trust Region Policy Optimization.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Trust Region Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.694082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.694082Z digest=sha256:949d9d3f7570bf823116d159fae1af9ac4e995b99f3d40a779c4350a4ff64c24

Observation 1294ccb0-cd74-46f5-b3e4-88d81c668ee6 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.698954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.698954Z digest=sha256:3588fe0dbd271182434bacf24a4f30cfe9f7c552a639404c209613ff5ed53846

Observation cce74b20-cde2-4ce9-946f-f98bd742c5a1 · outbound

This paper cites Sikchi, W.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Sikchi, W

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.472133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.703425Z digest=sha256:9be6e4b1e70bfdade8923b76937ba44483204f6331a85b128f79bdc6aff8bba4

Observation 95bbe34e-3893-4893-a074-af28b11983b2 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:38:31.321794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.707899Z digest=sha256:b390104a88d4fbc0c0545425e67dff8b5f8a39fda95a402d1a3bd914968f8b85

Observation e714c118-1a4f-4184-81fb-10adb79753f1 · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.711895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.711895Z digest=sha256:f3d945e81ba989e42745ea7ca283aca4086d91b8c3a6d3e05968c3ca6cf68b5a

Observation 524f49e3-1015-4ff9-8ba5-bf0351f5b9fb · outbound

This paper cites an unresolved cited work.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.716371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.716371Z digest=sha256:0e62df3f16ea45121c410235741bb04e6da2ae919878e416a4adf828aa7b5f7a

Observation 03b7e3bf-b9ad-4a06-9d3d-d41cd7ab7fc9 · outbound

This paper cites DeepMind Control Suite.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint DeepMind Control Suite

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.724743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.724743Z digest=sha256:395456316c029597f64032d4dcbb3a68c5b1103ba82bfae540fdabbcf6a164c4

Observation 8823a377-3131-4459-bc00-2a4113ebc097 · outbound

This paper cites Thrun and A.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Thrun and A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.265681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.730792Z digest=sha256:0352a3593361b4db25e14ce019ed4a7e7bda36c6986b8969855abbc78b5cb42c

Observation aff67faf-75cd-4427-8aa2-a7851e2150eb · outbound

This paper cites Williams, N.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Williams, N

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:38:31.250709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:38:30.734739Z digest=sha256:12bbc5368a3ebe72b5095b660a3d739bd76894a1eb01466fc2604b7275705ddd

Observation 5e5b0fcc-f58c-4872-b15f-75026891178c · outbound

This paper cites Diffusion Model Predictive Control.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Diffusion Model Predictive Control

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.745532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.745532Z digest=sha256:fda299e2c79d6a7735d977eea7127d125c902b8671d57ac650bc096f86e04364

Pith citing papers

Observation a772cadb-ca4b-44c8-844c-52e155ff9de9 · inbound

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion cites this paper.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.170831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.170831Z digest=sha256:4b9c944c9af7e440a7b783ab2f5ad9394c25573870a61f458289fac95a0daf18

Observation 75bf9623-f153-4404-a424-e19b8526d9cf · inbound

RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC cites this paper.

RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:27.805022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:43:35.563513Z digest=sha256:128bca70a11699ecd461a17bf785a98ade5626d5d91ad4fc462e4bb605837dea

Observation 7c62a9e8-584d-419c-b612-6138dcb44332 · inbound

Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination cites this paper.

Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:11:05.433394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:31:35.396809Z digest=sha256:b445e93601bc287a3e65077ce2bf20a214b6bb915879ccea99291f5ce9ded7f2