Pith. sign in

Paper Citation Record · LEDGER

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2501.04228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04228 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:45:25.203483Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:19:07.791607Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:16:33.829615Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92d3b4b9-1155-4786-a4eb-542f55e02170 · outbound

This paper cites Learning bipedal robot locomotion from hu- man movement,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Learning bipedal robot locomotion from hu- man movement,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.655529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.078299Z digest=sha256:1f8804457100586ef0d651224b59665d250f8fb3c7d0cec360f63d7b3eeb4a83

Observation 782e5d4a-c734-4655-a320-abe655d516ed · outbound

This paper cites Champion-level drone racing using deep reinforce- ment learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Champion-level drone racing using deep reinforce- ment learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.085999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.085999Z digest=sha256:0c3bd271d7904b458cb0f553fd32a874af9c0eae3f70d80eda7f337fd8d97e3b

Observation 47c9e932-1ae8-4a95-b932-98cd1696c037 · outbound

This paper cites Robot parkour learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Robot parkour learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.621452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.092430Z digest=sha256:c9da0b5315b3f11df236fb6739a35ca9d942e038f5390865590ec4fbd620a83f

Observation 05ee9a45-eb3b-418c-b5e4-1cff474b55bc · outbound

This paper cites Cat: Constraints as terminations for legged locomotion reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Cat: Constraints as terminations for legged locomotion reinforcement learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.600547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.099871Z digest=sha256:f0cd52b612d7794435e96cc75223a703eea56701188fec49b8ee4601f7510d48

Observation 4d3065f9-c25e-4600-baba-621c1bdd7c54 · outbound

This paper cites Learning agile and dynamic motor skills for legged robots,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Learning agile and dynamic motor skills for legged robots,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.576349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.105866Z digest=sha256:fb81cc71c70becab83b0fa6b3d60e27fc40616424e4d05c99fd4808f4770f97a

Observation 4d97ea8c-2be4-4b20-9e3f-d2041527e3a9 · outbound

This paper cites Boyd and L.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Boyd and L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.111129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.111129Z digest=sha256:2181e26134d57956976600af931f93e5eb3a0b8f72685009e4145e065f1eead5

Observation 30f58679-35a3-45c8-9960-f469244af536 · outbound

This paper cites Outracing champion gran turismo drivers with deep reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Outracing champion gran turismo drivers with deep reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.542087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.117310Z digest=sha256:091cc8bae44c0a8f66de77f9732014c42d181190d98db52da4448faab48b696e

Observation 3e4678e9-4e2d-4507-8ed2-909b1a95fe8f · outbound

This paper cites Benchmarking safe exploration in deep reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Benchmarking safe exploration in deep reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.524122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.123540Z digest=sha256:68fa0bc1053cbcf360c23c594fc8723b53098e3dee22c3e9a7ed668d72b57f83

Observation b3d1eaa7-b688-478c-a441-deebc05e5eda · outbound

This paper cites Learning to walk in the real world with minimal human effort,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Learning to walk in the real world with minimal human effort,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.503746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.128084Z digest=sha256:7eb654caffeb0a2f643d34650bcdc05ba57374253effbbdb5b3ddd14ba17bf59

Observation df298fa6-1a0d-4429-b293-1291f4804733 · outbound

This paper cites Real-time perceptive motion control using control barrier functions with analytical smoothing for six-wheeled- telescopic-legged robot tachyon 3,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Real-time perceptive motion control using control barrier functions with analytical smoothing for six-wheeled- telescopic-legged robot tachyon 3,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.481481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.132783Z digest=sha256:b2c4bca8239a2a7238e2d38646408df97f74231d4d59d46a5a90851e979850f0

Observation 3fa21293-a985-42d6-8719-ee76c72cf5a8 · outbound

This paper cites A comprehensive survey on safe reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions A comprehensive survey on safe reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.464542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.137072Z digest=sha256:d5b4a8ba91d776a680cca953ad0f9588f2f25e52f863d5ae15baa29d49bd12da

Observation 1ed1db81-1518-4574-945c-70f195a06df1 · outbound

This paper cites Penalized proximal policy optimization for safe reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Penalized proximal policy optimization for safe reinforcement learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.447305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.144171Z digest=sha256:a3ea3f00f148891008cd30fa43e033a3dc02379f741c5b8443e398d254bd521b

Observation 1d220833-10d0-48ae-827a-f031901482a4 · outbound

This paper cites Robot reinforcement learning on the constraint manifold,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Robot reinforcement learning on the constraint manifold,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.429887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.150252Z digest=sha256:8f5a7fd67cc9a5ce4ce98d8624747e9c987ad628e7635029c0246d6e2f72737a

Observation c8e2244b-9576-455f-b932-b0162063d5fd · outbound

This paper cites Trust region policy optimization,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Trust region policy optimization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.411592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.155601Z digest=sha256:de13f286bc68b43b1a77f6fc6fbe0e995a3f4cc300476a7fc6c85423bba2003d

Observation 8f508334-3915-499e-b54b-e4974a34f184 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.161838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.161838Z digest=sha256:1257c16756219ed23fcc91e195f12a3762ee6bde5201c947bb3439d4db77c842

Observation 6926126d-7005-4594-a378-1e7494faa967 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.394471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.167951Z digest=sha256:ac4091be61396439a7b966057a79e3d4cb0f9c52f92d631154c00a51d9ea8072

Observation 969e9470-17bd-4080-bae2-80fc4a7e5ba6 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Soft Actor-Critic Algorithms and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.174248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.174248Z digest=sha256:7e735cd99bf21cdbc5c03a5b578f8c109184531a3e0b290a336c1f5409e52ceb

Observation bcd72484-cc04-45d8-bd31-f07be145c0e0 · outbound

This paper cites Evaluation of Constrained Reinforcement Learning Algorithms for Legged Locomotion.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Evaluation of Constrained Reinforcement Learning Algorithms for Legged Locomotion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.180371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.180371Z digest=sha256:f35c8aee787676453065be5c10cceb7826767b87fd077d3e32fe0d0262625ef3

Observation 9583058d-0c6f-461e-b780-f2560d46d777 · outbound

This paper cites Learning agile soccer skills for a bipedal robot with deep reinforcement learning,.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Learning agile soccer skills for a bipedal robot with deep reinforcement learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.377392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.187371Z digest=sha256:f33ce2b9412f6f091a93e79d8f0184a5efa663c7d0f90506c62cb3b9b377ccae

Observation 4a8ee6c1-620f-470a-9e91-3a2feda6f2ac · outbound

This paper cites an unresolved cited work.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:45:25.357798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.192744Z digest=sha256:d4d07254a60ad0a78b40cc68cff688f615e7dad950a89524463495f340aa3722

Observation 65c9b896-690b-4eb7-a4d9-829137c734de · outbound

This paper cites Mujoco: A physics engine for model-based control.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions Mujoco: A physics engine for model-based control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:45:25.339663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:45:25.198253Z digest=sha256:716f73e49231540e0da2d10e75784fea6d9793e97247d3fccab0a6168e59b53c

Observation e4227e1d-006e-42d8-92af-336324b21b7a · outbound

This paper cites OpenAI Gym.

Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions OpenAI Gym

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:25.203483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:25.203483Z digest=sha256:cd05ce78e2e05de99fb78f04d1417c7d643221fc2927a81fd96cc7a38d067e4e

Pith citing papers

Observation dbf013d3-9484-45d5-972c-d17b4d67a52c · inbound

Gain Tuning Is Not What You Need: Reward Gain Adaptation for Constrained Locomotion Learning cites this paper.

Gain Tuning Is Not What You Need: Reward Gain Adaptation for Constrained Locomotion Learning Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:19:07.791607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:19:07.791607Z digest=sha256:b18984fac0c984bcb50009d053ccb4b71f6cf145df68e6456e56b6fa712a21a5

Observation d4d97c39-604e-4831-912e-05562f8fc141 · inbound

ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control cites this paper.

ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:33.831410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:13:29.399114Z digest=sha256:f1deb41e12c341dadddd88bf54dc2e30ab83729107e9996edabadc91dfcbd43f