Pith. sign in

REVIEW 4 cited by

RobotKeyframing: Learning Locomotion with High-Level Objectives via Mixture of Dense and Sparse Rewards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11562 v2 pith:IO3XMQO2 submitted 2024-07-16 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords frameworkhigh-levelobjectivestargetsdenseeffectivelyexperimentslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a novel learning-based control framework that uses keyframing to incorporate high-level objectives in natural locomotion for legged robots. These high-level objectives are specified as a variable number of partial or complete pose targets that are spaced arbitrarily in time. Our proposed framework utilizes a multi-critic reinforcement learning algorithm to effectively handle the mixture of dense and sparse rewards. Additionally, it employs a transformer-based encoder to accommodate a variable number of input targets, each associated with specific time-to-arrivals. Throughout simulation and hardware experiments, we demonstrate that our framework can effectively satisfy the target keyframe sequence at the required times. In the experiments, the multi-critic method significantly reduces the effort of hyperparameter tuning compared to the standard single-critic alternative. Moreover, the proposed transformer-based architecture enables robots to anticipate future goals, which results in quantitative improvements in their ability to reach their targets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Non-conflicting Energy Minimization in Reinforcement Learning based Robot Control

    cs.RO 2025-09 conditional novelty 6.0 of 10

    PEGrad projects energy-minimization gradients orthogonal to task-reward gradients in RL, achieving 64% torque reduction in simulation and reduced battery draw on a Unitree Go2 without sacrificing task reward.

  2. Arnold: a generalist muscle transformer policy

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A single transformer policy with a compositional sensorimotor vocabulary achieves expert or super-expert performance on 14 musculoskeletal control tasks spanning four embodiments.

  3. Multi-critic Learning for Whole-body End-effector Twist Tracking

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A multi-critic and twist-based RL controller achieves accurate whole-body loco-manipulation on a quadruped with an arm, outperforming pose-based baselines and MPC in tracking error.

  4. Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why

    cs.LG 2025-07 conditional novelty 3.0 of 10

    Feature-based and GAN-based imitation learning should be selected by task priorities (fidelity, diversity, interpretability, adaptability), not by paradigm loyalty.

Pith tools