Pith. sign in

REVIEW 2 cited by

Symmetry Considerations for Learning Task Symmetric Robot Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04359 v1 pith:ZN4TNJWZ submitted 2024-03-07 cs.RO cs.AI

classification cs.ROcs.AI
keywords symmetrytaskslearnedroboticapproachesbehaviorsfaillearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Symmetry is a fundamental aspect of many real-world robotic tasks. However, current deep reinforcement learning (DRL) approaches can seldom harness and exploit symmetry effectively. Often, the learned behaviors fail to achieve the desired transformation invariances and suffer from motion artifacts. For instance, a quadruped may exhibit different gaits when commanded to move forward or backward, even though it is symmetrical about its torso. This issue becomes further pronounced in high-dimensional or complex environments, where DRL methods are prone to local optima and fail to explore regions of the state space equally. Past methods on encouraging symmetry for robotic tasks have studied this topic mainly in a single-task setting, where symmetry usually refers to symmetry in the motion, such as the gait patterns. In this paper, we revisit this topic for goal-conditioned tasks in robotics, where symmetry lies mainly in task execution and not necessarily in the learned motions themselves. In particular, we investigate two approaches to incorporate symmetry invariance into DRL -- data augmentation and mirror loss function. We provide a theoretical foundation for using augmented samples in an on-policy setting. Based on this, we show that the corresponding approach achieves faster convergence and improves the learned behaviors in various challenging robotic tasks, from climbing boxes with a quadruped to dexterous manipulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning

    cs.RO 2025-05 conditional novelty 7.0 of 10

    A single legged-robot policy, trained by distilling nine expert skills and then fine-tuning with reinforcement learning, matches or beats each expert and generalizes to unseen unstructured terrain.

  2. Learning Accurate Whole-body Throwing with High-frequency Residual Policy and Pullback Tube Acceleration

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A legged mobile manipulator throws grasped objects to 4-6 m targets with mean landing errors around 0.28-0.43 m, using a learned nominal policy, a 400 Hz residual policy, and closed-loop pullback tube acceleration.

Pith tools