Pith. sign in

REVIEW 10 cited by

URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11656 v3 pith:PCD6KUYJ submitted 2024-05-19 cs.RO cs.AI

classification cs.ROcs.AI
keywords simulationscenesmodelsimagespipelineproblemrealistictraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Constructing simulation scenes that are both visually and physically realistic is a problem of practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large data-hungry learning methods seek new sources of training data for physical decision-making systems. However, building simulation models is often still done by hand. A graphic designer and a simulation engineer work with predefined assets to construct rich scenes with realistic dynamic and kinematic properties. While this may scale to small numbers of scenes, to achieve the generalization properties that are required for data-driven robotic control, we require a pipeline that is able to synthesize large numbers of realistic scenes, complete with 'natural' kinematic and dynamic structures. To attack this problem, we develop models for inferring structure and generating simulation scenes from natural images, allowing for scalable scene generation from web-scale datasets. To train these image-to-simulation models, we show how controllable text-to-image generative models can be used in generating paired training data that allows for modeling of the inverse problem, mapping from realistic images back to complete scene models. We show how this paradigm allows us to build large datasets of scenes in simulation with semantic and physical realism. We present an integrated end-to-end pipeline that generates simulation scenes complete with articulated kinematic and dynamic structures from real-world images and use these for training robotic control policies. We then robustly deploy in the real world for tasks like articulated object manipulation. In doing so, our work provides both a pipeline for large-scale generation of simulation environments and an integrated system for training robust robotic control policies in the resulting environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fail2Progress: Learning from Real-World Robot Failures with Stein Variational Inference

    cs.RO 2025-09 conditional novelty 7.0 of 10

    Fail2Progress generates failure-targeted simulation data via Stein variational inference and fine-tunes skill effect models, improving long-horizon manipulation success rates and generalizing to unseen object counts a...

  2. SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    SimFoundry automates zero-shot real-to-sim scene generation from video, producing digital twins and cousins that enable policy training with 0.911 mean Pearson correlation to real-world results and 17-40% success gain...

  3. PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations

    cs.RO 2026-02 conditional novelty 6.0 of 10

    PokeNet estimates joint types, axes, ranges, and operation order of articulated objects directly from a single-view point cloud video of a human demonstration.

  4. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  5. GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training

    cs.RO 2025-07 conditional novelty 6.0 of 10

    GraspGen shows that training a grasp-scoring discriminator on the generator's own simulated outputs, plus a large new multi-gripper dataset, improves 6-DOF grasping across simulation and a real robot.

  6. 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A self-improving vision-language-model policy iteratively crafts 3D environments from text, and renderings of those environments serve as effective synthetic pretraining data for vision models.

  7. AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

    cs.RO 2025-06 conditional novelty 6.0 of 10

    AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...

  8. UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified pipeline and LLM-based model that jointly predicts articulation and physical properties of 3D assets, plus a 40K-object dataset and verified benchmark.

  9. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

  10. Generating Actionable Robot Knowledge Bases by Combining 3D Scene Graphs with Robot Ontologies

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A pipeline standardizes heterogeneous robot scene formats into USD, maps them through semantic reporting into an ontology-backed knowledge graph, and a robot uses the graph to answer task queries for breakfast table-setting.

Pith tools