Pith. sign in

REVIEW 7 cited by

Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13882 v4 pith:TXAN6YSY submitted 2024-10-03 cs.CV

classification cs.CV
keywords articulate-anythingobjectssystemarticulatedcomplexextensiveinputspolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Interactive 3D simulated objects are crucial in AR/VR, animations, and robotics, driving immersive experiences and advanced automation. However, creating these articulated objects requires extensive human effort and expertise, limiting their broader applications. To overcome this challenge, we present Articulate-Anything, a system that automates the articulation of diverse, complex objects from many input modalities, including text, images, and videos. Articulate-Anything leverages vision-language models (VLMs) to generate code that can be compiled into an interactable digital twin for use in standard 3D simulators. Our system exploits existing 3D asset datasets via a mesh retrieval mechanism, along with an actor-critic system that iteratively proposes, evaluates, and refines solutions for articulating the objects, self-correcting errors to achieve a robust outcome. Qualitative evaluations demonstrate Articulate-Anything's capability to articulate complex and even ambiguous object affordances by leveraging rich grounded inputs. In extensive quantitative experiments on the standard PartNet-Mobility dataset, Articulate-Anything substantially outperforms prior work, increasing the success rate from 8.7-11.6% to 75% and setting a new bar for state-of-the-art performance. We further showcase the utility of our system by generating 3D assets from in-the-wild video inputs, which are then used to train robotic policies for fine-grained manipulation tasks in simulation that go beyond basic pick and place. These policies are then transferred to a real robotic system.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    SimFoundry automates zero-shot real-to-sim scene generation from video, producing digital twins and cousins that enable policy training with 0.911 mean Pearson correlation to real-world results and 17-40% success gain...

  2. ScrewSplat: An End-to-End Method for Articulated Object Recognition

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A method that recovers the 3D shape and the rotation or sliding axis of each movable part of an object from RGB video alone, by jointly optimizing randomly initialized screw axes with Gaussian Splatting.

  3. VLMgineer: Vision Language Models as Robotic Toolsmiths

    cs.RO 2025-07 conditional novelty 6.0 of 10

    VLMgineer combines VLM-generated URDF tool designs with evolutionary search to co-design tools and action plans, outperforming human-specified and existing tools on a new simulated manipulation benchmark.

  4. SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting

    cs.GR 2025-06 conditional novelty 6.0 of 10

    SplArt estimates revolute or prismatic joint parameters and part-level 3D Gaussian geometry from two sets of posed RGB images using self-supervised multi-stage optimization.

  5. Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels

    cs.CV 2025-08 reject novelty 5.0 of 10

    A supervised 3D U-Net predicts per-voxel material fields from CLIP feature grids, enabling fast MPM-based animation, but the reported evidence depends on pseudo-labels and a VLM judge from the same model family as the...

  6. Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation

    cs.RO 2025-02 conditional novelty 5.0 of 10

    A reconstruction and neural-rendering pipeline converts real tabletop scenes into photorealistic robot simulations, and policies trained only on simulated data transfer zero-shot to the real robot with an average succ...

  7. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Pith tools