Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Is Single-View Mesh Reconstruction Ready for Robotics?

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Current single-view 3D mesh models fail robotics-specific requirements for digital twin creation, an empirical evaluation on real robotics data finds.

desk verdict A careful, honest benchmark showing single-view mesh reconstruction is not ready for physics-sim manipulation; the thresholds are debatable but the negative trend is robust. read the letter →

arxiv 2505.17966 v2 pith:XKYJLQXQ submitted 2025-05-23 cs.RO cs.CV

classification cs.ROcs.CV
keywords single-view3Dreconstructionreal-to-simdigitaltwinroboticmanipulationmeshphysicssimulationobjectstabilityocclusionhandling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether today's single-view, category-agnostic mesh reconstruction models can serve as the perception front-end for instant digital twin creation in robotic manipulation: mapping one RGB(-D) image of a tabletop scene into complete object meshes ready for a physics simulator. To answer it, the authors define five robotics-specific desiderata, including surface accuracy within 2 mm Chamfer distance, collision-free geometry, stability within 5 degrees of the scene pose, occlusion error no more than 10 percent above visible-region error, and full-scene reconstruction within 2 seconds. They then evaluate eight object-level and three scene-level models on YCB-Video and Aria Digital Twin, two real robotics datasets with ground-truth meshes. Most models sit about 5 mm from the true surface on YCB-Video and about 1 cm on Aria, only about half of computed grasps transfer to the target meshes, and most reconstructed objects lack stable poses near their observed scene pose. The paper concludes that, despite strong computer vision benchmark results, existing approaches do not yet meet robotics-specific requirements for physics-based planning.

What carries the argument

The evaluative machinery combines five desiderata with a physics-aware measurement protocol on two annotated datasets. Reconstruction accuracy is measured as bidirectional Chamfer distance between 10,000 sampled surface points after ICP alignment seeded with 512 quaternion initialisations; grasp transfer is tested with antipodal grasp sampling from a parallel-jaw gripper, checking collisions, a 22.5-degree surface-normal alignment bound, and contact retention after shaking in simulation. Collisions are detected with FCL on objects placed at their estimated scene poses; stability is tested by perturbing each object 5 degrees away from candidate stable poses in PyBullet and checking whether it returns; occlusion handling compares Chamfer distance on visible versus occluded surface regions using a mask-recomputed ICP. The thresholds of 2 mm, zero collisions, 5 degrees stability, 10 percent occlusion degradation, and 2 seconds per scene are the yardstick that turns raw reconstruction error into a robotics-readiness verdict.

What would settle it

Run the same eight object-level models on a task whose success criterion is 5 mm surface distance instead of 2 mm; if grasp transfer to ground-truth meshes then exceeds roughly 90 percent, the paper's not-ready verdict would not generalise to that task, while success near 50 percent would confirm that reconstruction accuracy is the load-bearing constraint.

Watch

Extended reading notes

Core claim

The central claim is that single-view, category-agnostic full-mesh reconstruction models, used out-of-the-box on realistic robotics inputs, fail the requirements a digital twin for manipulation imposes. On YCB-Video, reconstructed surfaces are typically 5 mm from the closest ground-truth surface against a 2 mm target, and on Aria Digital Twin the median error is roughly twice as large. Grasp poses computed on the reconstructions transfer to the ground-truth meshes only about half the time, collisions occur in a majority of scenes for single-object models, and both object- and scene-level reconstructions are mostly unstable within 5 degrees of their observed pose. Occluded object regions raise Chamfer error by 40 to 95 percent for object-level models, while scene-level models that inpaint or jointly denoise objects stay near the 10 percent bound. Only SF3D and ZeroShape reconstruct a single object within roughly one second, and scene-level models take an order of magnitude longer, so the 2-second per scene target is met by none of them.

Load-bearing premise

The verdict depends on the five thresholds the authors set in Section II—2 mm surface distance, zero collisions, stability within 5 degrees, at most 10 percent extra error on occluded regions, and 2 seconds per scene—which are presented as robotics requirements but are not derived from a specific manipulation task.

Editorial extensions

If this is right

  • If correct, single-view reconstruction cannot currently serve as the perception front-end for real-time, physics-based manipulation planning at the tolerances the authors specify.
  • Practitioners should prefer scene-level models over single-object models when objects are physically close, because single-object models put the reconstructed meshes in mutual collision in a majority of evaluated scenes.
  • Occlusion handling improves sharply when models use scene context, either image inpainting before reconstruction or joint multi-object denoising, suggesting a concrete design direction for future reconstruction models.
  • Computational cost is the one desideratum with clear winners, with SF3D and ZeroShape producing objects in about 0.5 and 1 second, respectively, yet even they cannot handle a multi-object scene within the 2-second target.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 2 mm surface-distance threshold is not derived from a specific manipulation task; a task with 5 mm clearance, such as suction grasping of large objects, could succeed with current models even though the paper's headline verdict says not ready.
  • Grasp transfer rates depend on the chosen gripper and evaluation thresholds, so re-running the transfer test with a compliant or suction gripper could change the ranking and could raise success rates above the reported 50 percent.
  • A direct test of the paper's proposed remedies would be to fine-tune a fast model such as SF3D with physics-informed stability training while feeding it ground-truth depth, then re-running the same five-desiderata protocol to measure how much of the gap closes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper asks whether current single-view, category-agnostic mesh reconstruction models are ready to serve as the perception front-end for robot manipulation through instant digital twin creation. It proposes five robotics desiderata—reconstruction accuracy within 2 mm Chamfer distance, collision-free geometry, physical stability within 5° of the scene pose, bounded occlusion error, and scene reconstruction within 2 seconds—and evaluates eleven single-view reconstruction models on the YCB-Video and Aria Digital Twin datasets. The evaluation measures Chamfer distance, mesh collisions, object stability in PyBullet, occlusion resilience, latency, and memory, and adds an external grasp-transfer test in simulation. The main finding is that current models fail these desiderata by large margins, with typical Chamfer errors around 5 mm on YCB-Video and worse on Aria, frequent collisions and instabilities, substantially degraded occlusion handling, and inference times that exceed the 2 s target for most models.

Significance. If the empirical results hold, this is a valuable and timely negative result for the real-to-sim community: it documents, across two datasets, eleven models, and multiple complementary metrics, that single-view reconstruction performance on computer vision benchmarks does not transfer to robotics-grade physical simulation. The grasp-transfer experiment is a particular strength because it is an external, task-relevant validation against ground-truth meshes rather than a re-fit evaluation metric. The evaluation is also conservative in several respects—giving ground-truth masks to segmentation-dependent methods and manually selecting frames likely favors the models, so the observed failures are not easily explained away by pipeline artifacts. The main weakness is that the headline 'not ready' conclusion is stronger than the evidence directly supports, because the threshold values (notably 2 mm Chamfer) are asserted rather than derived from a downstream manipulation success criterion, a point the authors themselves concede in Section IV-D.

major comments (3)
  1. [§II, Desiderata 1–5; §IV-D] The thresholds of 2 mm Chamfer distance, 5° stability, and 2 s latency are presented as 'robotics-specific requirements', but they are not derived from a concrete manipulation objective. Section IV-D explicitly concedes that permissible reconstruction error is highly context-dependent and that deriving a systematic error-tolerance relationship is left for future work. The headline claim that 'existing approaches fail to meet robotics-specific requirements' is therefore stronger than the measurements alone establish: a task tolerating 5 mm surface error, or a pipeline using a robust grasp sampler, might accept the very reconstructions reported here. Please either add a controlled study that ties the thresholds to task outcomes (for example, corrupting ground-truth meshes to controlled Chamfer levels and measuring grasp-transfer or pick-and-place success), or rephrase the central claim as failure against the proposed desiderata rather than an unconditional statement about robotics readiness.
  2. [§IV-C, Fig. 3] The grasp-transfer experiment reports per-grasp transfer success rates near 50% without stating a required success rate, without a ground-truth-to-ground-truth baseline, and without an end-to-end pick-and-place evaluation. Since a manipulation planner can sample, rank, and filter many candidate grasps, a 50% per-grasp transfer rate is not by itself sufficient evidence that the reconstructions are unusable for manipulation. Please either report a task-level success metric with an explicit acceptance threshold and a baseline (e.g., grasps computed and evaluated on ground-truth meshes), or soften the conclusion drawn from this figure.
  3. [§IV-A and §IV-B] The evaluation uses manually selected objects and frames and provides ground-truth masks to all segmentation-dependent methods. Providing ground-truth masks is a conservative choice that strengthens the negative result, but the manual selection protocol is not quantified: the paper does not report the distribution of object categories, view-points, occlusion levels, or pose diversity, nor does it analyze whether the chosen frames are representative of a deployment distribution. Please document the selection protocol in detail and, if possible, release the full object and frame lists so that other researchers can reproduce or extend the benchmark.
minor comments (5)
  1. [§IV-A] InstantMesh is cited as [123] in the model description, but reference [123] is Instant3D; the rest of the paper and Figures 2, 3, 4, etc. cite InstantMesh as [27]. Please reconcile the citation.
  2. [§II, abstract] There are several typos: 'This is ensures physical stability' in Desideratum 2 should be 'This ensures physical stability', and the abstract uses 'quantitively' instead of 'quantitatively'.
  3. [§IV-B] The scale-estimation procedure based on ratios of principal-component standard deviations should include a caveat for objects with near-degenerate principal components (e.g., flat or axially symmetric objects), where the median ratio may be numerically unstable.
  4. [Fig. 10] The legend and axis of Figure 10 both repeat 'Relative Error Increase (%)'; please clean up the caption and axis labels.
  5. [General] The authors do not state whether evaluation code, selected frames, or model configuration files will be released; providing these would substantially increase the benchmark's reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper measures external models against author-set desiderata; the negative verdict is a threshold-dependent judgment, not a quantity that reduces to its inputs.

full rationale

This paper does not derive a prediction from a fitted parameter. It defines five robotics desiderata in Section II (2mm Chamfer, no collisions, stability within 5 degrees, at most 10% occlusion-error increase, and 2s latency) and then evaluates eleven existing single-view reconstruction models on YCB-Video and Aria Digital Twin. The measured Chamfer distances, collision frequencies, stability counts, occlusion-error increases, and runtimes are computed from model outputs and ground-truth meshes via independent pipelines (multi-initialisation ICP alignment, FCL collision queries, PyBullet stability simulation, MetaGraspNet grasp transfer). None of these quantities is defined in terms of the desideratum it is compared against: Fig. 2 reports raw surface distances before applying the 2mm bar, and Fig. 3 reports grasp-transfer success as an external task-level check. The only self-citation in the measurement chain is [107] (Occupancy Networks, which shares an author) as the source of the Chamfer-distance implementation; that is a standard, parameter-free metric and is not the load-bearing content of the paper's claim. No uniqueness theorem or ansatz is imported from the authors' prior work. The one in-scope limitation is explicitly stated in Section IV-D: "the effect of reconstruction errors on the performance of robot manipulation tasks in general and on physical simulation is highly context-dependent... the question of precisely what error magnitude is permissible in a certain situation and whether any systematic relationship can be drawn up is an entire research question of itself and left for future work." This concession shows the 2mm/5-degree/2s thresholds are normative choices rather than derived consequences; it weakens the strength of the 'not ready' conclusion, but it is a validity caveat, not a circular step. The raw measurements are reported transparently and would support the weaker claim that the models fail the authors' stated bars.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

All counted dependencies are evaluation-protocol choices rather than fitted model parameters. The central claim depends mainly on the four author-set thresholds in Section II and on the representativeness of the manually selected data, so the negative verdict is conditional. No new entities are introduced, and no numerical constants are fitted to make the conclusion hold.

free parameters (4)
  • reconstruction accuracy threshold = 2mm Chamfer distance
    Set in Desideratum 1 as the target for household robot manipulation; not derived from a specific task, and the paper's own limitations section admits permissible error is task-dependent.
  • stability tolerance = 5 degrees
    Desideratum 3 defines stable poses as within 5 degrees tilt of the scene pose, with 8 perturbation simulations; this tolerance determines the stability pass/fail counts.
  • occlusion degradation threshold = 10% relative Chamfer increase
    Desideratum 4: occluded regions should be no more than 10% worse than visible regions; used to judge occlusion handling.
  • scene latency target = 2 seconds per scene
    Desideratum 5: complete scene reconstruction within 2 seconds; the paper notes this is for less time-critical household applications, so it is an author-chosen operational target.
assumptions (5)
  • domain assumption PCA-based scale estimation recovers the metric scale of the reconstruction before alignment.
    Section IV-B: scale is estimated as the median ratio of PCA component standard deviations between ground truth and reconstruction; if reconstructions are partial or have hallucinated extents, this can distort the Chamfer error.
  • domain assumption 512 quaternion initializations suffice for ICP to find the globally best alignment.
    Section IV-B: a fixed pool of 512 quaternion seeds is used; no criterion certifies global optimality.
  • domain assumption Manual selection of objects and frames in YCB-Video and Aria Digital Twin is representative of robotics manipulation inputs.
    Section IV-A: 'we manually select objects and frames to ensure coverage', with no release of the selected subsets.
  • domain assumption Observed scenes are static, so reconstructed objects should be stable near their observed pose.
    Desideratum 3 and Section II-c: stability is defined relative to the assumption that all observed objects are at rest.
  • domain assumption Chamfer distance on 10,000 sampled surface points captures manipulation-relevant geometry.
    Section IV-C-a: used as the primary accuracy metric; 10,000 points is a sampling choice and CD does not directly measure grasp-relevant local features.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is Single-View Mesh Reconstruction Ready for Robotics?." pith.science (2026). https://pith.science/paper/XKYJLQXQ

@misc{pith2026250517966,
  author       = {Pith},
  title        = {Pith review of: Is Single-View Mesh Reconstruction Ready for Robotics?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XKYJLQXQ}},
  note         = {Machine review of arXiv:2505.17966}
}
read the original abstract

This paper evaluates single-view mesh reconstruction models for their potential in enabling instant digital twin creation for real-time planning and dynamics prediction using physics simulators for robotic manipulation. Recent single-view 3D reconstruction advances offer a promising avenue toward an automated real-to-sim pipeline: directly mapping a single observation of a scene into a simulation instance by reconstructing scene objects as individual, complete, and physically plausible 3D meshes. However, their suitability for physics simulations and robotics applications under immediacy, physical fidelity, and simulation readiness remains underexplored. We establish robotics-specific benchmarking criteria for 3D reconstruction, including handling typical inputs, collision-free and stable geometry, occlusions robustness, and meeting computational constraints. Our empirical evaluation using realistic robotics datasets shows that despite success on computer vision benchmarks, existing approaches fail to meet robotics-specific requirements. We quantitively examine limitations of single-view reconstruction for practical robotics implementation, in contrast to prior work that focuses on multi-view approaches. Our findings highlight critical gaps between computer vision advances and robotics needs, guiding future research at this intersection.

Figures

Figures reproduced from arXiv: 2505.17966 by the authors.

Figure 1
Figure 1. Illustration of a general Real2Sim2Real pipeline for robotic manipu [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Chamfer distances on YCB-Video [38]. CD is averaged across 10,000 sampled surface points from both target and reconstruction. Indicators for min, max, and median are shown. Most methods achieve 5mm surface. SF3D [80] InstantMesh [27] One2345 [26] LGM [96] Michelangelo [28] ZeroShape [142] Real3D [71] DSO [190] 0 20 40 60 80 100 Success Rate (%) 311 343 171 248 602 646 209 377 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 5
Figure 5. Frequency of reconstructed scenes where reconstructed meshes are [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figures from the paper (10 more)
Figure 6
Figure 6. Figure 6: Reconstruction error of scene reconstruction models on the YCB-Video [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 8
Figure 8. Figure 8: Total number of objects with physically stable poses within 5 [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 10
Figure 10. Figure 10: Chamfer distances (CD) of unoccluded object parts versus occluded [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Relative object position error of scene-level reconstruction models [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 13
Figure 13. Figure 13: Collision frequency in scene reconstruction on YCB-Video [ [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]
Figure 14
Figure 14. Figure 14: Chamfer distances on Aria Digital Twin [ [PITH_FULL_IMAGE:figures/full_fig_p019_14.png]
Figure 15
Figure 15. Figure 15: Object stability on Aria Digital Twin [39] dataset. Most reconstructions lack stability within 5∘ of ground truth poses. In physics simulation, unstable geometries would cause scene reconstructions to collapse. 0 1 2 3 4 Chamfer Distance (cm) [154] [153] [196] -20 -15…
Figure 16
Figure 16. Figure 16: Occlusion handling by scene-level models. Objects are individually [PITH_FULL_IMAGE:figures/full_fig_p019_16.png]
Figure 17
Figure 17. Figure 17: Sample reconstructions on the YCB-Video [ [PITH_FULL_IMAGE:figures/full_fig_p020_17.png]
Figure 18
Figure 18. Figure 18: YCB-Video [38] scene reconstructions using different 3D reconstruction methods that process multi-object scenes. The objects that are reconstructed are shown contoured and in colour. The scene renderings are solely for context as the reconstruction models are only giv…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.

Reference graph

Works this paper leans on

221 extracted references · 79 canonical work pages · cited by 1 Pith paper

  1. [1]

    World models,

    D. Ha and J. Schmidhuber, “World models,” arXiv preprint arXiv:1803.10122, vol. 2, no. 3, 2018

  2. [2]

    Model Based Reinforce- ment Learning for Atari,

    Ł. Kaiser, M. Babaeizadeh, P. Miłos, et al., “Model Based Reinforce- ment Learning for Atari,” 2019

  3. [3]

    Learning Latent Dynamics for Planning from Pixels,

    D. Hafner, T. Lillicrap, I. Fischer, et al., “Learning Latent Dynamics for Planning from Pixels,” in Proceedings of the 36th International Conference on Machine Learning , PMLR, 2019

  4. [4]

    Dream to Control: Learning Behaviors by Latent Imagination,

    D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to Control: Learning Behaviors by Latent Imagination,” 2019

  5. [5]

    T. Wang, X. Bao, I. Clavera, et al. , Benchmarking Model-Based Reinforcement Learning, 2019

  6. [6]

    Model- based Reinforcement Learning: A Survey,

    T. M. Moerland, J. Broekens, A. Plaat, and C. M. Jonker, “Model- based Reinforcement Learning: A Survey,” Foundations and Trends in Machine Learning , 2023

  7. [7]

    A survey of industrial model predictive control technology,

    S. J. Qin and T. A. Badgwell, “A survey of industrial model predictive control technology,” Control Engineering Practice , 2003

  8. [8]

    Temporal Difference Learning for Model Predictive Control,

    N. A. Hansen, H. Su, and X. Wang, “Temporal Difference Learning for Model Predictive Control,” in Proceedings of the 39th International Conference on Machine Learning , 2022

Show all 221 references
  1. [9]

    Robot Planning in the Real World: Research Challenges and Opportunities,

    R. Alterovitz, S. Koenig, and M. Likhachev, “Robot Planning in the Real World: Research Challenges and Opportunities,” AI Magazine, 2016

  2. [10]

    A survey of robot manipulation in contact,

    M. Suomalainen, Y . Karayiannidis, and V . Kyrki, “A survey of robot manipulation in contact,” Robotics and Autonomous Systems , 2022

  3. [11]

    Safe Model-based Reinforcement Learning with Stability Guarantees,

    F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause, “Safe Model-based Reinforcement Learning with Stability Guarantees,” in Advances in Neural Information Processing Systems , 2017

  4. [12]

    Scalable End-to-End Autonomous Vehicle Testing via Rare-event Simulation,

    M. O’ Kelly, A. Sinha, H. Namkoong, R. Tedrake, and J. C. Duchi, “Scalable End-to-End Autonomous Vehicle Testing via Rare-event Simulation,” in Advances in Neural Information Processing Systems , 2018

  5. [13]

    Demonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning,

    L. Smith, I. Kostrikov, and S. Levine, “Demonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning,” Robotics: Science and Systems (RSS) Demo , 2023

  6. [14]

    Bohlinger, J

    N. Bohlinger, J. Kinzel, D. Palenicek, L. Antczak, and J. Peters, Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion, 2025

  7. [15]

    Barcellona, A

    L. Barcellona, A. Zadaianchuk, D. Allegro, S. Papa, S. Ghidoni, and E. Gavves, Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination , 2024

  8. [16]

    X. Han, M. Liu, Y . Chen, et al. , Re$^3$Sim: Generating High- Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation, 2025

  9. [17]

    X. Li, J. Li, Z. Zhang, et al., RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator , 2024

  10. [18]

    H. Lou, Y . Liu, Y . Pan,et al., Robo-GS: A Physics Consistent Spatial- Temporal Model for Robotic Arm with Hybrid Representation , 2024

  11. [19]

    Y . Wu, L. Pan, W. Wu, G. Wang, Y . Miao, and H. Wang,RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning, 2024

  12. [20]

    M. N. Qureshi, S. Garg, F. Yandun, D. Held, G. Kantor, and A. Silwal, SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting , 2024

  13. [21]

    S. Zhu, L. Mou, D. Li, B. Ye, R. Huang, and H. Zhao, VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion, 2025

  14. [22]

    Y . Jia, G. Wang, Y . Dong, et al. , DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments , 2024

  15. [23]

    Torne, A

    M. Torne, A. Simeonov, Z. Li, et al. , Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation , 2024

  16. [24]

    A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards,

    S. Patel, X. Yin, W. Huang, et al., “A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards,” 2024

  17. [25]

    Pfaff, E

    N. Pfaff, E. Fu, J. Binagia, P. Isola, and R. Tedrake, Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups , 2025

  18. [26]

    One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization,

    M. Liu, C. Xu, H. Jin, et al., “One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization,” Advances in Neural Information Processing Systems , 2023

  19. [27]

    J. Xu, W. Cheng, Y . Gao, X. Wang, S. Gao, and Y . Shan,InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models , 2024

  20. [28]

    Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation,

    Z. Zhao, W. Liu, X. Chen, et al., “Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation,” in Advances in Neural Information Processing Systems , 2023

  21. [29]

    A Two-Stage Optimized Next-View Planning Framework for 3-D Unknown Environment Exploration, and Structural Reconstruction,

    Z. Meng, H. Qin, Z. Chen, et al., “A Two-Stage Optimized Next-View Planning Framework for 3-D Unknown Environment Exploration, and Structural Reconstruction,” IEEE Robotics and Automation Letters , 2017

  22. [30]

    Closed-Loop Next- Best-View Planning for Target-Driven Grasping,

    M. Breyer, L. Ott, R. Siegwart, and J. J. Chung, “Closed-Loop Next- Best-View Planning for Target-Driven Grasping,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2022

  23. [31]

    Interactive Perception: Leveraging Action in Perception and Perception in Action,

    J. Bohg, K. Hausman, B. Sankaran, et al. , “Interactive Perception: Leveraging Action in Perception and Perception in Action,” IEEE Transactions on Robotics , 2017

  24. [32]

    Planning Robotic Manipulation with Tight Environment Constraints,

    G. J. Pollayil, G. Grioli, M. Bonilla, and A. Bicchi, “Planning Robotic Manipulation with Tight Environment Constraints,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2021

  25. [33]

    Safe, Occlusion-Aware Manip- ulation for Online Object Reconstruction in Confined Spaces,

    Y . Miao, R. Wang, and K. Bekris, “Safe, Occlusion-Aware Manip- ulation for Online Object Reconstruction in Confined Spaces,” in Robotics Research, 2023

  26. [34]

    Y . Mu, T. Chen, S. Peng,et al., RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version) , 2024

  27. [35]

    Agarwal, G

    A. Agarwal, G. Singh, B. Sen, T. Lozano-Pérez, and L. P. Kaelbling, SceneComplete: Open-World 3D Scene Completion in Complex Real World Environments for Robot Manipulation , 2024

  28. [36]

    K. Yao, L. Zhang, X. Yan, et al., CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image , 2025. 16

  29. [37]

    Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models,

    P. Katara, Z. Xian, and K. Fragkiadaki, “Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) , 2024

  30. [38]

    PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,

    Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox, “PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,” 2018

  31. [39]

    Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception,

    X. Pan, N. Charron, Y . Yang, et al. , “Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception,” 2023

  32. [40]

    MuJoCo: A physics engine for model-based control,

    E. Todorov, T. Erez, and Y . Tassa, “MuJoCo: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012

  33. [41]

    Coumans and Y

    E. Coumans and Y . Bai, Pybullet, a python module for physics simulation for games, robotics and machine learning , 2016

  34. [42]

    Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning,

    V . Makoviychuk, L. Wawrzyniak, Y . Guo,et al., “Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning,” 2021

  35. [43]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering,

    B. Kerbl, G. Kopanas, T. Leimkuehler, and G. Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering,” ACM Trans. Graph., 2023

  36. [44]

    T. Xie, Z. Zong, Y . Qiu, et al., PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics , 2024

  37. [45]

    A. X. Chang, T. Funkhouser, L. Guibas, et al. , ShapeNet: An Information-Rich 3D Model Repository , 2015

  38. [46]

    ABO: Dataset and Benchmarks for Real-World 3D Object Understanding,

    J. Collins, S. Goel, K. Deng, et al., “ABO: Dataset and Benchmarks for Real-World 3D Object Understanding,” 2022

  39. [47]

    OmniObject3D: Large-V ocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation,

    T. Wu, J. Zhang, X. Fu, et al., “OmniObject3D: Large-V ocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation,” 2023

  40. [48]

    Objaverse-XL: A Universe of 10M+ 3D Objects,

    M. Deitke, R. Liu, M. Wallingford, et al., “Objaverse-XL: A Universe of 10M+ 3D Objects,” in Advances in Neural Information Processing Systems, 2023

  41. [49]

    Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation,

    H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich, “Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation,” 2023

  42. [50]

    DreamFusion: Text-to-3D using 2D Diffusion,

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “DreamFusion: Text-to-3D using 2D Diffusion,” 2022

  43. [51]

    A survey on the current state of the art on deep learning 3D reconstruction,

    B. Maxim and S. Nedevschi, “A survey on the current state of the art on deep learning 3D reconstruction,” in 2021 IEEE 17th International Conference on Intelligent Computer Communication and Processing (ICCP), 2021

  44. [52]

    A Survey of 3D Object Reconstruction Methods,

    M. G. Kantarci, B. Gökberk, and L. Akarun, “A Survey of 3D Object Reconstruction Methods,” in 2022 30th Signal Processing and Communications Applications Conference (SIU) , 2022

  45. [53]

    Y . Bai, L. Wong, and T. Twan,Survey on Fundamental Deep Learning 3D Reconstruction Techniques, 2024

  46. [54]

    Deep learning-based 3D reconstruction from multiple images: A survey,

    C. Wang, M. A. Reza, V . Vats, et al. , “Deep learning-based 3D reconstruction from multiple images: A survey,” Neurocomputing, 2024

  47. [55]

    A Critical Analysis of NeRF-Based 3D Reconstruction,

    F. Remondino, A. Karami, Z. Yan, G. Mazzacca, S. Rigon, and R. Qin, “A Critical Analysis of NeRF-Based 3D Reconstruction,” Remote Sensing, 2023

  48. [56]

    3D Gaussian as a New Era: A Survey,

    B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y . He, “3D Gaussian as a New Era: A Survey,” IEEE Transactions on Visualization and Computer Graphics, 2024

  49. [57]

    S. Zhu, G. Wang, X. Kong, D. Kong, and H. Wang, 3D Gaussian Splatting in Robotics: A Survey , 2024

  50. [58]

    Chen, A Review of Deep Learning-Powered Mesh Reconstruction Methods, 2023

    Z. Chen, A Review of Deep Learning-Powered Mesh Reconstruction Methods, 2023

  51. [59]

    What’s the Situation With Intelligent Mesh Generation: A Survey and Perspectives,

    N. Lei, Z. Li, Z. Xu, Y . Li, and X. Gu, “What’s the Situation With Intelligent Mesh Generation: A Survey and Perspectives,” IEEE Transactions on Visualization and Computer Graphics , 2024

  52. [60]

    M. Z. Irshad, M. Comi, Y .-C. Lin, et al., Neural Fields in Robotics: A Survey, 2024

  53. [61]

    Text-to-3D Generative AI on Mobile Devices: Measurements and Optimizations,

    X. Zhang, Z. Li, S. Oymak, and J. Chen, “Text-to-3D Generative AI on Mobile Devices: Measurements and Optimizations,” in Proceedings of the 2023 Workshop on Emerging Multimedia Systems , ser. EMS ’23, 2023

  54. [62]

    Vision-driven Compliant Manipulation for Reliable, High- Precision Assembly Tasks,

    A. S. Morgan, B. Wen, J. Liang, A. Boularias, A. M. Dollar, and K. Bekris, “Vision-driven Compliant Manipulation for Reliable, High- Precision Assembly Tasks,” in 17th Robotics: Science and Systems, RSS 2021, 2021

  55. [63]

    Semantic 3D Reconstruction for Robotic Manipulators with an Eye-In-Hand Vision System,

    F. Zha, Y . Fu, P. Wang, et al. , “Semantic 3D Reconstruction for Robotic Manipulators with an Eye-In-Hand Vision System,” Applied Sciences, 2020

  56. [64]

    Autonomous Robotic Assembly: From Part Singulation to Precise Assembly,

    K. Ota, D. K. Jha, S. Jain, et al., “Autonomous Robotic Assembly: From Part Singulation to Precise Assembly,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2024

  57. [65]

    Tolerance-Guided Policy Learning for Adaptable and Transferrable Delicate Industrial Insertion,

    B. Niu, C. Wang, and C. Liu, “Tolerance-Guided Policy Learning for Adaptable and Transferrable Delicate Industrial Insertion,” in Proceedings of the 2020 Conference on Robot Learning , 2021

  58. [66]

    J. X.-Y . Lim and Q.-C. Pham, Grasping, Part Identification, and Pose Refinement in One Shot with a Tactile Gripper , 2023

  59. [67]

    Perspectives of RealSense and ZED Depth Sensors for Robotic Vision Applications,

    V . Tadic, A. Toth, Z. Vizvari, et al., “Perspectives of RealSense and ZED Depth Sensors for Robotic Vision Applications,” Machines, 2022

  60. [68]

    Tochilkin, D

    D. Tochilkin, D. Pankratz, Z. Liu, et al., TripoSR: Fast 3D Object Reconstruction from a Single Image , 2024

  61. [69]

    SyncDreamer: Generating Multiview- consistent Images from a Single-view Image,

    Y . Liu, C. Lin, Z. Zeng, et al., “SyncDreamer: Generating Multiview- consistent Images from a Single-view Image,” 2023

  62. [70]

    Zero-1-to-3: Zero-shot One Image to 3D Object,

    R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick, “Zero-1-to-3: Zero-shot One Image to 3D Object,” 2023

  63. [71]

    Jiang, Q

    H. Jiang, Q. Huang, and G. Pavlakos, Real3D: Scaling Up Large Reconstruction Models with Real-World Images , 2024

  64. [72]

    PhyRecon: Physically Plausible Neural Scene Reconstruction,

    J. Ni, Y . Chen, B. Jing, et al., “PhyRecon: Physically Plausible Neural Scene Reconstruction,” 2024

  65. [73]

    Physically Compatible 3D Object Modeling from a Single Image,

    M. Guo, B. Wang, P. Ma, et al., “Physically Compatible 3D Object Modeling from a Single Image,” in Advances in Neural Information Processing Systems, 2024

  66. [74]

    Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication,

    Y . Chen, T. Xie, Z. Zong, et al., “Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication,” 2024

  67. [75]

    H. Yan, M. Zhang, Y . Li, C. Ma, and P. Ji, PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image , 2024

  68. [76]

    Learning Dual-Arm Push and Grasp Synergy in Dense Clutter,

    Y . Wang and H. Kasaei, “Learning Dual-Arm Push and Grasp Synergy in Dense Clutter,” IEEE Robotics and Automation Letters , 2025

  69. [77]

    vMAP: Vectorised Object Mapping for Neural Field SLAM,

    X. Kong, S. Liu, M. Taher, and A. J. Davison, “vMAP: Vectorised Object Mapping for Neural Field SLAM,” 2023

  70. [78]

    RICO: Regu- larizing the Unobservable for Indoor Compositional Reconstruction,

    Z. Li, X. Lyu, Y . Ding, M. Wang, Y . Liao, and Y . Liu, “RICO: Regu- larizing the Unobservable for Indoor Compositional Reconstruction,” 2023

  71. [79]

    Magic123: One Image to High- Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors,

    G. Qian, J. Mai, A. Hamdi, et al., “Magic123: One Image to High- Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors,” 2023

  72. [80]

    M. Boss, Z. Huang, A. Vasishta, and V . Jampani, SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement, 2024

  73. [81]

    One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion,

    M. Liu, R. Shi, L. Chen, et al., “One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion,” 2024

  74. [82]

    Real-time control in robotic systems,

    A. Simpkins, “Real-time control in robotic systems,” in Robotic Systems-Applications, Control and Programming , 2012, p. 231

  75. [83]

    Real-time camera tracking and 3D reconstruction using signed distance functions,

    E. Bylow, J. Sturm, C. Kerl, F. Kahl, and D. Cremers, “Real-time camera tracking and 3D reconstruction using signed distance functions,” in Robotics: Science and systems (RSS) conference 2013 , 2013

  76. [84]

    T. P. Swaminathan, C. Silver, and T. Akilan, Benchmarking Deep Learning Models on NVIDIA Jetson Nano for Real-Time Systems: An Empirical Investigation, 2024

  77. [85]

    Security and Privacy in Cloud Computing: Technical Review,

    Y . S. Abdulsalam and M. Hedabou, “Security and Privacy in Cloud Computing: Technical Review,” Future Internet, 2022

  78. [86]

    Surface Recon- struction From Point Clouds: A Survey and a Benchmark,

    Z. Huang, Y . Wen, Z. Wang, J. Ren, and K. Jia, “Surface Recon- struction From Point Clouds: A Survey and a Benchmark,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024

  79. [87]

    Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments,

    M. Mittal, C. Yu, Q. Yu, et al. , “Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments,” IEEE Robotics and Automation Letters , 2023

  80. [88]

    Image2Mesh: A Learning Framework for Single Image 3D Reconstruction,

    J. K. Pontes, C. Kong, S. Sridharan, S. Lucey, A. Eriksson, and C. Fookes, “Image2Mesh: A Learning Framework for Single Image 3D Reconstruction,” in Computer Vision – ACCV 2018 , 2019

  81. [89]

    Learning Free-Form Deformations for 3D Object Reconstruction,

    D. Jack, J. K. Pontes, S. Sridharan, et al. , “Learning Free-Form Deformations for 3D Object Reconstruction,” in Computer Vision – ACCV 2018, 2019

  82. [90]

    Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images,

    N. Wang, Y . Zhang, Z. Li, Y . Fu, W. Liu, and Y .-G. Jiang, “Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images,” 2018

  83. [91]

    MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers,

    Y . Siddiqui, A. Alliegro, A. Artemov, et al., “MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers,” 2024

  84. [92]

    Attention is All you Need,

    A. Vaswani, N. Shazeer, N. Parmar, et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems , 2017

  85. [93]

    A Papier-Mâché Approach to Learning 3D Surface Generation,

    T. Groueix, M. Fisher, V . G. Kim, B. C. Russell, and M. Aubry, “A Papier-Mâché Approach to Learning 3D Surface Generation,” 2018

  86. [94]

    Physically-aware Generative Network for 3D Shape Modeling,

    M. Mezghanni, M. Boulkenafed, A. Lieutier, and M. Ovsjanikov, “Physically-aware Generative Network for 3D Shape Modeling,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 17

  87. [95]

    Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation,

    R. Chen, Y . Chen, N. Jiao, and K. Jia, “Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation,” 2023

  88. [96]

    LGM: Large Multi-view Gaussian Model for High-Resolution 3D Content Creation,

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “LGM: Large Multi-view Gaussian Model for High-Resolution 3D Content Creation,” in Computer Vision – ECCV 2024 , 2025

  89. [97]

    Gaussian Splatting: 3D Reconstruction and Novel View Synthesis: A Review,

    A. Dalal, D. Hagen, K. G. Robbersmyr, and K. M. Knausgård, “Gaussian Splatting: 3D Reconstruction and Novel View Synthesis: A Review,” IEEE Access, 2024

  90. [98]

    IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Genera- tion,

    L. Melas-Kyriazi, I. Laina, C. Rupprecht, et al. , “IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Genera- tion,” in Proceedings of the 41st International Conference on Machine Learning, 2024

  91. [99]

    Splatter Image: Ultra-Fast Single-View 3D Reconstruction,

    S. Szymanowicz, C. Rupprecht, and A. Vedaldi, “Splatter Image: Ultra-Fast Single-View 3D Reconstruction,” 2024

  92. [100]

    GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation,

    Y . Xu, Z. Shi, W. Yifan,et al., “GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation,” in Computer Vision – ECCV 2024 , 2025

  93. [101]

    AGG: Amortized Generative 3D Gaussians for Single Image to 3D,

    D. Xu, Y . Yuan, M. Mardani, et al., “AGG: Amortized Generative 3D Gaussians for Single Image to 3D,” Transactions on Machine Learning Research, 2024

  94. [102]

    Szymanowicz, E

    S. Szymanowicz, E. Insafutdinov, C. Zheng, et al., Flash3D: Feed- Forward Generalisable 3D Scene Reconstruction from a Single Image , 2024

  95. [103]

    L. Liu, X. Wang, J. Qiu, T. Lin, X. Zhou, and Z. Su, Gaussian Object Carver: Object-Compositional Gaussian Splatting with surfaces completion, 2024

  96. [104]

    2D Gaussian Splatting for Geometrically Accurate Radiance Fields,

    B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2D Gaussian Splatting for Geometrically Accurate Radiance Fields,” in ACM SIGGRAPH 2024 Conference Papers , ser. SIGGRAPH ’24, 2024

  97. [105]

    Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes,

    Z. Yu, T. Sattler, and A. Geiger, “Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes,” ACM Trans. Graph., 2024

  98. [106]

    DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation,

    J. J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove, “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation,” 2019

  99. [107]

    Occupancy Networks: Learning 3D Reconstruction in Function Space,

    L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy Networks: Learning 3D Reconstruction in Function Space,” 2019

  100. [108]

    The SLAM prob- lem: A survey,

    J. Aulinas, Y . Petillot, J. Salvi, Llado, and Xavier, “The SLAM prob- lem: A survey,” in Artificial Intelligence Research and Development , 2008

  101. [109]

    Structure-From-Motion Revisited,

    J. L. Schonberger and J.-M. Frahm, “Structure-From-Motion Revisited,” 2016

  102. [110]

    GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting,

    C. Yan, D. Qu, D. Xu, et al., “GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting,” 2024

  103. [111]

    SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM,

    N. Keetha, J. Karhade, K. M. Jatavallabhula, et al., “SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM,” 2024

  104. [112]

    Gaussian Splatting SLAM,

    H. Matsuki, R. Murai, P. H. J. Kelly, and A. J. Davison, “Gaussian Splatting SLAM,” 2024

  105. [113]

    Beyond Point Clouds: Scene Understanding by Reasoning Geometry and Physics,

    B. Zheng, Y . Zhao, J. C. Yu, K. Ikeuchi, and S.-C. Zhu, “Beyond Point Clouds: Scene Understanding by Reasoning Geometry and Physics,” in 2013 IEEE Conference on Computer Vision and Pattern Recognition , 2013

  106. [114]

    MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction,

    Z. Yu, S. Peng, M. Niemeyer, T. Sattler, and A. Geiger, “MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction,” in Advances in Neural Information Processing Systems, 2022

  107. [115]

    3D- R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction,

    C. B. Choy, D. Xu, J. Gwak, K. Chen, and S. Savarese, “3D- R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction,” in Computer Vision – ECCV 2016 , 2016

  108. [116]

    Learning a Multi-View Stereo Machine,

    A. Kar, C. Häne, and J. Malik, “Learning a Multi-View Stereo Machine,” in Advances in Neural Information Processing Systems , 2017

  109. [117]

    Pix2V ox: Context- Aware 3D Reconstruction From Single and Multi-View Images,

    H. Xie, H. Yao, X. Sun, S. Zhou, and S. Zhang, “Pix2V ox: Context- Aware 3D Reconstruction From Single and Multi-View Images,” 2019

  110. [118]

    Auto-Encoding Variational Bayes,

    D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in Proceedings of the International Conference on Learning Repre- sentations, 2014

  111. [119]

    pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction,

    D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction,” 2024

  112. [120]

    Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction,

    H. Chen, J. Gu, A. Chen, et al., “Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction,” 2023

  113. [121]

    A Point Set Generation Network for 3D Object Reconstruction From a Single Image,

    H. Fan, H. Su, and L. J. Guibas, “A Point Set Generation Network for 3D Object Reconstruction From a Single Image,” 2017

  114. [122]

    Transformers as Meta-learners for Implicit Neural Representations,

    Y . Chen and X. Wang, “Transformers as Meta-learners for Implicit Neural Representations,” in Computer Vision – ECCV 2022 , 2022

  115. [123]

    Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model,

    J. Li, H. Tan, K. Zhang, et al. , “Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model,” 2023

  116. [124]

    Shape, Pose, and Appearance From a Single Image via Bootstrapped Radiance Field Inversion,

    D. Pavllo, D. J. Tan, M.-J. Rakotosaona, and F. Tombari, “Shape, Pose, and Appearance From a Single Image via Bootstrapped Radiance Field Inversion,” 2023

  117. [125]

    Lee and A

    H.-H. Lee and A. X. Chang, Understanding Pure CLIP Guidance for Voxel Grid NeRF Models , 2022

  118. [126]

    Zero- Shot Text-Guided Object Generation With Dream Fields,

    A. Jain, B. Mildenhall, J. T. Barron, P. Abbeel, and B. Poole, “Zero- Shot Text-Guided Object Generation With Dream Fields,” 2022

  119. [127]

    MVDream: Multi-view Diffusion for 3D Generation,

    Y . Shi, P. Wang, J. Ye, L. Mai, K. Li, and X. Yang, “MVDream: Multi-view Diffusion for 3D Generation,” 2023

  120. [128]

    Y . Chen, R. Xie, Q. Ye, et al., 2L3: Lifting Imperfect Generated 2D Images into Accurate 3D , 2024

  121. [129]

    Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image,

    K. Wu, F. Liu, Z. Cai, et al., “Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image,” 2024

  122. [130]

    Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era,

    X.-F. Han, H. Laga, and M. Bennamoun, “Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021

  123. [131]

    U-Net: Convolutional Networks for Biomedical Image Segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 , 2015

  124. [132]

    DINOv2: Learning Robust Visual Features without Supervision,

    M. Oquab, T. Darcet, T. Moutakanni, et al. , “DINOv2: Learning Robust Visual Features without Supervision,” Transactions on Machine Learning Research, 2023

  125. [133]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” 2016

  126. [134]

    Jun and A

    H. Jun and A. Nichol, Shap-E: Generating Conditional 3D Implicit Functions, 2023

  127. [135]

    LRM: Large Reconstruction Model for Single Image to 3D,

    Y . Hong, K. Zhang, J. Gu, et al., “LRM: Large Reconstruction Model for Single Image to 3D,” 2023

  128. [136]

    Sharf: Shape- conditioned Radiance Fields from a Single View,

    K. Rematas, R. Martin-Brualla, and V . Ferrari, “Sharf: Shape- conditioned Radiance Fields from a Single View,” in Proceedings of the 38th International Conference on Machine Learning , 2021

  129. [137]

    Generative Adversarial Nets,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, et al. , “Generative Adversarial Nets,” in Advances in Neural Information Processing Systems, 2014

  130. [138]

    HoloGAN: Unsupervised Learning of 3D Representations From Natural Images,

    T. Nguyen-Phuoc, C. Li, L. Theis, C. Richardt, and Y . -L. Yang, “HoloGAN: Unsupervised Learning of 3D Representations From Natural Images,” 2019

  131. [139]

    Image GANs meet Differen- tiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering,

    Y . Zhang, W. Chen, H. Ling, et al. , “Image GANs meet Differen- tiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering,” 2020

  132. [140]

    Pix2Scene: Learning Implicit 3D Representations from Images,

    S. Rajeswar, F. Mannan, F. Golemo, D. Vazquez, D. Nowrouzezahrai, and A. Courville, “Pix2Scene: Learning Implicit 3D Representations from Images,” 2018

  133. [141]

    ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Niessner, “ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes,” 2017

  134. [142]

    ZeroShape: Regression-based Zero-shot Shape Reconstruction,

    Z. Huang, S. Stojanov, A. Thai, V . Jampani, and J. M. Rehg, “ZeroShape: Regression-based Zero-shot Shape Reconstruction,” 2024

  135. [143]

    DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction Model,

    Y . Xu, H. Tan, F. Luan, et al. , “DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction Model,” 2023

  136. [144]

    PF-LRM: Pose-Free Large Recon- struction Model for Joint Pose and Shape Prediction,

    P. Wang, H. Tan, S. Bi, et al. , “PF-LRM: Pose-Free Large Recon- struction Model for Joint Pose and Shape Prediction,” 2023

  137. [145]

    Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers,

    Z.-X. Zou, Z. Yu, Y . -C. Guo, et al. , “Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers,” 2024

  138. [146]

    Learning Transferable Visual Models From Natural Language Supervision,

    A. Radford, J. W. Kim, C. Hallacy, et al. , “Learning Transferable Visual Models From Natural Language Supervision,” in Proceedings of the 38th International Conference on Machine Learning , 2021

  139. [147]

    AvatarCLIP: Zero-shot text-driven generation and animation of 3D avatars,

    F. Hong, M. Zhang, L. Pan, Z. Cai, L. Yang, and Z. Liu, “AvatarCLIP: Zero-shot text-driven generation and animation of 3D avatars,” ACM Trans. Graph., 2022

  140. [148]

    CLIP-Mesh: Generating textured meshes from text using pretrained image-text models,

    N. M. Khalid, T. Xie, E. Belilovsky, and T. Popa, “CLIP-Mesh: Generating textured meshes from text using pretrained image-text models,” in SIGGRAPH Asia 2022 Conference Papers , 2022

  141. [149]

    When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?,

    M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?,” 2022

  142. [150]

    NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors,

    J. Wang, P. Wang, X. Long, et al., “NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors,” in Computer Vision – ECCV 2022, 2022

  143. [151]

    DN-Splatter: Depth and Normal Priors for Gaussian 18 Splatting and Meshing,

    M. Turkulainen, X. Ren, I. Melekhov, O. Seiskari, E. Rahtu, and J. Kannala, “DN-Splatter: Depth and Normal Priors for Gaussian 18 Splatting and Meshing,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , 2025

  144. [152]

    Szymanowicz, J

    S. Szymanowicz, J. Y . Zhang, P. Srinivasan,et al., Bolt3D: Generating 3D Scenes in Seconds , 2025

  145. [153]

    Dogaru, M

    A. Dogaru, M. Özer, and B. Egger, Generalizable 3D Scene Recon- struction via Divide and Conquer from a Single View , 2024

  146. [154]

    Huang, Y .-C

    Z. Huang, Y .-C. Guo, X. An, et al., MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation , 2024

  147. [155]

    H. Han, R. Yang, H. Liao, et al., REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment , 2024

  148. [156]

    Object-Compositional Neural Implicit Surfaces,

    Q. Wu, X. Liu, Y . Chen,et al., “Object-Compositional Neural Implicit Surfaces,” in Computer Vision – ECCV 2022 , 2022

  149. [157]

    Holistic 3D Scene Understanding From a Single Image With Implicit Representation,

    C. Zhang, Z. Cui, Y . Zhang, B. Zeng, M. Pollefeys, and S. Liu, “Holistic 3D Scene Understanding From a Single Image With Implicit Representation,” 2021

  150. [158]

    Graph- Dreamer: Compositional 3D Scene Synthesis from Scene Graphs,

    G. Gao, W. Liu, A. Chen, A. Geiger, and B. Schölkopf, “Graph- Dreamer: Compositional 3D Scene Synthesis from Scene Graphs,” 2024

  151. [159]

    3D-Scene-Former: 3D scene generation from a single RGB image using Transformers,

    J. Chatterjee and M. Torres Vega, “3D-Scene-Former: 3D scene generation from a single RGB image using Transformers,” The Visual Computer, 2024

  152. [160]

    ClusteringSDF: Self- Organized Neural Implicit Surfaces for 3D Decomposition,

    T. Wu, C. Zheng, Q. Wu, and T. -J. Cham, “ClusteringSDF: Self- Organized Neural Implicit Surfaces for 3D Decomposition,” in Computer Vision – ECCV 2024 , 2025

  153. [161]

    Hassena, J

    G. Hassena, J. Moon, R. Fujii, et al., ObjectCarver: Semi-automatic segmentation, reconstruction and separation of 3D objects , 2024

  154. [162]

    Physical Simulation Layer for Accurate 3D Modeling,

    M. Mezghanni, T. Bodrito, M. Boulkenafed, and M. Ovsjanikov, “Physical Simulation Layer for Accurate 3D Modeling,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  155. [163]

    Z. Chen, A. Walsman, M. Memmel, et al., URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images, 2024

  156. [164]

    L. Le, J. Xie, W. Liang, et al. , Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model, 2024

  157. [165]

    DragAPart: Learning a Part-Level Motion Prior for Articulated Objects,

    R. Li, C. Zheng, C. Rupprecht, and A. Vedaldi, “DragAPart: Learning a Part-Level Motion Prior for Articulated Objects,” in Computer Vision – ECCV 2024 , 2025

  158. [166]

    PIE-NeRF: Physics-based Interactive Elastodynamics with NeRF,

    Y . Feng, Y . Shang, X. Li, T. Shao, C. Jiang, and Y . Yang, “PIE-NeRF: Physics-based Interactive Elastodynamics with NeRF,” 2024

  159. [167]

    Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items,

    L. Downs, A. Francis, N. Koenig, et al., “Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items,” in 2022 International Conference on Robotics and Automation (ICRA) , 2022

  160. [168]

    Jaunet, G

    T. Jaunet, G. Bono, R. Vuillemot, and C. Wolf, SIM2REALVIZ: Visualizing the Sim2Real Gap in Robot Ego-Pose Estimation , 2021

  161. [169]

    Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation,

    M. Khanna, Y . Mao, H. Jiang,et al., “Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation,” 2024

  162. [170]

    Kolve, R

    E. Kolve, R. Mottaghi, W. Han, et al., AI2-THOR: An Interactive 3D Environment for Visual AI , 2022

  163. [171]

    ProcTHOR: Large-Scale Embodied AI Using Procedural Generation,

    M. Deitke, E. VanderBilt, A. Herrasti, et al., “ProcTHOR: Large-Scale Embodied AI Using Procedural Generation,” Advances in Neural Information Processing Systems , 2022

  164. [172]

    Matterport3D: Learning from RGB-D Data in Indoor Environments,

    A. Chang, A. Dai, T. Funkhouser, et al. , “Matterport3D: Learning from RGB-D Data in Indoor Environments,” 2017

  165. [173]

    ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data,

    G. Baruch, Z. Chen, A. Dehghan, et al., “ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data,” 2021

  166. [174]

    The YCB object and Model set: Towards common benchmarks for manipulation research,

    B. Calli, A. Singh, A. Walsman, S. Srinivasa, P. Abbeel, and A. M. Dollar, “The YCB object and Model set: Towards common benchmarks for manipulation research,” in 2015 International Conference on Advanced Robotics (ICAR) , 2015

  167. [175]

    Self-supervised 6D Object Pose Estimation for Robot Manipulation,

    X. Deng, Y . Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox, “Self-supervised 6D Object Pose Estimation for Robot Manipulation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), 2020

  168. [176]

    Systematic object- invariant in-hand manipulation via reconfigurable underactuation: Introducing the RUTH gripper,

    Q. Lu, N. Baron, A. B. Clark, and N. Rojas, “Systematic object- invariant in-hand manipulation via reconfigurable underactuation: Introducing the RUTH gripper,” The International Journal of Robotics Research, 2021

  169. [177]

    Robotic bin-picking: Benchmarking robotics grippers with modified YCB object and model set,

    T. Lerher, P. Bencak, D. Hercog, B. Jerman, and L. Bizjak, “Robotic bin-picking: Benchmarking robotics grippers with modified YCB object and model set,” Progress in Material Handling Research , 2023

  170. [178]

    LiDAR- based SLAM for robotic mapping: State of the art and new frontiers,

    X. Yue, Y . Zhang, J. Chen, J. Chen, X. Zhou, and M. He, “LiDAR- based SLAM for robotic mapping: State of the art and new frontiers,” Industrial Robot: the international journal of robotics research and application, 2024

  171. [179]

    A review of visual SLAM for robotics: Evolution, properties, and future applications,

    B. Al-Tawil, T. Hempel, A. Abdelrahman, and A. Al-Hamadi, “A review of visual SLAM for robotics: Evolution, properties, and future applications,” Frontiers in Robotics and AI , 2024

  172. [180]

    Neural radiance fields in the industrial and robotics domain: Applica- tions, research opportunities and use cases,

    E. Šlapak, E. Pardo, M. Dopiriak, T. Maksymyuk, and J. Gazda, “Neural radiance fields in the industrial and robotics domain: Applica- tions, research opportunities and use cases,” Robotics and Computer- Integrated Manufacturing, 2024

  173. [181]

    Abou-Chakra, K

    J. Abou-Chakra, K. Rana, F. Dayoub, and N. Sünderhauf, Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics, 2024

  174. [182]

    OpenAI et al., GPT-4 Technical Report, 2024

  175. [183]

    ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation,

    G. Lu, S. Zhang, Z. Wang, C. Liu, J. Lu, and Y . Tang, “ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation,” in Computer Vision – ECCV 2024 , 2025

  176. [184]

    T. Wu, C. Zheng, F. Guan, A. Vedaldi, and T. -J. Cham, Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images , 2025

  177. [185]

    BUOL: A Bottom-Up Framework With Occupancy-Aware Lifting for Panoptic 3D Scene Reconstruction From a Single Image,

    T. Chu, P. Zhang, Q. Liu, and J. Wang, “BUOL: A Bottom-Up Framework With Occupancy-Aware Lifting for Panoptic 3D Scene Reconstruction From a Single Image,” 2023

  178. [186]

    CLAY: A Controllable Large- scale Generative Model for Creating High-quality 3D Assets,

    L. Zhang, Z. Wang, Q. Zhang, et al., “CLAY: A Controllable Large- scale Generative Model for Creating High-quality 3D Assets,” ACM Trans. Graph., 2024

  179. [187]

    H. Wu, M. G. Karumuri, C. Zou, et al. , Direct and Explicit 3D Generation from a Single Image , 2024

  180. [188]

    DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior,

    J. Sun, B. Zhang, R. Shao, et al., “DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior,” 2023

  181. [189]

    DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation,

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng, “DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation,” 2023

  182. [190]

    R. Li, C. Zheng, C. Rupprecht, and A. Vedaldi, DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness , 2025

  183. [191]

    Q. Shen, Z. Wu, X. Yi, et al., Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction , 2024

  184. [192]

    Large Point-to-Gaussian Model for Image-to-3D Generation,

    L. Lu, H. Gao, T. Dai, et al. , “Large Point-to-Gaussian Model for Image-to-3D Generation,” in Proceedings of the 32nd ACM International Conference on Multimedia , ser. MM ’24, 2024

  185. [193]

    X. Wei, K. Zhang, S. Bi, et al. , MeshLRM: Large Reconstruction Model for High-Quality Meshes , 2025

  186. [194]

    Multiview Compressive Coding for 3D Reconstruction,

    C.-Y . Wu, J. Johnson, J. Malik, C. Feichtenhofer, and G. Gkioxari, “Multiview Compressive Coding for 3D Reconstruction,” 2023

  187. [195]

    Object-Aware 3D Scene Reconstruction from Single 2D Images of Indoor Scenes,

    M. Wen and K. Cho, “Object-Aware 3D Scene Reconstruction from Single 2D Images of Indoor Scenes,” Mathematics, 2023

  188. [196]

    B. Chen, H. Jiang, S. Liu, et al., PhysGen3D: Crafting a Miniature Interactive World from a Single Image , 2025

  189. [197]

    Single-view 3D Scene Reconstruction with High-fidelity Shape and Texture,

    Y . Chen, J. Ni, N. Jiang, Y . Zhang, Y . Zhu, and S. Huang, “Single-view 3D Scene Reconstruction with High-fidelity Shape and Texture,” in 2024 International Conference on 3D Vision (3DV) , 2024

  190. [198]

    SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image Using Latent Video Diffusion,

    V . V oleti, C.-H. Yao, M. Boss, et al. , “SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image Using Latent Video Diffusion,” in Computer Vision – ECCV 2024 , 2025

  191. [199]

    VPP: Efficient Conditional 3D Generation via V oxel-Point Progressive Representation,

    Z. Qi, M. Yu, R. Dong, and K. Ma, “VPP: Efficient Conditional 3D Generation via V oxel-Point Progressive Representation,” in Advances in Neural Information Processing Systems , 2023

  192. [200]

    Wonder3D: Single Image to 3D using Cross-Domain Diffusion,

    X. Long, Y .-C. Guo, C. Lin, et al., “Wonder3D: Single Image to 3D using Cross-Domain Diffusion,” 2024

  193. [201]

    Efficient Geometry-Aware 3D Generative Adversarial Networks,

    E. R. Chan, C. Z. Lin, M. A. Chan, et al., “Efficient Geometry-Aware 3D Generative Adversarial Networks,” 2022

  194. [202]

    Deep Marching Tetrahedra: A Hybrid Representation for High-Resolution 3D Shape Synthesis,

    T. Shen, J. Gao, K. Yin, M. -Y . Liu, and S. Fidler, “Deep Marching Tetrahedra: A Hybrid Representation for High-Resolution 3D Shape Synthesis,” in Advances in Neural Information Processing Systems , 2021

  195. [203]

    NeRF: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , 2022

  196. [204]

    Xiang, Z

    J. Xiang, Z. Lv, S. Xu, et al., Structured 3D Latents for Scalable and Versatile 3D Generation, 2025

  197. [205]

    Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,

    M. FISCHLER AND, “Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,” Commun. ACM, 1981

  198. [206]

    Online 3D Scene Reconstruction Using Neural Object Priors,

    T. Chabal, S. Chen, J. Ponce, and C. Schmid, “Online 3D Scene Reconstruction Using Neural Object Priors,” in 3DV 2025 - 12th International Conference on 3D Vision 2025 , 2025

  199. [207]

    VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality,

    Y . Jiang, C. Yu, T. Xie, et al., “VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality,” in ACM SIGGRAPH 2024 Conference Papers , ser. SIGGRAPH ’24, 2024. 19

  200. [208]

    Complete Object-Compositional Neural Implicit Surfaces With 3D Pseudo Supervision,

    W. Kim, J. Park, and K. Cho, “Complete Object-Compositional Neural Implicit Surfaces With 3D Pseudo Supervision,” IEEE Access, 2025

  201. [209]

    ReconFusion: 3D Recon- struction with Diffusion Priors,

    R. Wu, B. Mildenhall, P. Henzler, et al., “ReconFusion: 3D Recon- struction with Diffusion Priors,” 2024

  202. [210]

    MetaGraspNet: A Large-Scale Benchmark Dataset for Scene-Aware Ambidextrous Bin Picking via Physics-based Metaverse Synthesis,

    M. Gilles, Y . Chen, T. Robin Winter, E. Zhixuan Zeng, and A. Wong, “MetaGraspNet: A Large-Scale Benchmark Dataset for Scene-Aware Ambidextrous Bin Picking via Physics-based Metaverse Synthesis,” in 2022 IEEE 18th International Conference on Automation Science and Engineering ...

  203. [211]

    FCL: A general purpose library for collision and proximity queries,

    J. Pan, S. Chitta, and D. Manocha, “FCL: A general purpose library for collision and proximity queries,” in 2012 IEEE International Conference on Robotics and Automation , 2012

  204. [212]

    Indoor Segmenta- tion and Support Inference from RGBD Images,

    N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor Segmenta- tion and Support Inference from RGBD Images,” in Computer Vision – ECCV 2012 , 2012

  205. [213]

    D$^3$RoMa: Disparity Diffusion- based Depth Sensing for Material-Agnostic Robotic Manipulation,

    S. Wei, H. Geng, J. Chen, et al., “D$^3$RoMa: Disparity Diffusion- based Depth Sensing for Material-Agnostic Robotic Manipulation,” 2024

  206. [214]

    Visual Robotic Manipulation with Depth-Aware Pretraining,

    J. Li, W. Wang, Y . Peng, C. Shen, Y . Zhu, and Z. Xu, “Visual Robotic Manipulation with Depth-Aware Pretraining,” in 2024 IEEE International Conference on Robotics and Biomimetics (ROBIO), 2024

  207. [215]

    Leveraging depth data in remote robot teleoperation interfaces for general object manipulation,

    D. Kent, C. Saldanha, and S. Chernova, “Leveraging depth data in remote robot teleoperation interfaces for general object manipulation,” The International Journal of Robotics Research , 2020

  208. [216]

    Deep Depth Completion of a Single RGB-D Image,

    Y . Zhang and T. Funkhouser, “Deep Depth Completion of a Single RGB-D Image,” 2018

  209. [217]

    A comprehensive survey of depth completion approaches,

    M. A. U. Khan, D. Nazir, A. Pagani, et al., “A comprehensive survey of depth completion approaches,” Sensors, 2022. VI. B IOGRAPHY SECTION Frederik Nolte is an ELLIS PhD student currently pursuing a DPhil in Engineering Science in the Applied AI Lab at the Oxford Robotics Inst...

  210. [218]

    [153] [196] No Collision Collision Fig. 13. Collision frequency in scene reconstruction on YCB-Video [38]. Even though scene reconstruction models have access to information about the entire scene, mesh collisions in 3D reconstruction are common. PhysGen3D

  211. [219]

    shows fewer collisions, potentially due to overestimating object distances (see Fig. 11)

  212. [220]

    [27] [26] [96] [28] [142] [71] [190] 0 1 2 3 4 15Chamfer Distance (cm) Fig. 14. Chamfer distances on Aria Digital Twin [39] dataset. Reconstruction errors on this dataset are noticeably higher than those shown in Fig. 2. Reconstructed surfaces tend to be 1cm distant from the c...

  213. [221]

    [27] [26] [96] [28] [142] [71] [190] 0 10 20 30 40 100 160Number of Meshes Stable Unstable Fig. 15. Object stability on Aria Digital Twin [39] dataset. Most reconstructions lack stability within 5 ∘ of ground truth poses. In physics simulation, unstable geometries would cause ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.