Pith. sign in

REVIEW 3 major objections 5 minor 59 references

DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read DLO-Splatting tracks a rope's 3D shape by combining physics prediction with Gaussian-splatting image updates.

desk verdict Novel PBD-plus-3DGS combination for DLO tracking, but the central prediction equation is dimensionally off, and the single demo doesn't back the claims. read the letter →

arxiv 2505.08644 v2 pith:KPY7WGVQ submitted 2025-05-13 cs.CV cs.RO

classification cs.CVcs.RO
keywords deformablelinearobjects3DGaussiansplattingposition-baseddynamicsstateestimationknottyingmulti-viewRGBprediction-updatefiltering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

DLO-Splatting estimates the 3D shape of a deformable linear object such as a rope from multi-view RGB images and the gripper pose, using a prediction-update loop that needs no learned dynamics and no fiducial markers. The prediction step evolves a chain of nodes with position-based dynamics under gravity, planar contact, friction, and fixed segment-length constraints. The update step renders the predicted rope as spherical 3D Gaussians, compares the render to each camera image, and adjusts the node positions by stochastic gradient descent on the rendering loss. The demonstration, a cross move used in knot tying, shows the algorithm tracking the grasped rope tip more accurately than a vision-only baseline, though neither method recovers the rope's final shape and topology.

What carries the argument

The central machinery is a Bayesian-style prediction-update filter whose state is a node chain $X_t \in \mathbb{R}^{N\times 3}$. The prediction step uses position-based dynamics: Verlet integration with gravity, a planar contact normal, friction, and a length-constraint projection that keeps every segment at length $L$. The update step places $d$ spherical 3D Gaussians per segment along the centerline, renders them into each camera with $\alpha$-blending, and minimizes the squared image difference $L_{\text{obs}}$ by stochastic gradient descent on the node positions. The Gaussians' covariance is tied to the rope diameter and their positions to the nodes, so the visual loss directly drives geometric correction of the predicted rope.

What would settle it

Record a full knot-tying sequence in which one segment of the rope is marked with a distinctive color, so the over/under order at each crossing is visible in the images. If, at the moment a crossing is formed, the tracked node chain places the marked segment on the wrong side while the rendered rope still matches all three camera views, then the claim that Gaussian-splatting rendering disambiguates topology is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that combining a first-principles physics prediction with a differentiable Gaussian-splatting renderer produces a state estimate for a deformable linear object that is robust to visual ambiguities that defeat vision-only trackers, specifically dense crossings and self-occlusion during knot tying. The algorithm represents the rope as $N$ nodes, predicts the next node configuration $\hat{X}_{t+1}$ with position-based dynamics, then refines it by minimizing the observation loss $L_{\text{obs}} = \|\mathbf{I}^t_k - \tilde{\mathbf{I}}^t_k\|^2$ against all $K$ camera views, where the rendered image $\tilde{\mathbf{I}}^t_k$ is computed by $\alpha$-blending spherical Gaussians placed along the rope centerline. Because the Gaussians are tied to the nodes, the gradient of the rendering loss updates the 3D geometry directly. The authors claim this lets the filter track visually complex topologies that cannot be disambiguated by vision alone, and they demonstrate the advantage on the grasped rope tip during a knot-tying cross move.

Load-bearing premise

The load-bearing premise is that a simple physics prediction—gravity, table contact, friction, and fixed segment lengths—moves the rope estimate close enough to the true state that the image-based correction can converge, even while the rope crosses over itself during knot tying.

Editorial extensions

If this is right

  • A robot could estimate its rope's 3D state during manipulation using only RGB cameras and its own gripper pose, with no learned dynamics model and no markers.
  • The same prediction-update loop should transfer to new rope materials or environments without retraining, provided mass, friction, and contact parameters are known.
  • If the rendering loss truly disambiguates crossing topology, multi-view RGB alone can maintain correct over/under structure where 2D or single-view trackers fail.
  • The demonstrated 1 Hz update rate and lack of self-intersection modeling currently restrict the method to slow motions and sparse self-contact; faster rendering and explicit self-contact prediction would be needed for full-speed knot tying.
  • Because the update uses object masking, the method can in principle be extended to simultaneous tracking of several DLOs once multi-instance segmentation is added.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct ablation varying the number of cameras (for example, two vs three) would test whether the reported tip-tracking advantage comes from multi-view disambiguation or from the physics prediction alone; the paper does not report this.
  • The fixed segment-length position-based dynamics model behaves like a discrete inextensible rod without bending stiffness; comparing against a rod model with bending and twisting energy would isolate whether the final topology collapse is a prediction error or an update error.
  • The 1 Hz bottleneck is likely shared between the physics time step and the stochastic gradient descent rendering update; a test that runs prediction at 100 Hz while keeping update at 1 Hz would show which side limits the filter.
  • The claim that vision alone cannot disambiguate these topologies would be sharpened by a controlled experiment with two visually identical crossings, where the only distinguishing information is the camera viewpoint; the paper's single cross-move demonstration is suggestive but not exhaustive.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes DLO-Splatting, a prediction-update filter for estimating the 3D state of a deformable linear object from multi-view RGB images and gripper pose. The prediction step uses position-based dynamics with gravity, planar contact, friction, and fixed segment-length constraints; the update step renders the predicted node chain with 3D Gaussian splatting and optimizes a rendering loss with SGD. The method is evaluated in one qualitative knot-tying 'cross move' with three cameras and compared against TrackDLO. The authors report that DLO-Splatting tracks the grasped rope tip more accurately than TrackDLO, but that both methods fail to recover the rope's final shape and topology.

Significance. If fully validated, the contribution would be useful: DLO-Splatting combines a training-free physics prediction with a differentiable rendering update, which is a reasonable way to address occlusion and multi-view disambiguation for DLO tracking. The paper also clearly enumerates its limitations, including unmodeled self-intersections, a 1 Hz update rate, and sensitivity to occlusion. However, the central claim that the method tracks visually complex topologies through dense knotting is not supported by the current evidence: Eq. (2) is not reproducible as written, the evaluation is a single qualitative trial without metrics, and the paper itself admits both methods fail on the final topology. The idea merits further work, but the manuscript as submitted is not ready for publication.

major comments (3)
  1. [Section III-A, Eq. (2)] Equation (2) is not Verlet integration and is dimensionally inconsistent as written: X^{t+1} = X^t + (X^t - X^{t-1})/Δt + (1/2)F^t Δt^2 adds a velocity-like term to positions, and the force term has units of mass·length unless F^t is already divided by node mass, which Eq. (3) contradicts because the forces include m_i g. The standard Verlet update for a node of mass m_i is x^{t+1} = 2x^t - x^{t-1} + (F^t/m_i)Δt^2, or equivalently x^{t+1} = x^t + v^t Δt + (1/2)(F^t/m_i)Δt^2. This is load-bearing because the PBD output is the initial condition for the rendering update in Eqs. (14)-(15), so any change in the integration formula changes the entire filter. The one qualitative success reported in Section IV (tracking the grasped tip) is controlled by the directly imposed gripper position in Eq. (1), so it does not validate the physics integration. The authors need to provide a corrected, dimensionally consistent integration step and state the mass normalization explicitly.
  2. [Section IV and Figure 2] The evaluation does not support the abstract's and Introduction's claim that DLO-Splatting tracks visually complex topologies that cannot be disambiguated by vision-based tracking alone. The demonstration is a single cross move with three cameras, presented only qualitatively: there are no quantitative metrics such as node-to-ground-truth distance or topology error, no error bars, no ablations, and no sensitivity analysis. The paper's own conclusion states that both DLO-Splatting and TrackDLO fail to recover the rope's final shape and topology, and Section V lists unmodeled self-intersections, a 1 Hz update rate, and occlusion from the gripper as limitations. The only reported advantage, tracking the grasped tip during one move, is too narrow to establish the central contribution, especially because the grasped tip is directly constrained by the gripper pose.
  3. [Section III-A and Section IV] The algorithm is not reproducible from the text because the values of the free parameters used in the demonstration are not reported: friction coefficient μ_f, integration time step Δt, Gaussian diameter σ, node count N and segment length L, Gaussians per segment d, and total rope mass m. The optimizer hyperparameters for the rendering update (learning rate, number of SGD iterations, initialization) are also missing. In addition, the relationship between the constraint projection in Eqs. (4)-(6), which is written as an update to X^t, and the length correction in Eqs. (8)-(9), which is applied to X^{t+1}, is unclear, so a reader cannot reimplement the exact algorithm. Without these details, the single demonstration cannot be checked or extended.
minor comments (5)
  1. [Section III-B, Eq. (11)] The rendering function h is written as a function of X^t_PBD and the camera projection P_k, but Eqs. (12)-(13) also depend on per-Gaussian colors c_j and opacities o_j; the text should specify how these are initialized and whether they are optimized or fixed during the update.
  2. [Section III-A, Eq. (3)] The velocity v^t_{i,xy} in the friction term is not defined; the text should state whether it is computed from the current and previous node positions and how the planar component is extracted.
  3. [Section III-A, Eq. (8)] Equation (8) is ambiguous without parentheses; it should read Δl^{t+1}_i = (L - l^{t+1}_i) · (x^{t+1}_{i+1} - x^{t+1}_i)/l^{t+1}_i to clarify that the scalar length difference multiplies the unit direction vector.
  4. [Section III-A, Eq. (2)] The gripper action a_t is introduced in the text but does not appear explicitly in Eq. (2); the authors should clarify how the action influences the free nodes as opposed to the grasped node.
  5. [Section IV] The paper should report the image resolution used for the rendering loss, the number of rendering iterations per update, and the actual per-step runtime; Section V mentions a 1 Hz update rate, but the evaluation section does not provide these details.

Circularity Check

1 steps flagged · score 6.0 of 10

Positive demo result (grasped-tip tracking) is forced by Eq. 1 gripper-input construction; core filter derivation is not circular.

  1. fitted input called prediction [Section III-A Eq. (1) and Section IV / Fig. 2 caption]
    "The position of the grasped node, xt g, is computed as the closest node in the set of nodes describing the shape of the rope, Xt, to the position of the center of the gripper, pt g ... When the grasped node moves to the new gripper position, the position correction xt+1 g = pt g + atΔt must also be propagated through the rope. ... DLO-Splatting succeeds in tracking the grasped tip."

    Equation (1) defines the grasped node as the rope node nearest the gripper center, and the algorithm directly imposes its next position from the gripper pose. The demonstration's only reported success is that DLO-Splatting 'succeeds in tracking the grasped tip' (Fig. 2). Because the tip trajectory is prescribed by the gripper state input, this success is guaranteed by construction: the output 'tracked tip' is the input gripper pose, not an inference from the rendering loss or the dynamics. It therefore cannot validate the claim that multi-view synthesis disambiguates topology, especially since the paper states both methods 'failed to estimate the correct topology.'

full rationale

The core estimator is a genuine predict-update filter: PBD predicts a candidate state, and the rendering loss Lobs = ||I_k^t - \tilde I_k^t||^2 against external camera images is minimized over a state correction ΔX (Eqs. 14-15). This uses observations that are not constructed from the target, so the derivation of the filter is not circular. The only circular element is the demonstration's positive result. The grasped tip is defined as the closest node to the gripper (Eq. 1) and its motion is directly imposed from gripper pose, so reporting that DLO-Splatting 'tracks the grasped tip' is an input-output identity, not an independent tracking success. The paper's own conclusion narrows the achievement: self-intersections are not modeled, update runs at 1 Hz, and both methods fail to recover the rope's final shape and topology, so the headline topology claims are unsupported but that is a scope/correctness issue. The TrackDLO initialization is a self-citation, but it is used symmetrically as a starting point for both methods and does not enter the derivation. Eq. (2) appears dimensionally inconsistent with Verlet integration; that is a reproducibility/correctness defect, not circularity. Overall, partial circularity exists only in the evaluation's 'grasped tip' claim, not in the state-estimation derivation.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The method rests on standard simulation and rendering components, but several hand-tuned parameters (friction, time step, Gaussian diameter, discretization) and domain assumptions (accurate gripper kinematics, reliable segmentation, no self-intersection) are not validated. The central result, a qualitative tracking improvement, is therefore fragile and not benchmarked against an external standard.

free parameters (6)
  • Friction coefficient μf
    Appears in the friction term of Eq. 3; no value is specified, and contact behavior depends on it.
  • Integration time step Δt
    Used in the Verlet integration of Eq. 2; no value is given.
  • Gaussian diameter σ
    Set equal to the rope diameter in Section III.B; the diameter value is not reported.
  • Node count N and segment length L
    The rope discretization determines the PBD length constraint and the number of Gaussians; not reported.
  • Gaussians per segment d
    Number of Gaussians placed on each rope segment in Section III.B; not reported.
  • Total rope mass m
    Used in Eq. 3 to compute per-node masses; not reported.
assumptions (4)
  • domain assumption Gripper forward kinematics gives the true grasp point on the rope.
    Eq. 1 takes the closest node to the gripper center as the grasped node, so any kinematic error or slipping directly enters the prediction.
  • domain assumption The PBD force model with gravity, normal, friction, and fixed segment length approximates the rope dynamics at the 1 Hz update rate.
    Eqs. 2-9 include only these effects; the conclusion admits self-intersections are not modeled, so the prediction is only an approximation.
  • domain assumption Reliable foreground segmentation of the rope is available in every camera view.
    Section III.B initializes Gaussian colors from segmentation and uses object masking in the update, so tracking depends on clean masks.
  • domain assumption Spherical Gaussians with covariance σI are a sufficient visual model of the rope's appearance for the rendering loss.
    Section III.B uses equal-sized spheres and ignores rotation, which simplifies optimization but limits how faithfully the rope is rendered.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting." pith.science (2026). https://pith.science/paper/KPY7WGVQ

@misc{pith2026250508644,
  author       = {Pith},
  title        = {Pith review of: DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KPY7WGVQ}},
  note         = {Machine review of arXiv:2505.08644}
}
read the original abstract

This work presents DLO-Splatting, an algorithm for estimating the 3D shape of Deformable Linear Objects (DLOs) from multi-view RGB images and gripper state information through prediction-update filtering. The DLO-Splatting algorithm uses a position-based dynamics model with shape smoothness and rigidity dampening corrections to predict the object shape. Optimization with a 3D Gaussian Splatting-based rendering loss iteratively renders and refines the prediction to align it with the visual observations in the update step. Initial experiments demonstrate promising results in a knot tying scenario, which is challenging for existing vision-only methods.

Figures

Figures reproduced from arXiv: 2505.08644 by the authors.

Figure 1
Figure 1. The DLO-Splatting Algorithm. The DLO-Splatting algorithm estimates the 3D state of a DLO using a prediction-update framework akin to Bayesian filtering. The DLO state is predicted using position-based dynamics and is iteratively updated using 3D Gaussian Splatting-based rendering. B. Update with 3D Gaussian Splatting-Based Rendering In 3D Gaussian Splatting, the scene is represented as a set of M 3D Gaussian distrib… view at source ↗
Figure 2
Figure 2. Qualitative Results. The DLO-Splatting algorithm is compared to TrackDLO on qualitative tracking during a cross move commonly performed in knot tying. During this move, one tip of the DLO is moved through a loop to create an additional crossing in the topology. After eight seconds, both DLO-Splatting and TrackDLO failed to estimate the correct topology of the DLO, however DLO-Splatting succeeds in tracking the grasp… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 58 canonical work pages

  1. [1]

    Learning Topological Motion Primitives for Knot Planning,

    M. Yan, G. Li, Y . Zhu, and J. Bohg, “Learning Topological Motion Primitives for Knot Planning,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2020. 1

  2. [2]

    Automatic Shape Control of Deformable Wires Based on Model-Free Visual Servoing,

    R. Lagneau, A. Krupa, and M. Marchal, “Automatic Shape Control of Deformable Wires Based on Model-Free Visual Servoing,” IEEE Robot. Autom. Lett. , vol. 5, no. 4, pp. 5252– 5259, 2020

  3. [3]

    Modeling, Learning, Per- ception, and Control Methods for Deformable Object Manipu- lation,

    H. Yin, A. Varava, and D. Kragic, “Modeling, Learning, Per- ception, and Control Methods for Deformable Object Manipu- lation,” in Sci. Robot. , vol. 6, May 2021, pp. 1–16. 1

  4. [4]

    Shape Control of Deformable Linear Objects with Offline and Online Learning of Local Linear Deformation Models,

    M. Yu, H. Zhong, and X. Li, “Shape Control of Deformable Linear Objects with Offline and Online Learning of Local Linear Deformation Models,” in IEEE Int. Conf. Robot. Autom. (ICRA), 2022, pp. 1337–1343

  5. [5]

    Robotic Cable Routing with Spatial Representation,

    S. Jin, W. Lian, C. Wang, M. Tomizuka, and S. Schaal, “Robotic Cable Routing with Spatial Representation,” IEEE Robot. Autom. Lett. , vol. 7, no. 2, pp. 5687–5694, 2022. 1

  6. [6]

    Dynamic Trajectory Planning for Robotic Knot Tying,

    B. Lu, H. K. Chu, and L. Cheng, “Dynamic Trajectory Planning for Robotic Knot Tying,” in IEEE Int. Conf. Real-Time Comput. Robot. (RCAR) , 2016, pp. 180–185. 1

  7. [7]

    Disentangling Dense Multi-Cable Knots,

    V . Viswanath, J. Grannen, P. Sundaresan, B. Thananjeyan, A. Balakrishna, E. Novoseller, J. Ichnowski, M. Laskey, J. E. Gonzalez, and K. Goldberg, “Disentangling Dense Multi-Cable Knots,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2021, pp. 3731–3738

  8. [8]

    Efficient Spatial Representation and Routing of Deformable One-Dimensional Objects for Manipulation,

    A. Keipour, M. Bandari, and S. Schaal, “Efficient Spatial Representation and Routing of Deformable One-Dimensional Objects for Manipulation,” IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS), pp. 211–216, 2022

Show all 59 references
  1. [9]

    An Algorithm Based on Bidirectional Searching and Geometric Constrained Sampling for Automatic Manipulation Planning in Aircraft Cable Assembly,

    J. Guo, J. Zhang, D. Wu, Y . Gai, and K. Chen, “An Algorithm Based on Bidirectional Searching and Geometric Constrained Sampling for Automatic Manipulation Planning in Aircraft Cable Assembly,” J. Manuf. Syst. , vol. 57, pp. 158–168, 2020

  2. [10]

    Model-Based Manipulation of Linear Flexible Objects: Task Automation in Simulation and Real World,

    P. Chang and T. Padır, “Model-Based Manipulation of Linear Flexible Objects: Task Automation in Simulation and Real World,” Machines, vol. 8, 2020. 1

  3. [11]

    DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation,

    B. P. Duisterhof, Z. Mandi, Y . Yao, J.-W. Liu, J. Seidenschwarz, M. Z. Shou, R. Deva, S. Song, S. Birchfield, B. Wen, and J. Ichnowski, “DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation,” IEEE Int. Work- shop Algorithmic F ound. Robot. (WAFR...

  4. [12]

    Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision,

    A. Longhini, M. B ¨usching, B. P. Duisterhof, J. Lundell, J. Ich- nowski, M. Bj ¨orkman, and D. Kragic, “Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision,” in Conf. Robot. Learn. (CoRL) , 2024. 1

  5. [13]

    TrackDLO: Tracking Deformable Linear Objects Under Occlusion With Motion Coherence,

    J. Xiang, H. Dinkel, H. Zhao, N. Gao, B. Coltin, T. Smith, and T. Bretl, “TrackDLO: Tracking Deformable Linear Objects Under Occlusion With Motion Coherence,” IEEE Robot. Autom. Lett., vol. 8, no. 10, pp. 6179–6186, 2023. 1, 3

  6. [14]

    Occlusion-Robust Deformable Object Tracking Without Physics Simulation,

    C. Chi and D. Berenson, “Occlusion-Robust Deformable Object Tracking Without Physics Simulation,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2019, pp. 6443–6450

  7. [15]

    Tracking Partially- Occluded Deformable Objects While Enforcing Geometric Con- straints,

    Y . Wang, D. McConachie, and D. Berenson, “Tracking Partially- Occluded Deformable Objects While Enforcing Geometric Con- straints,” in IEEE Int. Conf. Robot. Autom. (ICRA) , 2021, pp. 14 199–14 205

  8. [16]

    A Framework for Manipulating Deformable Linear Objects by Coherent Point Drift,

    T. Tang, C. Wang, and M. Tomizuka, “A Framework for Manipulating Deformable Linear Objects by Coherent Point Drift,” IEEE Robot. Autom. Lett. , vol. 3, no. 4, pp. 3426–3433, 2018

  9. [17]

    Track Deformable Objects from Point Clouds with Structure Preserved Registration,

    T. Tang and M. Tomizuka, “Track Deformable Objects from Point Clouds with Structure Preserved Registration,” Int. J. Robot. Res. , vol. 41, no. 6, pp. 599–614, 2022

  10. [18]

    Learning to Propagate Interaction Effects for Modeling Deformable Linear Objects Dynamics,

    Y . Yang, J. A. Stork, and T. Stoyanov, “Learning to Propagate Interaction Effects for Modeling Deformable Linear Objects Dynamics,” IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , pp. 4056–4062, 2021

  11. [19]

    Deformable Linear Object Prediction Using Locally Linear Latent Dynamics,

    W. Zhang, K. Schmeckpeper, P. Chaudhari, and K. Daniilidis, “Deformable Linear Object Prediction Using Locally Linear Latent Dynamics,” in IEEE Int. Conf. Robot. Autom. (ICRA) , June 2021, pp. 13 503–13 509

  12. [20]

    Tracking De- formable Objects with Point Clouds,

    J. Schulman, A. Lee, J. Ho, and P. Abbeel, “Tracking De- formable Objects with Point Clouds,” in IEEE Int. Conf. Robot. Autom. (ICRA) , 2013, pp. 1130–1137. 1

  13. [21]

    Non-Rigid Point Set Registration with Global-Local Topology Preservation,

    S. Ge, G. Fan, and M. Ding, “Non-Rigid Point Set Registration with Global-Local Topology Preservation,”IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW) , pp. 245– 251, 2014. 1

  14. [22]

    Simultaneous Shape Tracking of Mul- tiple Deformable Linear Objects with Global-Local Topology Preservation,

    J. Xiang and H. Dinkel, “Simultaneous Shape Tracking of Mul- tiple Deformable Linear Objects with Global-Local Topology Preservation,” in IEEE Int. Conf. Robot. Autom. (ICRA) Work- shop on Representing and Manipulating Deformable Objects , May 2023. 1

  15. [23]

    Dense- PhysNet: Learning Dense Physical Object Representations via Multi-Step Dynamic Interactions,

    Z. Xu, J. Wu, A. Zeng, J. B. Tenenbaum, and S. Song, “Dense- PhysNet: Learning Dense Physical Object Representations via Multi-Step Dynamic Interactions,” in Robot. Sci. Syst. (RSS) ,

  16. [24]

    Simultaneous Learning of Contact and Continuous Dynamics,

    B. Bianchini, M. Halm, and M. Posa, “Simultaneous Learning of Contact and Continuous Dynamics,” in Conf. Robot. Learn. (CoRL), 2023. 1

  17. [25]

    Identification of Spring Parameters for Deformable Object Simulation,

    B. Lloyd, G. Sz ´ekely, and M. Harders, “Identification of Spring Parameters for Deformable Object Simulation,” IEEE Trans. Vis. Comput. Graphics , vol. 13, no. 5, pp. 1081–1094, 2007. 1

  18. [26]

    Uni- fied Particle Physics for Real-Time Applications,

    M. Macklin, M. M ¨uller, N. Chentanez, and T.-Y . Kim, “Uni- fied Particle Physics for Real-Time Applications,” ACM Trans. Graph. (TOG), vol. 33, no. 4, Jul. 2014

  19. [27]

    Position-Based Simu- lation Methods in Computer Graphics

    J. Bender, M. M ¨uller, and M. Macklin, “Position-Based Simu- lation Methods in Computer Graphics.” in Eurographics (Tuto- rials), 2015, pp. 1–32. 1

  20. [28]

    State Estimation for Deformable Objects by Point Registration and Dynamic Simulation,

    T. Tang, Y . Fan, H.-C. Lin, and M. Tomizuka, “State Estimation for Deformable Objects by Point Registration and Dynamic Simulation,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2017, pp. 2427–2433. 1

  21. [29]

    Physically Embodied Gaussian Splatting: A Visually Learnt and Physically Grounded 3D Representation for Robotics,

    J. Abou-Chakra, K. Rana, F. Dayoub, and N. Suenderhauf, “Physically Embodied Gaussian Splatting: A Visually Learnt and Physically Grounded 3D Representation for Robotics,” in Conf. Robot. Learn. (CoRL) , 2024. 3

  22. [30]

    Dynamic 3D Gaussian Track- ing for Graph-Based Neural Dynamics Modeling,

    M. Zhang, K. Zhang, and Y . Li, “Dynamic 3D Gaussian Track- ing for Graph-Based Neural Dynamics Modeling,” in Conf. Robot. Learn. (CoRL) , 2024. 1

  23. [31]

    CoTracker: It is Better to Track Together,

    N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “CoTracker: It is Better to Track Together,” Eur . Conf. Comput. Vis. (ECCV) , 2024. 1

  24. [32]

    VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation,

    X. Shi, Z. Huang, W. Bian, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li, “VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2023, pp. 12 435– 12 446

  25. [33]

    TAP-Vid: A Benchmark for Tracking Any Point in a Video,

    C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y . Aytar, J. Carreira, A. Zisserman, and Y . Yang, “TAP-Vid: A Benchmark for Tracking Any Point in a Video,” in Adv. Neur . Inf. Proc. (NeurIPS) , 2023

  26. [34]

    Tracking Everything Everywhere All at Once,

    Q. Wang, Y .-Y . Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely, “Tracking Everything Everywhere All at Once,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2023, pp. 19 738– 19 749

  27. [35]

    RAFT: Recurrent All-Pairs Field Trans- forms for Optical Flow,

    Z. Teed and J. Deng, “RAFT: Recurrent All-Pairs Field Trans- forms for Optical Flow,” in Eur . Conf. Comput. Vis. (ECCV) , 2020, pp. 402–419. 1

  28. [36]

    Deformable Linear Objects 3D Shape Estimation and Tracking From Multiple 2D Views,

    A. Caporali, K. Galassi, and G. Palli, “Deformable Linear Objects 3D Shape Estimation and Tracking From Multiple 2D Views,” IEEE Robot. Autom. Lett. , vol. 8, no. 6, pp. 3852–3859,

  29. [37]

    NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” Commun. ACM , vol. 65, no. 1, p. 99–106, 2022. 1

  30. [38]

    D-NeRF: Neural Radiance Fields for Dynamic Scenes,

    A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer, “D-NeRF: Neural Radiance Fields for Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2021, pp. 10 318–10 327

  31. [39]

    Nerfies: Deformable Neural Radiance Fields,

    K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable Neural Radiance Fields,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2021, pp. 5845–5854

  32. [40]

    Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,

    Z. Li, S. Niklaus, N. Snavely, and O. Wang, “Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2021, pp. 6494–6504

  33. [41]

    DynIBaR: Neural Dynamic Image-Based Rendering,

    Z. Li, Q. Wang, F. Cole, R. Tucker, and N. Snavely, “DynIBaR: Neural Dynamic Image-Based Rendering,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2023, pp. 4273–

  34. [42]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering,” ACM Trans. Graph. (TOG) , vol. 42, no. 4, July 2023. 1, 2

  35. [43]

    Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis,

    J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis,” in IEEE Int. Conf. 3D Vis. (3DV) , 2024, pp. 800–809

  36. [44]

    4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 20 310–20 320. 1

  37. [45]

    Monocular Dynamic View Synthesis: A Reality Check,

    H. Gao, R. Li, S. Tulsiani, B. Russell, and A. Kanazawa, “Monocular Dynamic View Synthesis: A Reality Check,” in Adv. Neur . Inf. Proc. (NeurIPS) , 2022, pp. 33 768–33 780. 1

  38. [46]

    FlowIBR: Leveraging Pre-Training for Efficient Neural Image- Based Rendering of Dynamic Scenes,

    M. B ¨usching, J. Bengtson, D. Nilsson, and M. Bj ¨orkman, “FlowIBR: Leveraging Pre-Training for Efficient Neural Image- Based Rendering of Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW) , June 2024, pp. 8016–8026. 1

  39. [47]

    gsplat: An Open-Source Library for Gaussian Splatting,

    V . Ye, R. Li, J. Kerr, M. Turkulainen, B. Yi, Z. Pan, O. Seiskari, J. Ye, J. Hu, M. Tancik, and A. Kanazawa, “gsplat: An Open-Source Library for Gaussian Splatting,” arXiv preprint arXiv:2409.06765, 2024. 3

  40. [48]

    KnotDLO: Toward Interpretable Knot Tying,

    H. Dinkel, R. Navaratna, J. Xiang, B. Coltin, T. Smith, and T. Bretl, “KnotDLO: Toward Interpretable Knot Tying,” in IEEE Int. Conf. on Robot. and Autom. (ICRA) 3D Visual Representations for Manipulation Workshop , May 2024. 3

  41. [49]

    TieBot: Learning to Knot a Tie from Visual Demonstration through a Real-to-Sim-to-Real Approach,

    W. Peng, J. Lv, Y . Zeng, H. Chen, S. Zhao, J. Sun, C. Lu, and L. Shao, “TieBot: Learning to Knot a Tie from Visual Demonstration through a Real-to-Sim-to-Real Approach,” in Conf. Robot. Learn. (CoRL) , 2024. 3

  42. [50]

    General Hand–Eye Calibration Based on Reprojection Error Minimization,

    K. Koide and E. Menegatti, “General Hand–Eye Calibration Based on Reprojection Error Minimization,” IEEE Robot. Au- tom. Lett. , vol. 4, no. 2, pp. 1021–1028, 2019. 3

  43. [51]

    Robotic Operating System: Noetic Ninjemys,

    Stanford Artificial Intelligence Laboratory, “Robotic Operating System: Noetic Ninjemys,” https://www.ros.org, 2018. 3

  44. [52]

    Warp: A High-performance Python Framework for GPU Simulation and Graphics,

    M. Macklin, “Warp: A High-performance Python Framework for GPU Simulation and Graphics,” March 2022, nVIDIA GPU Technology Conference (GTC). 3

  45. [53]

    A Robust Deformable Linear Object Perception Pipeline in 3D: From Segmentation to Reconstruction,

    S. Zhaole, H. Zhou, L. Nanbo, L. Chen, J. Zhu, and R. B. Fisher, “A Robust Deformable Linear Object Perception Pipeline in 3D: From Segmentation to Reconstruction,” IEEE Robot. Autom. Lett., vol. 9, no. 1, pp. 843–850, 2024. 3

  46. [54]

    Wire Point Cloud Instance Segmentation from RGBD Imagery with Mask R-CNN,

    H. Dinkel, J. Xiang, H. Zhao, B. Coltin, T. Smith, and T. Bretl, “Wire Point Cloud Instance Segmentation from RGBD Imagery with Mask R-CNN,” in IEEE Int. Conf. Robot. Autom. (ICRA) Workshop on Representing and Manipulating Deformable Ob- jects, May 2022

  47. [55]

    Auto-Generated Wires Dataset for Semantic Segmen- tation with Domain Independence,

    R. Zanella, A. Caporali, K. Tadaka, D. De Gregorio, and G. Palli, “Auto-Generated Wires Dataset for Semantic Segmen- tation with Domain Independence,” in IEEE Int. Conf. Comput. Cont. Robot. (ICCCR) . IEEE, Jan. 2021, pp. 292–298

  48. [56]

    Ari- adne+: Deep Learning-Based Augmented Framework for the Instance Segmentation of Wires,

    A. Caporali, R. Zanella, D. De Gregorio, and G. Palli, “Ari- adne+: Deep Learning-Based Augmented Framework for the Instance Segmentation of Wires,” in IEEE Trans. Ind. Inf. , February 2022, pp. 1–11

  49. [57]

    FASTDLO: Fast Deformable Linear Objects Instance Segmentation,

    A. Caporali, K. Galassi, R. Zanella, and G. Palli, “FASTDLO: Fast Deformable Linear Objects Instance Segmentation,” IEEE Robot. Autom. Lett. , vol. 7, no. 4, pp. 9075–9082, 2022

  50. [58]

    RT-DLO: Real-Time Deformable Linear Objects Instance Segmentation,

    A. Caporali, K. Galassi, B. L. ˇZagar, R. Zanella, G. Palli, and A. C. Knoll, “RT-DLO: Real-Time Deformable Linear Objects Instance Segmentation,” IEEE Trans. Ind. Informat. , vol. 19, no. 11, pp. 11 333–11 342, 2023

  51. [59]

    HAND- LOOM: Learned Tracing of One-Dimensional Objects for In- spection and Manipulation,

    V . Viswanath, K. Shivakumar, M. Parulekar, J. Ajmera, J. Kerr, J. Ichnowski, R. Cheng, T. Kollar, and K. Goldberg, “HAND- LOOM: Learned Tracing of One-Dimensional Objects for In- spection and Manipulation,” in Conf. Robot. Learn. (CoRL) , 2023, pp. 341–357. 3

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.