REVIEW 3 major objections 5 minor 59 references
DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read DLO-Splatting tracks a rope's 3D shape by combining physics prediction with Gaussian-splatting image updates.
desk verdict Novel PBD-plus-3DGS combination for DLO tracking, but the central prediction equation is dimensionally off, and the single demo doesn't back the claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a Bayesian-style prediction-update filter whose state is a node chain $X_t \in \mathbb{R}^{N\times 3}$. The prediction step uses position-based dynamics: Verlet integration with gravity, a planar contact normal, friction, and a length-constraint projection that keeps every segment at length $L$. The update step places $d$ spherical 3D Gaussians per segment along the centerline, renders them into each camera with $\alpha$-blending, and minimizes the squared image difference $L_{\text{obs}}$ by stochastic gradient descent on the node positions. The Gaussians' covariance is tied to the rope diameter and their positions to the nodes, so the visual loss directly drives geometric correction of the predicted rope.
What would settle it
Record a full knot-tying sequence in which one segment of the rope is marked with a distinctive color, so the over/under order at each crossing is visible in the images. If, at the moment a crossing is formed, the tracked node chain places the marked segment on the wrong side while the rendered rope still matches all three camera views, then the claim that Gaussian-splatting rendering disambiguates topology is falsified.
Extended reading notes
Core claim
The paper's central claim is that combining a first-principles physics prediction with a differentiable Gaussian-splatting renderer produces a state estimate for a deformable linear object that is robust to visual ambiguities that defeat vision-only trackers, specifically dense crossings and self-occlusion during knot tying. The algorithm represents the rope as $N$ nodes, predicts the next node configuration $\hat{X}_{t+1}$ with position-based dynamics, then refines it by minimizing the observation loss $L_{\text{obs}} = \|\mathbf{I}^t_k - \tilde{\mathbf{I}}^t_k\|^2$ against all $K$ camera views, where the rendered image $\tilde{\mathbf{I}}^t_k$ is computed by $\alpha$-blending spherical Gaussians placed along the rope centerline. Because the Gaussians are tied to the nodes, the gradient of the rendering loss updates the 3D geometry directly. The authors claim this lets the filter track visually complex topologies that cannot be disambiguated by vision alone, and they demonstrate the advantage on the grasped rope tip during a knot-tying cross move.
Load-bearing premise
The load-bearing premise is that a simple physics prediction—gravity, table contact, friction, and fixed segment lengths—moves the rope estimate close enough to the true state that the image-based correction can converge, even while the rope crosses over itself during knot tying.
Editorial extensions
If this is right
- A robot could estimate its rope's 3D state during manipulation using only RGB cameras and its own gripper pose, with no learned dynamics model and no markers.
- The same prediction-update loop should transfer to new rope materials or environments without retraining, provided mass, friction, and contact parameters are known.
- If the rendering loss truly disambiguates crossing topology, multi-view RGB alone can maintain correct over/under structure where 2D or single-view trackers fail.
- The demonstrated 1 Hz update rate and lack of self-intersection modeling currently restrict the method to slow motions and sparse self-contact; faster rendering and explicit self-contact prediction would be needed for full-speed knot tying.
- Because the update uses object masking, the method can in principle be extended to simultaneous tracking of several DLOs once multi-instance segmentation is added.
Reading between the lines
- A direct ablation varying the number of cameras (for example, two vs three) would test whether the reported tip-tracking advantage comes from multi-view disambiguation or from the physics prediction alone; the paper does not report this.
- The fixed segment-length position-based dynamics model behaves like a discrete inextensible rod without bending stiffness; comparing against a rod model with bending and twisting energy would isolate whether the final topology collapse is a prediction error or an update error.
- The 1 Hz bottleneck is likely shared between the physics time step and the stochastic gradient descent rendering update; a test that runs prediction at 100 Hz while keeping update at 1 Hz would show which side limits the filter.
- The claim that vision alone cannot disambiguate these topologies would be sharpened by a controlled experiment with two visually identical crossings, where the only distinguishing information is the camera viewpoint; the paper's single cross-move demonstration is suggestive but not exhaustive.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DLO-Splatting, a prediction-update filter for estimating the 3D state of a deformable linear object from multi-view RGB images and gripper pose. The prediction step uses position-based dynamics with gravity, planar contact, friction, and fixed segment-length constraints; the update step renders the predicted node chain with 3D Gaussian splatting and optimizes a rendering loss with SGD. The method is evaluated in one qualitative knot-tying 'cross move' with three cameras and compared against TrackDLO. The authors report that DLO-Splatting tracks the grasped rope tip more accurately than TrackDLO, but that both methods fail to recover the rope's final shape and topology.
Significance. If fully validated, the contribution would be useful: DLO-Splatting combines a training-free physics prediction with a differentiable rendering update, which is a reasonable way to address occlusion and multi-view disambiguation for DLO tracking. The paper also clearly enumerates its limitations, including unmodeled self-intersections, a 1 Hz update rate, and sensitivity to occlusion. However, the central claim that the method tracks visually complex topologies through dense knotting is not supported by the current evidence: Eq. (2) is not reproducible as written, the evaluation is a single qualitative trial without metrics, and the paper itself admits both methods fail on the final topology. The idea merits further work, but the manuscript as submitted is not ready for publication.
major comments (3)
- [Section III-A, Eq. (2)] Equation (2) is not Verlet integration and is dimensionally inconsistent as written: X^{t+1} = X^t + (X^t - X^{t-1})/Δt + (1/2)F^t Δt^2 adds a velocity-like term to positions, and the force term has units of mass·length unless F^t is already divided by node mass, which Eq. (3) contradicts because the forces include m_i g. The standard Verlet update for a node of mass m_i is x^{t+1} = 2x^t - x^{t-1} + (F^t/m_i)Δt^2, or equivalently x^{t+1} = x^t + v^t Δt + (1/2)(F^t/m_i)Δt^2. This is load-bearing because the PBD output is the initial condition for the rendering update in Eqs. (14)-(15), so any change in the integration formula changes the entire filter. The one qualitative success reported in Section IV (tracking the grasped tip) is controlled by the directly imposed gripper position in Eq. (1), so it does not validate the physics integration. The authors need to provide a corrected, dimensionally consistent integration step and state the mass normalization explicitly.
- [Section IV and Figure 2] The evaluation does not support the abstract's and Introduction's claim that DLO-Splatting tracks visually complex topologies that cannot be disambiguated by vision-based tracking alone. The demonstration is a single cross move with three cameras, presented only qualitatively: there are no quantitative metrics such as node-to-ground-truth distance or topology error, no error bars, no ablations, and no sensitivity analysis. The paper's own conclusion states that both DLO-Splatting and TrackDLO fail to recover the rope's final shape and topology, and Section V lists unmodeled self-intersections, a 1 Hz update rate, and occlusion from the gripper as limitations. The only reported advantage, tracking the grasped tip during one move, is too narrow to establish the central contribution, especially because the grasped tip is directly constrained by the gripper pose.
- [Section III-A and Section IV] The algorithm is not reproducible from the text because the values of the free parameters used in the demonstration are not reported: friction coefficient μ_f, integration time step Δt, Gaussian diameter σ, node count N and segment length L, Gaussians per segment d, and total rope mass m. The optimizer hyperparameters for the rendering update (learning rate, number of SGD iterations, initialization) are also missing. In addition, the relationship between the constraint projection in Eqs. (4)-(6), which is written as an update to X^t, and the length correction in Eqs. (8)-(9), which is applied to X^{t+1}, is unclear, so a reader cannot reimplement the exact algorithm. Without these details, the single demonstration cannot be checked or extended.
minor comments (5)
- [Section III-B, Eq. (11)] The rendering function h is written as a function of X^t_PBD and the camera projection P_k, but Eqs. (12)-(13) also depend on per-Gaussian colors c_j and opacities o_j; the text should specify how these are initialized and whether they are optimized or fixed during the update.
- [Section III-A, Eq. (3)] The velocity v^t_{i,xy} in the friction term is not defined; the text should state whether it is computed from the current and previous node positions and how the planar component is extracted.
- [Section III-A, Eq. (8)] Equation (8) is ambiguous without parentheses; it should read Δl^{t+1}_i = (L - l^{t+1}_i) · (x^{t+1}_{i+1} - x^{t+1}_i)/l^{t+1}_i to clarify that the scalar length difference multiplies the unit direction vector.
- [Section III-A, Eq. (2)] The gripper action a_t is introduced in the text but does not appear explicitly in Eq. (2); the authors should clarify how the action influences the free nodes as opposed to the grasped node.
- [Section IV] The paper should report the image resolution used for the rendering loss, the number of rendering iterations per update, and the actual per-step runtime; Section V mentions a 1 Hz update rate, but the evaluation section does not provide these details.
Circularity Check
Positive demo result (grasped-tip tracking) is forced by Eq. 1 gripper-input construction; core filter derivation is not circular.
-
fitted input called prediction
[Section III-A Eq. (1) and Section IV / Fig. 2 caption]
"The position of the grasped node, xt g, is computed as the closest node in the set of nodes describing the shape of the rope, Xt, to the position of the center of the gripper, pt g ... When the grasped node moves to the new gripper position, the position correction xt+1 g = pt g + atΔt must also be propagated through the rope. ... DLO-Splatting succeeds in tracking the grasped tip."
Equation (1) defines the grasped node as the rope node nearest the gripper center, and the algorithm directly imposes its next position from the gripper pose. The demonstration's only reported success is that DLO-Splatting 'succeeds in tracking the grasped tip' (Fig. 2). Because the tip trajectory is prescribed by the gripper state input, this success is guaranteed by construction: the output 'tracked tip' is the input gripper pose, not an inference from the rendering loss or the dynamics. It therefore cannot validate the claim that multi-view synthesis disambiguates topology, especially since the paper states both methods 'failed to estimate the correct topology.'
full rationale
The core estimator is a genuine predict-update filter: PBD predicts a candidate state, and the rendering loss Lobs = ||I_k^t - \tilde I_k^t||^2 against external camera images is minimized over a state correction ΔX (Eqs. 14-15). This uses observations that are not constructed from the target, so the derivation of the filter is not circular. The only circular element is the demonstration's positive result. The grasped tip is defined as the closest node to the gripper (Eq. 1) and its motion is directly imposed from gripper pose, so reporting that DLO-Splatting 'tracks the grasped tip' is an input-output identity, not an independent tracking success. The paper's own conclusion narrows the achievement: self-intersections are not modeled, update runs at 1 Hz, and both methods fail to recover the rope's final shape and topology, so the headline topology claims are unsupported but that is a scope/correctness issue. The TrackDLO initialization is a self-citation, but it is used symmetrically as a starting point for both methods and does not enter the derivation. Eq. (2) appears dimensionally inconsistent with Verlet integration; that is a reproducibility/correctness defect, not circularity. Overall, partial circularity exists only in the evaluation's 'grasped tip' claim, not in the state-estimation derivation.
Assumptions & free parameters
free parameters (6)
- Friction coefficient μf
- Integration time step Δt
- Gaussian diameter σ
- Node count N and segment length L
- Gaussians per segment d
- Total rope mass m
assumptions (4)
- domain assumption Gripper forward kinematics gives the true grasp point on the rope.
- domain assumption The PBD force model with gravity, normal, friction, and fixed segment length approximates the rope dynamics at the 1 Hz update rate.
- domain assumption Reliable foreground segmentation of the rope is available in every camera view.
- domain assumption Spherical Gaussians with covariance σI are a sufficient visual model of the rope's appearance for the rendering loss.
Cite this review
Pith. "Pith review of DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting." pith.science (2026). https://pith.science/paper/KPY7WGVQ
@misc{pith2026250508644,
author = {Pith},
title = {Pith review of: DLO-Splatting: Tracking Deformable Linear Objects Using 3D Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/KPY7WGVQ}},
note = {Machine review of arXiv:2505.08644}
}
read the original abstract
This work presents DLO-Splatting, an algorithm for estimating the 3D shape of Deformable Linear Objects (DLOs) from multi-view RGB images and gripper state information through prediction-update filtering. The DLO-Splatting algorithm uses a position-based dynamics model with shape smoothness and rigidity dampening corrections to predict the object shape. Optimization with a 3D Gaussian Splatting-based rendering loss iteratively renders and refines the prediction to align it with the visual observations in the update step. Initial experiments demonstrate promising results in a knot tying scenario, which is challenging for existing vision-only methods.
Figures
Reference graph
Works this paper leans on
-
[1]
Learning Topological Motion Primitives for Knot Planning,
M. Yan, G. Li, Y . Zhu, and J. Bohg, “Learning Topological Motion Primitives for Knot Planning,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2020. 1
work page 2020
-
[2]
Automatic Shape Control of Deformable Wires Based on Model-Free Visual Servoing,
R. Lagneau, A. Krupa, and M. Marchal, “Automatic Shape Control of Deformable Wires Based on Model-Free Visual Servoing,” IEEE Robot. Autom. Lett. , vol. 5, no. 4, pp. 5252– 5259, 2020
work page 2020
-
[3]
Modeling, Learning, Per- ception, and Control Methods for Deformable Object Manipu- lation,
H. Yin, A. Varava, and D. Kragic, “Modeling, Learning, Per- ception, and Control Methods for Deformable Object Manipu- lation,” in Sci. Robot. , vol. 6, May 2021, pp. 1–16. 1
work page 2021
-
[4]
M. Yu, H. Zhong, and X. Li, “Shape Control of Deformable Linear Objects with Offline and Online Learning of Local Linear Deformation Models,” in IEEE Int. Conf. Robot. Autom. (ICRA), 2022, pp. 1337–1343
work page 2022
-
[5]
Robotic Cable Routing with Spatial Representation,
S. Jin, W. Lian, C. Wang, M. Tomizuka, and S. Schaal, “Robotic Cable Routing with Spatial Representation,” IEEE Robot. Autom. Lett. , vol. 7, no. 2, pp. 5687–5694, 2022. 1
work page 2022
-
[6]
Dynamic Trajectory Planning for Robotic Knot Tying,
B. Lu, H. K. Chu, and L. Cheng, “Dynamic Trajectory Planning for Robotic Knot Tying,” in IEEE Int. Conf. Real-Time Comput. Robot. (RCAR) , 2016, pp. 180–185. 1
work page 2016
-
[7]
Disentangling Dense Multi-Cable Knots,
V . Viswanath, J. Grannen, P. Sundaresan, B. Thananjeyan, A. Balakrishna, E. Novoseller, J. Ichnowski, M. Laskey, J. E. Gonzalez, and K. Goldberg, “Disentangling Dense Multi-Cable Knots,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2021, pp. 3731–3738
work page 2021
-
[8]
Efficient Spatial Representation and Routing of Deformable One-Dimensional Objects for Manipulation,
A. Keipour, M. Bandari, and S. Schaal, “Efficient Spatial Representation and Routing of Deformable One-Dimensional Objects for Manipulation,” IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS), pp. 211–216, 2022
work page 2022
Show all 59 references
-
[9]
An Algorithm Based on Bidirectional Searching and Geometric Constrained Sampling for Automatic Manipulation Planning in Aircraft Cable Assembly,
J. Guo, J. Zhang, D. Wu, Y . Gai, and K. Chen, “An Algorithm Based on Bidirectional Searching and Geometric Constrained Sampling for Automatic Manipulation Planning in Aircraft Cable Assembly,” J. Manuf. Syst. , vol. 57, pp. 158–168, 2020
2020
-
[10]
Model-Based Manipulation of Linear Flexible Objects: Task Automation in Simulation and Real World,
P. Chang and T. Padır, “Model-Based Manipulation of Linear Flexible Objects: Task Automation in Simulation and Real World,” Machines, vol. 8, 2020. 1
2020
-
[11]
DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation,
B. P. Duisterhof, Z. Mandi, Y . Yao, J.-W. Liu, J. Seidenschwarz, M. Z. Shou, R. Deva, S. Song, S. Birchfield, B. Wen, and J. Ichnowski, “DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation,” IEEE Int. Work- shop Algorithmic F ound. Robot. (WAFR...
2024
-
[12]
Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision,
A. Longhini, M. B ¨usching, B. P. Duisterhof, J. Lundell, J. Ich- nowski, M. Bj ¨orkman, and D. Kragic, “Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision,” in Conf. Robot. Learn. (CoRL) , 2024. 1
2024
-
[13]
TrackDLO: Tracking Deformable Linear Objects Under Occlusion With Motion Coherence,
J. Xiang, H. Dinkel, H. Zhao, N. Gao, B. Coltin, T. Smith, and T. Bretl, “TrackDLO: Tracking Deformable Linear Objects Under Occlusion With Motion Coherence,” IEEE Robot. Autom. Lett., vol. 8, no. 10, pp. 6179–6186, 2023. 1, 3
2023
-
[14]
Occlusion-Robust Deformable Object Tracking Without Physics Simulation,
C. Chi and D. Berenson, “Occlusion-Robust Deformable Object Tracking Without Physics Simulation,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2019, pp. 6443–6450
2019
-
[15]
Tracking Partially- Occluded Deformable Objects While Enforcing Geometric Con- straints,
Y . Wang, D. McConachie, and D. Berenson, “Tracking Partially- Occluded Deformable Objects While Enforcing Geometric Con- straints,” in IEEE Int. Conf. Robot. Autom. (ICRA) , 2021, pp. 14 199–14 205
2021
-
[16]
A Framework for Manipulating Deformable Linear Objects by Coherent Point Drift,
T. Tang, C. Wang, and M. Tomizuka, “A Framework for Manipulating Deformable Linear Objects by Coherent Point Drift,” IEEE Robot. Autom. Lett. , vol. 3, no. 4, pp. 3426–3433, 2018
2018
-
[17]
Track Deformable Objects from Point Clouds with Structure Preserved Registration,
T. Tang and M. Tomizuka, “Track Deformable Objects from Point Clouds with Structure Preserved Registration,” Int. J. Robot. Res. , vol. 41, no. 6, pp. 599–614, 2022
2022
-
[18]
Learning to Propagate Interaction Effects for Modeling Deformable Linear Objects Dynamics,
Y . Yang, J. A. Stork, and T. Stoyanov, “Learning to Propagate Interaction Effects for Modeling Deformable Linear Objects Dynamics,” IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , pp. 4056–4062, 2021
2021
-
[19]
Deformable Linear Object Prediction Using Locally Linear Latent Dynamics,
W. Zhang, K. Schmeckpeper, P. Chaudhari, and K. Daniilidis, “Deformable Linear Object Prediction Using Locally Linear Latent Dynamics,” in IEEE Int. Conf. Robot. Autom. (ICRA) , June 2021, pp. 13 503–13 509
2021
-
[20]
Tracking De- formable Objects with Point Clouds,
J. Schulman, A. Lee, J. Ho, and P. Abbeel, “Tracking De- formable Objects with Point Clouds,” in IEEE Int. Conf. Robot. Autom. (ICRA) , 2013, pp. 1130–1137. 1
2013
-
[21]
Non-Rigid Point Set Registration with Global-Local Topology Preservation,
S. Ge, G. Fan, and M. Ding, “Non-Rigid Point Set Registration with Global-Local Topology Preservation,”IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW) , pp. 245– 251, 2014. 1
2014
-
[22]
Simultaneous Shape Tracking of Mul- tiple Deformable Linear Objects with Global-Local Topology Preservation,
J. Xiang and H. Dinkel, “Simultaneous Shape Tracking of Mul- tiple Deformable Linear Objects with Global-Local Topology Preservation,” in IEEE Int. Conf. Robot. Autom. (ICRA) Work- shop on Representing and Manipulating Deformable Objects , May 2023. 1
2023
-
[23]
Dense- PhysNet: Learning Dense Physical Object Representations via Multi-Step Dynamic Interactions,
Z. Xu, J. Wu, A. Zeng, J. B. Tenenbaum, and S. Song, “Dense- PhysNet: Learning Dense Physical Object Representations via Multi-Step Dynamic Interactions,” in Robot. Sci. Syst. (RSS) ,
-
[24]
Simultaneous Learning of Contact and Continuous Dynamics,
B. Bianchini, M. Halm, and M. Posa, “Simultaneous Learning of Contact and Continuous Dynamics,” in Conf. Robot. Learn. (CoRL), 2023. 1
2023
-
[25]
Identification of Spring Parameters for Deformable Object Simulation,
B. Lloyd, G. Sz ´ekely, and M. Harders, “Identification of Spring Parameters for Deformable Object Simulation,” IEEE Trans. Vis. Comput. Graphics , vol. 13, no. 5, pp. 1081–1094, 2007. 1
2007
-
[26]
Uni- fied Particle Physics for Real-Time Applications,
M. Macklin, M. M ¨uller, N. Chentanez, and T.-Y . Kim, “Uni- fied Particle Physics for Real-Time Applications,” ACM Trans. Graph. (TOG), vol. 33, no. 4, Jul. 2014
2014
-
[27]
Position-Based Simu- lation Methods in Computer Graphics
J. Bender, M. M ¨uller, and M. Macklin, “Position-Based Simu- lation Methods in Computer Graphics.” in Eurographics (Tuto- rials), 2015, pp. 1–32. 1
2015
-
[28]
State Estimation for Deformable Objects by Point Registration and Dynamic Simulation,
T. Tang, Y . Fan, H.-C. Lin, and M. Tomizuka, “State Estimation for Deformable Objects by Point Registration and Dynamic Simulation,” in IEEE/RSJ Int. Conf. Intell. Robot. Sys. (IROS) , 2017, pp. 2427–2433. 1
2017
-
[29]
Physically Embodied Gaussian Splatting: A Visually Learnt and Physically Grounded 3D Representation for Robotics,
J. Abou-Chakra, K. Rana, F. Dayoub, and N. Suenderhauf, “Physically Embodied Gaussian Splatting: A Visually Learnt and Physically Grounded 3D Representation for Robotics,” in Conf. Robot. Learn. (CoRL) , 2024. 3
2024
-
[30]
Dynamic 3D Gaussian Track- ing for Graph-Based Neural Dynamics Modeling,
M. Zhang, K. Zhang, and Y . Li, “Dynamic 3D Gaussian Track- ing for Graph-Based Neural Dynamics Modeling,” in Conf. Robot. Learn. (CoRL) , 2024. 1
2024
-
[31]
CoTracker: It is Better to Track Together,
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “CoTracker: It is Better to Track Together,” Eur . Conf. Comput. Vis. (ECCV) , 2024. 1
2024
-
[32]
VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation,
X. Shi, Z. Huang, W. Bian, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li, “VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2023, pp. 12 435– 12 446
2023
-
[33]
TAP-Vid: A Benchmark for Tracking Any Point in a Video,
C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y . Aytar, J. Carreira, A. Zisserman, and Y . Yang, “TAP-Vid: A Benchmark for Tracking Any Point in a Video,” in Adv. Neur . Inf. Proc. (NeurIPS) , 2023
2023
-
[34]
Tracking Everything Everywhere All at Once,
Q. Wang, Y .-Y . Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely, “Tracking Everything Everywhere All at Once,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2023, pp. 19 738– 19 749
2023
-
[35]
RAFT: Recurrent All-Pairs Field Trans- forms for Optical Flow,
Z. Teed and J. Deng, “RAFT: Recurrent All-Pairs Field Trans- forms for Optical Flow,” in Eur . Conf. Comput. Vis. (ECCV) , 2020, pp. 402–419. 1
2020
-
[36]
Deformable Linear Objects 3D Shape Estimation and Tracking From Multiple 2D Views,
A. Caporali, K. Galassi, and G. Palli, “Deformable Linear Objects 3D Shape Estimation and Tracking From Multiple 2D Views,” IEEE Robot. Autom. Lett. , vol. 8, no. 6, pp. 3852–3859,
-
[37]
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” Commun. ACM , vol. 65, no. 1, p. 99–106, 2022. 1
2022
-
[38]
D-NeRF: Neural Radiance Fields for Dynamic Scenes,
A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer, “D-NeRF: Neural Radiance Fields for Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2021, pp. 10 318–10 327
2021
-
[39]
Nerfies: Deformable Neural Radiance Fields,
K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable Neural Radiance Fields,” in IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2021, pp. 5845–5854
2021
-
[40]
Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,
Z. Li, S. Niklaus, N. Snavely, and O. Wang, “Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2021, pp. 6494–6504
2021
-
[41]
DynIBaR: Neural Dynamic Image-Based Rendering,
Z. Li, Q. Wang, F. Cole, R. Tucker, and N. Snavely, “DynIBaR: Neural Dynamic Image-Based Rendering,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2023, pp. 4273–
2023
-
[42]
3D Gaussian Splatting for Real-Time Radiance Field Rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian Splatting for Real-Time Radiance Field Rendering,” ACM Trans. Graph. (TOG) , vol. 42, no. 4, July 2023. 1, 2
2023
-
[43]
Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis,
J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis,” in IEEE Int. Conf. 3D Vis. (3DV) , 2024, pp. 800–809
2024
-
[44]
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,
G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 20 310–20 320. 1
2024
-
[45]
Monocular Dynamic View Synthesis: A Reality Check,
H. Gao, R. Li, S. Tulsiani, B. Russell, and A. Kanazawa, “Monocular Dynamic View Synthesis: A Reality Check,” in Adv. Neur . Inf. Proc. (NeurIPS) , 2022, pp. 33 768–33 780. 1
2022
-
[46]
FlowIBR: Leveraging Pre-Training for Efficient Neural Image- Based Rendering of Dynamic Scenes,
M. B ¨usching, J. Bengtson, D. Nilsson, and M. Bj ¨orkman, “FlowIBR: Leveraging Pre-Training for Efficient Neural Image- Based Rendering of Dynamic Scenes,” in IEEE/CVF Int. Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW) , June 2024, pp. 8016–8026. 1
2024
-
[47]
gsplat: An Open-Source Library for Gaussian Splatting,
V . Ye, R. Li, J. Kerr, M. Turkulainen, B. Yi, Z. Pan, O. Seiskari, J. Ye, J. Hu, M. Tancik, and A. Kanazawa, “gsplat: An Open-Source Library for Gaussian Splatting,” arXiv preprint arXiv:2409.06765, 2024. 3
2024 arXiv
-
[48]
KnotDLO: Toward Interpretable Knot Tying,
H. Dinkel, R. Navaratna, J. Xiang, B. Coltin, T. Smith, and T. Bretl, “KnotDLO: Toward Interpretable Knot Tying,” in IEEE Int. Conf. on Robot. and Autom. (ICRA) 3D Visual Representations for Manipulation Workshop , May 2024. 3
2024
-
[49]
TieBot: Learning to Knot a Tie from Visual Demonstration through a Real-to-Sim-to-Real Approach,
W. Peng, J. Lv, Y . Zeng, H. Chen, S. Zhao, J. Sun, C. Lu, and L. Shao, “TieBot: Learning to Knot a Tie from Visual Demonstration through a Real-to-Sim-to-Real Approach,” in Conf. Robot. Learn. (CoRL) , 2024. 3
2024
-
[50]
General Hand–Eye Calibration Based on Reprojection Error Minimization,
K. Koide and E. Menegatti, “General Hand–Eye Calibration Based on Reprojection Error Minimization,” IEEE Robot. Au- tom. Lett. , vol. 4, no. 2, pp. 1021–1028, 2019. 3
2019
-
[51]
Robotic Operating System: Noetic Ninjemys,
Stanford Artificial Intelligence Laboratory, “Robotic Operating System: Noetic Ninjemys,” https://www.ros.org, 2018. 3
2018
-
[52]
Warp: A High-performance Python Framework for GPU Simulation and Graphics,
M. Macklin, “Warp: A High-performance Python Framework for GPU Simulation and Graphics,” March 2022, nVIDIA GPU Technology Conference (GTC). 3
2022
-
[53]
A Robust Deformable Linear Object Perception Pipeline in 3D: From Segmentation to Reconstruction,
S. Zhaole, H. Zhou, L. Nanbo, L. Chen, J. Zhu, and R. B. Fisher, “A Robust Deformable Linear Object Perception Pipeline in 3D: From Segmentation to Reconstruction,” IEEE Robot. Autom. Lett., vol. 9, no. 1, pp. 843–850, 2024. 3
2024
-
[54]
Wire Point Cloud Instance Segmentation from RGBD Imagery with Mask R-CNN,
H. Dinkel, J. Xiang, H. Zhao, B. Coltin, T. Smith, and T. Bretl, “Wire Point Cloud Instance Segmentation from RGBD Imagery with Mask R-CNN,” in IEEE Int. Conf. Robot. Autom. (ICRA) Workshop on Representing and Manipulating Deformable Ob- jects, May 2022
2022
-
[55]
Auto-Generated Wires Dataset for Semantic Segmen- tation with Domain Independence,
R. Zanella, A. Caporali, K. Tadaka, D. De Gregorio, and G. Palli, “Auto-Generated Wires Dataset for Semantic Segmen- tation with Domain Independence,” in IEEE Int. Conf. Comput. Cont. Robot. (ICCCR) . IEEE, Jan. 2021, pp. 292–298
2021
-
[56]
Ari- adne+: Deep Learning-Based Augmented Framework for the Instance Segmentation of Wires,
A. Caporali, R. Zanella, D. De Gregorio, and G. Palli, “Ari- adne+: Deep Learning-Based Augmented Framework for the Instance Segmentation of Wires,” in IEEE Trans. Ind. Inf. , February 2022, pp. 1–11
2022
-
[57]
FASTDLO: Fast Deformable Linear Objects Instance Segmentation,
A. Caporali, K. Galassi, R. Zanella, and G. Palli, “FASTDLO: Fast Deformable Linear Objects Instance Segmentation,” IEEE Robot. Autom. Lett. , vol. 7, no. 4, pp. 9075–9082, 2022
2022
-
[58]
RT-DLO: Real-Time Deformable Linear Objects Instance Segmentation,
A. Caporali, K. Galassi, B. L. ˇZagar, R. Zanella, G. Palli, and A. C. Knoll, “RT-DLO: Real-Time Deformable Linear Objects Instance Segmentation,” IEEE Trans. Ind. Informat. , vol. 19, no. 11, pp. 11 333–11 342, 2023
2023
-
[59]
HAND- LOOM: Learned Tracing of One-Dimensional Objects for In- spection and Manipulation,
V . Viswanath, K. Shivakumar, M. Parulekar, J. Ajmera, J. Kerr, J. Ichnowski, R. Cheng, T. Kollar, and K. Goldberg, “HAND- LOOM: Learned Tracing of One-Dimensional Objects for In- spection and Manipulation,” in Conf. Robot. Learn. (CoRL) , 2023, pp. 341–357. 3
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.