REVIEW 2 major objections 4 minor 104 references
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that fusing 2D motion segmentation masks into layered radiance fields, with test-time refinement, makes 3D-rendered dynamic-object segmentations surpass the 2D baseline in egocentric videos.
desk verdict Solid empirical step on a real benchmark, but the '3D helps 2D' claim needs a 2D test-time-adaptation control before it fully lands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a layered radiance field paired with two mask-fusion losses. The scene is decomposed into static, semi-static, and dynamic layers, each an opacity and color field, with the dynamic layer defined in camera coordinates. Positive motion fusion (PMF) renders a mask from the dynamic layer and pulls it toward the 2D motion mask with an L2 loss; negative motion fusion (NMF) renders a mask from the semi-static layer and penalizes it wherever the 2D model marked a pixel as dynamic, using a binarized mask to select pixels. A self-calibrated RGB loss with learned uncertainty supervises appearance. Test-time refinement re-optimizes only the semi-static and dynamic layers on the selected frames (with an optional temporal window, which the experiments show does not help), so the geometry is fitted to the frames of interest.
What would settle it
Track, per video frame, whether the dynamic object's geometry is captured (for instance by comparing rendered depth or mask intersection-over-union against the image) and compare that to the segmentation gain over the 2D baseline; if the mAP improvement does not concentrate in frames whose geometry is reconstructed, the claim that test-time refinement closes the geometry gap is falsified.
Extended reading notes
Core claim
The central claim is that a layered radiance field can act as a denoising fusion network for 2D motion segmentation. Motion Grouping's masks are incomplete but high-precision, and fusing them into the dynamic layer (positive motion fusion) while penalizing the semi-static layer for agreeing with them (negative motion fusion) yields a single 3D representation whose rendered masks are cleaner and more temporally coherent than the 2D inputs. Because long egocentric videos are too complex for the field to learn complete geometry, test-time refinement over the frames to be analyzed closes the geometry gap; the refinement and the fusion reinforce each other. With NeuralDiff as the backbone, the full method reaches 72.51 mAP on the dynamic UDOS split, surpassing the 2D baseline by 8.2 mAP and reversing the gap documented in EPIC Fields.
Load-bearing premise
Fusion only works when the radiance field has already learned the geometry of the moving object; if that geometry is missing from the reconstruction, the motion labels have nothing to attach to and the gains vanish.
Editorial extensions
If this is right
- The 3D model's rendered dynamic masks beat the 2D Motion Grouping baseline (72.51 vs 64.27 mAP), so the previously documented gap between 3D and 2D dynamic segmentation is closed.
- Semi-static segmentation also improves (27.70 vs 25.55 mAP), even though only dynamic labels from the 2D model are used for fusion.
- The method transfers to other layered architectures: adding TR and PMF to NeRF-W and NeRF-T raises dynamic mAP by roughly 15 to 20 percent, and adding a semi-static layer plus NMF improves them further.
- Combining the approach with a supervised hand-object segmentation method (EgoHOS) boosts that method's mAP from 71.20 to 77.31, while the same fusion with unsupervised Motion Grouping alone slightly exceeds EgoHOS.
- Adding temporal context frames to test-time refinement does not help; focusing optimization on the target frames is what matters.
Reading between the lines
- The mechanism suggests a general recipe: any noisy but high-precision 2D mask that is roughly view-consistent can be lifted into a 3D field and re-rendered cleaner, so the PMF/NMF structure may extend beyond motion to saliency, depth, or open-vocabulary masks.
- Because the bottleneck is geometry completeness, gains should scale with the quality of the underlying 3D model; adopting faster representations such as Gaussian splatting could make test-time refinement practical at interactive speeds.
- The negative fusion loss implicitly forces a competition between semi-static and dynamic layers, which suggests a testable prediction: the method will also sharpen the layer boundary when objects transition from moving to stationary within a video.
- The paper leaves open whether sparse first-frame annotations could replace the 2D segmenter entirely; a testable extension is to fuse a few human clicks instead of Motion Grouping masks and compare mAP.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Tschernezki et al. introduce Layered Motion Fusion (LMF), a method for improving dynamic object segmentation in egocentric videos by fusing motion segmentation predictions from a 2D model (Motion Grouping, MG) into a layered radiance field with static, semi-static, and dynamic layers. Positive motion fusion (PMF) supervises the rendered dynamic-layer mask with MG's prediction, and negative motion fusion (NMF) penalizes the semi-static layer for activating at MG-positive pixels. A test-time refinement (TR) step optimizes the semi-static and dynamic layers on the evaluation frames. On the EPIC Fields UDOS benchmark, the full system (ND+TR+LMF) reaches 72.51 mAP for dynamic segmentation, compared with 64.27 for MG and 55.58 for the NeuralDiff 3D baseline; ablations (Table 5) show that each component contributes. The authors also report gains when applying the fusion to NeRF-T and NeRF-W and when combining with EgoHOS.
Significance. If the comparison were fully controlled, the paper would be significant: it is the first demonstration on EPIC Fields that a 3D neural representation can outperform a strong 2D motion-segmentation model on dynamic objects, directly addressing the negative result reported in EPIC Fields. The paper has several concrete strengths: the ablation in Table 5 isolates PMF/NMF/TR and shows consistent additive gains; the method is applied to three different 3D architectures (Tables 2 and 6); the authors include an explicit discussion of the key failure mode (missing geometry, supporting material Section B); and the runtime analysis is transparent. These positive elements are, however, currently attached to a comparison protocol that does not isolate the claimed 3D advantage, so the main contribution is not yet established.
major comments (2)
- [§4.2, Table 1; §4.3, Table 5] The headline comparison that the 3D model 'surpasses the 2D baseline by a large margin' is not yet established because the 2D baseline is not given the same test-time opportunity. In Table 5, ND+TR without LMF reaches 62.29 mAP on the Dyn task, below the frozen MG baseline of 64.27; the jump to 72.51 comes when MG's own masks are fused back into the 3D model through PMF/NMF (Eqs. 11–12) during test-time refinement. This is not a circularity problem in the technical sense, because the evaluation is against external ground truth, but it means the headline comparison mixes the effect of 3D lifting with the effect of additional per-test-frame optimization and pseudo-label reuse. To support the paper's central claim, add a 2D control with comparable test-time adaptation, e.g., fine-tuning MG on the selected frames or applying temporal CRF/optical-flow propagation to MG masks, evaluated with the same frames and protocol. If such a control also lifts MG to roughly 72 mAP, the 3D fusion would not be the source of the reported gain. The same issue applies to the EgoHOS comparison in Table 4, where the supervised 2D method is frozen while the proposed pipeline receives test-time fitting.
- [§3.3 and §4.1] The test-time refinement procedure is not specified sufficiently for reproducibility or for judging the fairness of the comparison. Equation (13) defines the objective, but the paper does not report the number of optimization steps, the optimizer, the learning rate, the loss weights used during refinement, or whether the refinement uses the same lambda_PMF and lambda_NMF as the base training. It also does not state how many frames are in T for the main results; Table 3 analyzes the neighbor window on only 5 scenes. Since TR is responsible for a 6.71 mAP gain on the Dyn task (Table 5: ND 55.58 vs ND+TR 62.29), these missing details are load-bearing for the method's central evaluation.
minor comments (4)
- [§3.3] The sentence 'we only refine the semi-static and dynamic model Phi_st and Phi_dy' should read 'Phi_ss and Phi_dy'; the subscript 'ss' is used throughout for the semi-static layer, while 'st' denotes the static layer in Eqs. (5)–(7).
- [§4.1] The text 'learning rate of 5×10 4' appears to be a typo for 5×10^-4; please state the value clearly and consistently.
- [§3.2, Eq. (9)] The pseudo-color vectors p_ss=(0,1,0) and p_dy=(0,0,1) are introduced only in the text, and the notation p_ss,1, p_ss,2, p_ss,3 is used without explicit component definitions; define these components before Eq. (9).
- [§4.2, Table 4] The row label 'MG + Ours' is ambiguous because the proposed pipeline always uses MG masks as pseudo-labels; clarify whether this row is identical to ND+TR+LMF from Table 1 or is a separate configuration.
Circularity Check
No significant circularity: the 3D segmentation output is trained from, but not defined by, the 2D motion masks, and evaluation is against external ground truth.
full rationale
The paper's central claim is that a layered radiance field, refined at test time and trained with motion-fusion losses, produces dynamic-object segmentations that beat the 2D Motion Grouping baseline on the EPIC Fields UDOS benchmark. This claim does not reduce to the paper's own inputs by construction. The positive and negative motion fusion losses (Eqs. 11-12) use MG masks as targets, but the evaluated masks are rendered from the optimized neural field via the volume rendering equation (Eq. 8), so the prediction is not identical to the pseudo-label; in fact the final Dyn mAP of 72.51 exceeds the 64.27 of the MG input, which would be impossible if the method merely returned its training signal. Test-time refinement (Eq. 13) optimizes the semi-static and dynamic networks on selected frames using RGB frames and MG pseudo-labels, but ground truth is used only for evaluation, not for fitting; this is transductive adaptation rather than circularity. The paper self-cites prior work (NeuralDiff, EPIC Fields, N3F), but the derivation does not depend on those citations as an unverified premise: NeuralDiff is a released baseline architecture, EPIC Fields provides the benchmark and evaluation script, and the ablations in Table 5 show independent, additive contributions of PMF, NMF, and TR. The strongest legitimate criticism is experimental-design fairness: the MG baseline is not granted test-time refinement or pseudo-label reuse, so the headline gain may be partly attributable to extra test-time compute rather than to 3D lifting. That is a confounding-variable concern about the comparison, not a circular-derivation concern. No equation in the paper defines the claimed output in terms of the quantity it is supposed to predict, and no fitted parameter is renamed as a prediction. The method is therefore not circular; the confound should be evaluated as a correctness or experimental-design risk, not as circularity.
Assumptions & free parameters
free parameters (3)
- lambda_PMF =
1.1
- lambda_NMF =
1.0
- neighboring frame window N =
0 (default)
assumptions (4)
- domain assumption The scene can be decomposed into three layers: static, semi-static, and dynamic, and the dynamic layer is best modeled in camera coordinates.
- domain assumption The 2D motion segmentation model (Motion Grouping) provides masks with low false-positive rate in egocentric videos.
- domain assumption Camera poses for each frame are available and accurate.
- standard math Volume rendering equations (emission-absorption) correctly model the layered scene combination.
Cite this review
Pith. "Pith review of Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos." pith.science (2026). https://pith.science/paper/2WA5LD7G
@misc{pith2026250605546,
author = {Pith},
title = {Pith review of: Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/2WA5LD7G}},
note = {Machine review of arXiv:2506.05546}
}
read the original abstract
Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown that 3D techniques can at last improve outputs extracted from independent 2D views, by fusing them into 3D and denoising them. This is particularly helpful in egocentric videos, where the camera motion is significant, but only under the assumption that the scene itself is static. In fact, as shown in the recent analysis conducted by EPIC Fields, 3D techniques are ineffective when it comes to studying dynamic phenomena, and, in particular, when segmenting moving objects. In this paper, we look into this issue in more detail. First, we propose to improve dynamic segmentation in 3D by fusing motion segmentation predictions from a 2D-based model into layered radiance fields (Layered Motion Fusion). However, the high complexity of long, dynamic videos makes it challenging to capture the underlying geometric structure, and, as a result, hinders the fusion of motion cues into the (incomplete) scene geometry. We address this issue through test-time refinement, which helps the model to focus on specific frames, thereby reducing the data complexity. This results in a synergy between motion fusion and the refinement, and in turn leads to segmentation predictions of the 3D model that surpass the 2D baseline by a large margin. This demonstrates that 3D techniques can enhance 2D analysis even for dynamic phenomena in a challenging and realistic setting.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
N2f2: Hierarchical scene under- standing with nested neural feature fields
Yash Bhalgat, Iro Laina, Jo ˜ao F Henriques, Andrew Zisser- man, and Andrea Vedaldi. N2f2: Hierarchical scene under- standing with nested neural feature fields. InProceedings of the European Conference on Computer Vision (ECCV), pages 197–214. Springer, 2024. 3
2024
-
[2]
Hen- riques, Andrea Vedaldi, and Andrew Zisserman
Yash Bhalgat, Vadim Tschernezki, Iro Laina, Joao F. Hen- riques, Andrea Vedaldi, and Andrew Zisserman. 3d-aware instance segmentation and tracking in egocentric videos. In Proceedings of the Asian Conference on Computer Vision (ACCV). IEEE, 2024. 14
2024
-
[3]
Henriques, Andrea Vedaldi, and Andrew Zisserman
Yash Sanjay Bhalgat, Iro Laina, Joao F. Henriques, Andrea Vedaldi, and Andrew Zisserman. Contrastive lift: 3d object instance segmentation by slow-fast contrastive fusion. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2023. 2, 5
2023
-
[4]
It’s moving! a prob- abilistic model for causal motion segmentation in moving camera videos
Pia Bideau and Erik Learned-Miller. It’s moving! a prob- abilistic model for causal motion segmentation in moving camera videos. InProceedings of the European Conference on Computer Vision (ECCV), 2016. 3
2016
-
[5]
Traditional and recent approaches in background modeling for foreground detection: An overview.Computer science review, 11:31–66, 2014
Thierry Bouwmans. Traditional and recent approaches in background modeling for foreground detection: An overview.Computer science review, 11:31–66, 2014. 3
2014
-
[6]
Object segmentation by long term analysis of point trajectories
Thomas Brox and Jitendra Malik. Object segmentation by long term analysis of point trajectories. InProceedings of the European Conference on Computer Vision (ECCV),
-
[7]
Hexplane: A fast representa- tion for dynamic scenes.Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes.Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[8]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the International Conference on Com- puter Vision (ICCV), 2021. 2
2021
Show all 104 references
-
[9]
Self-supervised learning with geometric constraints in monocular video: Connecting flow, depth, and camera
Yuhua Chen, Cordelia Schmid, and Cristian Sminchis- escu. Self-supervised learning with geometric constraints in monocular video: Connecting flow, depth, and camera. In Proceedings of the International Conference on Computer Vision (ICCV), pages 7063–7072, 2019. 3, 5
2019
-
[10]
Guess What Moves: Unsupervised Video and Image Segmentation by Anticipat- ing Motion
Subhabrata Choudhury, Laurynas Karazija, Iro Laina, An- drea Vedaldi, and Christian Rupprecht. Guess What Moves: Unsupervised Video and Image Segmentation by Anticipat- ing Motion. InProceedings of the British Machine Vision Conference (BMVC), 2022. 3
2022
-
[11]
Csurka, C
G. Csurka, C. R. Dance, L. Dan, J. Willamowski, and C. Bray. Visual categorization with bags of keypoints. InProc. ECCV Workshop on Stat. Learn. in Comp. Vision, 2004. 1
2004
-
[12]
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.Interna- tional Journal of Computer Vision (IJCV), 2022
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.Interna- tio...
2022
-
[13]
An im- age is worth 16×16 words: Transformers for image recog- nition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An im- age is worth 16×16 words: Transformers for image recog- nitio...
2021
-
[14]
Tenen- baum, and Jiajun Wu
Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenen- baum, and Jiajun Wu. Neural radiance flow for 4d view synthesis and video processing. InProceedings of the In- ternational Conference on Computer Vision (ICCV), 2021. 2
2021
-
[15]
NeRF-SOS: Any-view self- supervised object segmentation on complex scenes
Zhiwen Fan, Peihao Wang, Yifan Jiang, Xinyu Gong, De- jia Xu, and Zhangyang Wang. NeRF-SOS: Any-view self- supervised object segmentation on complex scenes. InPro- ceedings of the International Conference on Learning Rep- resentations (ICLR), 2023. 3
2023
-
[16]
Fast dynamic radiance fields with time-aware neural voxels
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xi- aopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. Fast dynamic radiance fields with time-aware neural voxels. InProc. SIGGRAPH Asia, 2022. 2
2022
-
[17]
K- planes: Explicit radiance fields in space, time, and appear- ance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K- planes: Explicit radiance fields in space, time, and appear- ance. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[18]
Panoptic NeRF: 3d-to-2d label transfer for panoptic urban scene segmentation.arXiv.cs, abs/2203.15224, 2022
Xiao Fu, Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou, Andreas Geiger, and Yiyi Liao. Panoptic NeRF: 3d-to-2d label transfer for panoptic urban scene segmentation.arXiv.cs, abs/2203.15224, 2022. 2
2022 arXiv
-
[19]
Dynamic view synthesis from dynamic monocu- lar video.arXiv.cs, abs/2105.06468, 2021
Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocu- lar video.arXiv.cs, abs/2105.06468, 2021. 2
2021 arXiv
-
[20]
Monocular dynamic view synthe- sis: A reality check
Hang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell, and Angjoo Kanazawa. Monocular dynamic view synthe- sis: A reality check. InProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2022. 6, 7, 16
2022
-
[21]
Monocular dynamic view synthe- sis: A reality check.arXiv.cs, abs/2210.13445, 2022
Hang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell, and Angjoo Kanazawa. Monocular dynamic view synthe- sis: A reality check.arXiv.cs, abs/2210.13445, 2022. 16
2022 arXiv
-
[22]
Girshick
Ross B. Girshick. Fast R-CNN. InProceedings of the In- ternational Conference on Computer Vision (ICCV), 2015. 1
2015
-
[23]
Neural deformable voxel grid for fast optimization of dynamic view synthe- sis
Xiang Guo, Guanying Chen, Yuchao Dai, Xiaoqing Ye, Ji- adai Sun, Xiao Tan, and Errui Ding. Neural deformable voxel grid for fast optimization of dynamic view synthe- sis. InProceedings of the Asian Conference on Computer Vision (ACCV), 2022. 2
2022
-
[24]
Deep residual learning for image recognition.arXiv preprint arXiv:1512.03385, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition.arXiv preprint arXiv:1512.03385, 2015. 1
2015 arXiv
-
[25]
Dense 3d semantic mapping of indoor scenes from rgb- d images
Alexander Hermans, Georgios Floros, and Bastian Leibe. Dense 3d semantic mapping of indoor scenes from rgb- d images. In2014 IEEE International Conference on Robotics and Automation (ICRA), pages 2631–2638. IEEE,
-
[26]
Sfm-ttr: Using struc- ture from motion for test-time refinement of single-view depth networks
Sergio Izquierdo and Javier Civera. Sfm-ttr: Using struc- ture from motion for test-time refinement of single-view depth networks. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 21466–21476, 2023. 2, 3, 5
2023
-
[27]
Fusion- seg: Learning to combine motion and appearance for fully automatic segmentation of generic objects in videos
Suyog Dutt Jain, Bo Xiong, and Kristen Grauman. Fusion- seg: Learning to combine motion and appearance for fully automatic segmentation of generic objects in videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 3
2017
-
[28]
What uncertainties do we need in Bayesian deep learning for computer vision?Proceed- ings of Advances in Neural Information Processing Systems (NeurIPS), 2017
Alex Kendall and Yarin Gal. What uncertainties do we need in Bayesian deep learning for computer vision?Proceed- ings of Advances in Neural Information Processing Systems (NeurIPS), 2017. 4
2017
-
[29]
3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4), 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4), 2023. 3, 17
2023
-
[30]
Lerf: Language embed- ded radiance fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embed- ded radiance fields. InProceedings of the International Conference on Computer Vision (ICCV), 2023. 1, 2, 3
2023
-
[31]
Garfield: Group anything with radiance fields.arXiv.cs, abs/2401.09419, 2024
Chung Min Kim, Mingxuan Wu, Justin Kerr, Ken Goldberg, Matthew Tancik, and Angjoo Kanazawa. Garfield: Group anything with radiance fields.arXiv.cs, abs/2401.09419, 2024. 1
2024 arXiv
-
[32]
Decomposing neRF for editing via feature field dis- tillation
Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitz- mann. Decomposing neRF for editing via feature field dis- tillation. InProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 2, 3, 5
2022
-
[33]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural net- works. InProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2012. 1
2012
-
[34]
Joint semantic segmentation and 3d re- construction from monocular video
Abhijit Kundu, Yin Li, Frank Dellaert, Fuxin Li, and James M Rehg. Joint semantic segmentation and 3d re- construction from monocular video. InComputer Vision– ECCV 2014: 13th European Conference, Zurich, Switzer- land, September 6-12, 2014, Proceedings, Part VI 13, pages 703–...
2014
-
[35]
Panoptic neural fields: A semantic object-aware neural scene rep- resentation
Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J Guibas, Andrea Tagliasac- chi, Frank Dellaert, and Thomas Funkhouser. Panoptic neural fields: A semantic object-aware neural scene rep- resentation. InProceedings of the IEEE Conference on Co...
2022
-
[36]
Divided attention: Unsuper- vised multiple-object discovery and segmentation with in- terpretable contextually separated slots, 2024
Dong Lao, Zhengyang Hu, Francesco Locatello, Yanchao Yang, and Stefano Soatto. Divided attention: Unsuper- vised multiple-object discovery and segmentation with in- terpretable contextually separated slots, 2024. 3
2024
-
[37]
Language-driven semantic seg- mentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. InProceedings of the International Conference on Learning Representations (ICLR), 2022. 2
2022
-
[38]
New- combe, and Zhaoyang Lv
Tianye Li, Mira Slavcheva, Michael Zollh ¨ofer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. New- combe, and Zhaoyang Lv. Neural 3D video synthesis from multi-view video. InProceedings of the IEEE Confer- ence on Co...
-
[39]
Neural scene flow fields for space-time view synthe- sis of dynamic scenes
Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthe- sis of dynamic scenes. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[40]
Dynibar: Neural dynamic image-based rendering
Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, and Noah Snavely. Dynibar: Neural dynamic image-based rendering. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[41]
Long Lian, Zhirong Wu, and Stella X. Yu. Bootstrapping objectness from videos by relaxed common fate and vi- sual grouping. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 14582–14591, 2023. 3
2023
-
[42]
A comprehensive sur- vey on test-time adaptation under distribution shifts, 2023
Jian Liang, Ran He, and Tieniu Tan. A comprehensive sur- vey on test-time adaptation under distribution shifts, 2023. 2, 3
2023
-
[43]
Semantic attention flow fields for monocular dynamic scene decomposition
Yiqing Liang, Eliot Laidlaw, Alexander Meyerowitz, Sri- nath Sridhar, and James Tompkin. Semantic attention flow fields for monocular dynamic scene decomposition. InPro- ceedings of the International Conference on Computer Vi- sion (ICCV), pages 21797–21806, 2023. 1, 2, 3
2023
-
[44]
Devrf: Fast deformable voxel radi- ance fields for dynamic scenes.Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 35: 36762–36775, 2022
Jia-Wei Liu, Yan-Pei Cao, Weijia Mao, Wenqiao Zhang, David Junhao Zhang, Jussi Keppo, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Devrf: Fast deformable voxel radi- ance fields for dynamic scenes.Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 35: 3...
2022
-
[45]
David G. Lowe. Object recognition from local scale- invariant features. InProceedings of the International Con- ference on Computer Vision (ICCV), 1999. 1
1999
-
[46]
Multi-view deep learning for consistent semantic mapping with RGB-D cameras
Lingni Ma, J ¨org St ¨uckler, Christian Kerl, and Daniel Cre- mers. Multi-view deep learning for consistent semantic mapping with RGB-D cameras. InProc.IROS, 2017. 3
2017
-
[47]
Marr.Vision: A computational investigation into the human representation and processing of visual information
D. Marr.Vision: A computational investigation into the human representation and processing of visual information. New Yorl: WH Freeman, 1982. 1
1982
-
[48]
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Saj- jadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. InProceedings of the IEEE Conference on Computer Vision and Pattern Recog- ni...
2021
-
[49]
Dif- fuser: Multi-view 2D-to-3D label diffusion for semantic scene segmentation
Ruben Mascaro, Lucas Teixeira, and Margarita Chli. Dif- fuser: Multi-view 2D-to-3D label diffusion for semantic scene segmentation. InProceedings of the IEEE Confer- ence on Robotics and Automation (ICRA), 2021. 3
2021
-
[50]
A review of motion segmentation: Approaches and ma- jor challenges
Jana Mattheus, Hans Grobler, and Adnan M Abu-Mahfouz. A review of motion segmentation: Approaches and ma- jor challenges. In2020 2nd International multidisci- plinary information technology and engineering conference (IMITEC), pages 1–8. IEEE, 2020. 3
2020
-
[51]
Scene chronology
Kevin Matzen and Noah Snavely. Scene chronology. In Proceedings of the European Conference on Computer Vi- sion (ECCV), 2014. 3
2014
-
[52]
Semanticfusion: Dense 3d seman- tic mapping with convolutional neural networks
John McCormac, Ankur Handa, Andrew Davison, and Stefan Leutenegger. Semanticfusion: Dense 3d seman- tic mapping with convolutional neural networks. In2017 IEEE International Conference on Robotics and automa- tion (ICRA), pages 4628–4635. IEEE, 2017. 3
2017
-
[53]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProceedings of the European Conference on Computer Vision (ECCV), 2020. 1, 2
2020
-
[54]
Laterf: Label and text driven object radi- ance fields
Ashkan Mirzaei, Yash Kant, Jonathan Kelly, and Igor Gilitschenski. Laterf: Label and text driven object radi- ance fields. InProceedings of the European Conference on Computer Vision (ECCV), 2022. 3
2022
-
[55]
Learning visual motion segmentation using event surfaces
Anton Mitrokhin, Zhiyuan Hua, Cornelia Fermuller, and Yiannis Aloimonos. Learning visual motion segmentation using event surfaces. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[56]
Panopticfusion: Online volumetric semantic map- ping at the level of stuff and things
Gaku Narita, Takashi Seno, Tomoya Ishikawa, and Yohsuke Kaji. Panopticfusion: Online volumetric semantic map- ping at the level of stuff and things. In2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4205–4212. IEEE, 2019. 3
2019
-
[57]
Captur- ing the geometry of object categories from video supervi- sion.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018
David Novotn ´y, Diane Larlus, and Andrea Vedaldi. Captur- ing the geometry of object categories from video supervi- sion.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018. 4
2018
-
[58]
P. Ochs, J. Malik, and T. Brox. Segmentation of moving objects by long term video analysis.PAMI, 36(6):1187 – 1200, 2014. Preprint. 3
2014
-
[59]
Neural scene graphs for dynamic scenes
Julian Ost, Fahim Mannan, Nils Thuerey, Julian Knodt, and Felix Heide. Neural scene graphs for dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[60]
Fast object seg- mentation in unconstrained video
Anestis Papazoglou and Vittorio Ferrari. Fast object seg- mentation in unconstrained video. InProceedings of the In- ternational Conference on Computer Vision (ICCV), 2013. 3
2013
-
[61]
Nerfies: Deformable neural radiance fields
Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. InProceedings of the International Conference on Com- puter Vision (ICCV), 2021. 2
2021
-
[62]
Perronnin and C
F. Perronnin and C. Dance. Fisher kernels on visual vo- cabularies for image categorizaton. InProceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), 2006. 1
2006
-
[63]
Spatial cognition from egocentric video: Out of sight, not out of mind
Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, and Dima Damen. Spatial cognition from egocentric video: Out of sight, not out of mind. In 2025 International Conference on 3D Vision (3DV), 2025. 14
2025
-
[64]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[65]
Langsplat: 3d language gaussian splatting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. Langsplat: 3d language gaussian splatting. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 20051–20060, 2024. 3
2024
-
[66]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. InProceed- ings of t...
2021
-
[67]
Freeman, Fredo Durand, Joshua B
Prafull Sharma, Ayush Tewari, Yilun Du, Sergey Zakharov, Rares Andrei Ambrus, Adrien Gaidon, William T. Freeman, Fredo Durand, Joshua B. Tenenbaum, and Vincent Sitz- mann. Neural groundplans: Persistent neural scene repre- sentations from a single image. InProceedings of the I...
-
[68]
Panoptic lifting for 3d scene understand- ing with neural fields
Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bul `o, Nor- man M ¨uller, Matthias Nießner, Angela Dai, and Peter Kontschieder. Panoptic lifting for 3d scene understand- ing with neural fields. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[69]
Very deep con- volutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep con- volutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR), 2015. 1
2015
-
[70]
Sivic and A
J. Sivic and A. Zisserman. Video Google: Efficient visual search of videos. InToward Category-Level Object Recog- nition. Springer, 2006. 1
2006
-
[71]
NeRFPlayer: A streamable dynamic scene representa- tion with decomposed neural radiance fields.arXiv.cs, abs/2210.15947, 2022
Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. NeRFPlayer: A streamable dynamic scene representa- tion with decomposed neural radiance fields.arXiv.cs, abs/2210.15947, 2022. 2
-
[72]
Dense point trajectories by gpu-accelerated large displace- ment optical flow
Narayanan Sundaram, Thomas Brox, and Kurt Keutzer. Dense point trajectories by gpu-accelerated large displace- ment optical flow. InProceedings of the European Confer- ence on Computer Vision (ECCV), 2010. 3
2010
-
[73]
Meaningful maps with object-oriented semantic mapping
Niko S ¨underhauf, Trung T Pham, Yasir Latif, Michael Mil- ford, and Ian Reid. Meaningful maps with object-oriented semantic mapping. In2017 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS), pages 5079–5085. IEEE, 2017. 3
2017
-
[74]
Cnn-slam: Real-time dense monocular slam with learned depth prediction
Keisuke Tateno, Federico Tombari, Iro Laina, and Nassir Navab. Cnn-slam: Real-time dense monocular slam with learned depth prediction. InProceedings of the IEEE con- ference on computer vision and pattern recognition, pages 6243–6252, 2017. 3
2017
-
[75]
Learning motion patterns in videos
Pavel Tokmakov, Karteek Alahari, and Cordelia Schmid. Learning motion patterns in videos. InProceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), 2017. 3
2017
-
[76]
Training data-efficient image transformers & distillation through at- tention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv´e J´egou. Training data-efficient image transformers & distillation through at- tention. InProceedings of the International Conference on Machine Learning (ICML), 2021. 2
2021
-
[77]
Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monoc- ular video
Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollh ¨ofer, Christoph Lassner, and Christian Theobalt. Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monoc- ular video. InProceedings of the International Conference ...
2021
-
[78]
NeuralDiff: Segmenting 3D objects that move in egocen- tric videos
Vadim Tschernezki, Diane Larlus, and Andrea Vedaldi. NeuralDiff: Segmenting 3D objects that move in egocen- tric videos. InProceedings of the International Conference on 3D Vision (3DV), 2021. 1, 2, 3, 4, 6, 7, 8, 16
2021
-
[79]
Neural Feature Fusion Fields: 3D distillation of self-supervised 2D image representation
Vadim Tschernezki, Iro Laina, Diane Larlus, and Andrea Vedaldi. Neural Feature Fusion Fields: 3D distillation of self-supervised 2D image representation. InProceedings of the International Conference on 3D Vision (3DV), 2022. 1, 2, 3, 5
2022
-
[80]
EPIC Fields: Marrying 3D Geometry and Video Understanding
Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Laina, Diane Larlus, Dima Damen, and Andrea Vedaldi. EPIC Fields: Marrying 3D Geometry and Video Understanding. InProceedings of Advances in Neural In- formation Processing Systems (NeurIPS), 2023. 1, 2, 4, 5, 6...
2023
-
[81]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InNIPS,
-
[82]
In- cremental dense semantic stereo fusion for large-scale se- mantic scene reconstruction
Vibhav Vineet, Ondrej Miksik, Morten Lidegaard, Matthias Nießner, Stuart Golodetz, Victor A Prisacariu, Olaf K¨ahler, David W Murray, Shahram Izadi, Patrick P ´erez, et al. In- cremental dense semantic stereo fusion for large-scale se- mantic scene reconstruction. In2015 IEEE ...
-
[83]
Suhani V ora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi S. M. Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth. NeSF: Neural se- mantic fields for generalizable semantic segmentation of 3D scenes.arXiv.cs, abs/2111.13260, 2021. 2
2021 arXiv
-
[84]
Dm-nerf: 3d scene geometry decomposition and manipulation from 2d images
Bing Wang, Lu Chen, and Bo Yang. Dm-nerf: 3d scene geometry decomposition and manipulation from 2d images. arXiv preprint arXiv:2208.07227, 2022. 2
2022 arXiv
-
[85]
Neural trajectory fields for dynamic novel view syn- thesis.arXiv.cs, abs/2105.05994, 2021
Chaoyang Wang, Ben Eckart, Simon Lucey, and Orazio Gallo. Neural trajectory fields for dynamic novel view syn- thesis.arXiv.cs, abs/2105.05994, 2021. 2
2021 arXiv
-
[86]
Fourier PlenOctrees for dynamic radiance field rendering in real-time
Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. Fourier PlenOctrees for dynamic radiance field rendering in real-time. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[87]
Zero-shot video object segmenta- tion via attentive graph neural networks
Wenguan Wang, Xiankai Lu, Jianbing Shen, David J Cran- dall, and Ling Shao. Zero-shot video object segmenta- tion via attentive graph neural networks. InProceedings of the International Conference on Computer Vision (ICCV),
-
[88]
Tianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, For- rester Cole, and Cengiz Oztireli. Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocu- lar video.Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 35:32653–32666, 2022...
2022
-
[89]
Space-time neural irradiance fields for free-viewpoint video
Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
2021
-
[90]
Object discovery in videos as foreground motion clus- tering
Christopher Xie, Yu Xiang, Zaid Harchaoui, and Dieter Fox. Object discovery in videos as foreground motion clus- tering. InProceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2019. 3
2019
-
[91]
Fig-nerf: Figure-ground neural radi- ance fields for 3d object category modelling
Christopher Xie, Keunhong Park, Ricardo Martin-Brualla, and Matthew Brown. Fig-nerf: Figure-ground neural radi- ance fields for 3d object category modelling. InProceed- ings of the International Conference on 3D Vision (3DV), pages 962–971. IEEE, 2021. 3
2021
-
[92]
Segmenting moving objects via an object-centric layered representation
Junyu Xie, Weidi Xie, and Andrew Zisserman. Segmenting moving objects via an object-centric layered representation. InProceedings of Advances in Neural Information Process- ing Systems (NeurIPS), 2022. 3
2022
-
[93]
Self-supervised video object segmen- tation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisser- man, and Weidi Xie. Self-supervised video object segmen- tation by motion grouping. InProceedings of the Interna- tional Conference on Computer Vision (ICCV), 2021. 2, 3, 5, 6, 8, 15
2021
-
[94]
Learning to segment rigid motions from two frames
Gengshan Yang and Deva Ramanan. Learning to segment rigid motions from two frames. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[95]
Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera
Jae Shin Yoon, Kihwan Kim, Orazio Gallo, Hyun Soo Park, and Jan Kautz. Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
2020
-
[96]
Unsuper- vised discovery of object radiance fields
Hong-Xing Yu, Leonidas Guibas, and Jiajun Wu. Unsuper- vised discovery of object radiance fields. InProceedings of the International Conference on Learning Representations (ICLR), 2022. 3
2022
-
[97]
Star: Self-supervised tracking and reconstruc- tion of rigid objects in motion with neural rendering
Wentao Yuan, Zhaoyang Lv, Tanner Schmidt, and Steven Lovegrove. Star: Self-supervised tracking and reconstruc- tion of rigid objects in motion with neural rendering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
2021
-
[98]
Motion segmentation: A review.Artificial Intelligence Research and Development, pages 398–407, 2008
Luca Zappella, Xavier Llad ´o, and Joaquim Salvi. Motion segmentation: A review.Artificial Intelligence Research and Development, pages 398–407, 2008. 3
2008
-
[99]
Fine-grained egocentric hand-object segmentation: Dataset, model, and applications
Lingzhi Zhang, Shenghao Zhou, Simon Stent, and Jianbo Shi. Fine-grained egocentric hand-object segmentation: Dataset, model, and applications. InProceedings of the European Conference on Computer Vision (ECCV), pages 127–145. Springer, 2022. 6, 8, 14
2022
-
[100]
Nerflets: Local radiance fields for efficient structure-aware 3d scene representation from 2d supervision
Xiaoshuai Zhang, Abhijit Kundu, Thomas Funkhouser, Leonidas Guibas, Hao Su, and Kyle Genova. Nerflets: Local radiance fields for efficient structure-aware 3d scene representation from 2d supervision. InProceedings of the IEEE Conference on Computer Vision and Pattern Recog- ni...
2023
-
[101]
Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and Andrew J. Davison. In-place scene labelling and under- standing with implicit scene representation. InProceed- ings of the International Conference on Computer Vision (ICCV), 2021. 1, 2, 3, 5, 14, 15
2021
-
[102]
Feature 3dgs: Su- percharging 3d gaussian splatting to enable distilled fea- ture fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Su- percharging 3d gaussian splatting to enable distilled fea- ture fields. InProceedings of the IEEE Conference on Computer ...
2024
-
[103]
Fmgs: Foundation model embedded 3d gaussian splatting for holistic 3d scene understanding,
Xingxing Zuo, Pouya Samangouei, Yunwen Zhou, Yan Di, and Mingyang Li. Fmgs: Foundation model embedded 3d gaussian splatting for holistic 3d scene understanding,
-
[2024]
First, we show how our proposed method improves the semi-static segmentation through negative motion fusion (NMF) in Section A
3 Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos Supplementary Material The supplementary material is structured as follows. First, we show how our proposed method improves the semi-static segmentation through negative motion fusion (NMF) in Sect...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.