REVIEW 4 major objections 5 minor 46 references
MONA: Moving Object Detection from Videos Shot by Dynamic Camera
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A moving-object detection and segmentation pipeline for dynamic-camera video, integrated with LEAP-VO, cuts trajectory errors by more than 60% on MPI Sintel.
desk verdict Clean new pipeline for masking moving objects, but the evaluation never measures detection quality directly and the trajectory gains are an uncontrolled proxy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the adaptive bounding-box filter with an area-scaled threshold: for each YOLO box $b_i$, the dynamic-point count $D_i$ is compared against $\tau_t^i = \tau_0 \times \mathrm{Area}(b_i)/\mathrm{Area}(b_u)$, where $b_u$ is the smallest box whose count exceeds the base threshold $\tau_0$. A second mechanism is the frame-adaptive dynamic-point threshold: the mean optical-flow magnitude $\bar{m}_t$ of a frame scales the threshold $\theta_t$ that classifies points as dynamic. These two scaling rules let the pipeline work without a fixed hyperparameter across frames with different camera motion and object sizes.
What would settle it
Run MONA on MPI Sintel while replacing its moving-object masks with masks of equal area placed at random, or with a fixed fraction of low-confidence tracked points removed; if trajectory error stays near MONA+LEAP-VO's 0.029 m ATE, the segmentation accuracy is not the cause. Alternatively, compute the IoU between MONA's masks and hand-annotated moving-object masks on any dynamic-camera video; low IoU alongside large trajectory gains would show the claimed mechanism is not what drives the improvement.
Extended reading notes
Core claim
MONA claims that moving objects in dynamic-camera video can be identified without a static background model. The dynamic-points module reuses LEAP-VO's anchor-based probability estimate that each queried pixel is part of a moving object, then refines it with RAFT's optical flow: the mean flow magnitude of a frame sets the threshold for declaring points dynamic. The segmentation module detects all objects with YOLO, keeps only bounding boxes whose dynamic-point count exceeds an area-scaled threshold, and prompts SAM with those boxes to produce object masks. Integrated into LEAP-VO's bundle adjustment, the masks filter out tracked points inside moving regions, yielding an ATE of 0.029 m, relative translation error of 0.013 m, and relative rotation error of 0.054 degrees on MPI Sintel, state-of-the-art among the compared methods. The paper interprets these numbers as evidence that MONA's detection and segmentation are accurate enough to improve downstream camera trajectory estimation.
Load-bearing premise
The claim collapses if the trajectory improvements on MPI Sintel come from generic filtering of unreliable tracked points rather than from MONA's moving-object detection and segmentation, because the paper reports no direct metrics for detection or mask quality.
Editorial extensions
If this is right
- Masking dynamically moving regions during bundle adjustment lets LEAP-VO select tracked points almost entirely from static scene structure, which is why trajectory errors drop by more than 60 percent on MPI Sintel.
- The same masked frames can be fed to other visual odometry and SLAM systems, potentially improving their accuracy in urban scenes with large pedestrians or vehicles.
- Reliable moving-object masks make markerless online video usable for generating pseudo-ground-truth camera trajectories for autonomous driving, UAV planning, and human motion recovery.
- Because MONA needs no training on moving-object annotations, it can be applied to new domains where dynamic objects appear but labeled video is unavailable.
Reading between the lines
- The reported trajectory gains could come from generic removal of outlier tracks rather than from semantically accurate masks; a control that filters the same number of points at random would separate these explanations.
- If the thresholding rules are the real driver, MONA may generalize to moving cameras and objects beyond the categories YOLO is trained on, since no fine-tuning is required.
- One testable extension is to report detection and mask metrics such as IoU and recall on a dynamic-camera benchmark; the current evidence for detection quality is indirect, through trajectory accuracy.
- The adaptive area-scaling threshold could be reused as a cheap prior for other prompt-based segmentation tasks where object size varies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MONA, a two-module framework for detecting and segmenting moving objects in videos captured by dynamic cameras. The first module extracts dynamic points by combining dynamic probabilities from LEAP-VO with RAFT optical flow and an adaptive threshold; the second module filters YOLO bounding boxes using those dynamic points and feeds them as prompts to SAM to obtain segmentation masks. The method is validated indirectly by integrating it with LEAP-VO for camera trajectory estimation on the MPI Sintel dataset, where Table 1 reports ATE 0.029 m, RPE trans 0.013 m, and RPE rot 0.054 deg, described as over 60% improvement over LEAP-VO. No direct metric for moving object detection boxes or segmentation masks is reported.
Significance. If MONA genuinely improves moving-object detection and segmentation in dynamic-camera videos, the contribution would be useful for visual odometry and downstream urban-scene applications. The integration design is reasonable: reusing an established dynamic-probability model, combining off-the-shelf components (RAFT, YOLO, SAM), and demonstrating a downstream trajectory improvement is a sensible way to show utility. I also see no circularity: MONA is not trained to optimize the reported ATE/RPE numbers, and LEAP-VO serves both as a component source and as an evaluation harness. However, the significance of the claimed detection capability rests entirely on an indirect proxy, and the paper currently does not establish that the downstream gain is caused by accurate moving-object segmentation rather than by generic point filtering or dataset-specific effects.
major comments (4)
- [Experiments and Results] The first paragraph of Experiments and Results states that no public dataset is available to directly evaluate detection accuracy and mask quality, and the only quantitative evidence presented is Table 1, which reports ATE and RPE from camera trajectory estimation. This is a load-bearing gap: the paper's title and abstract claim a moving object detection and segmentation capability, but no metric on detected boxes or predicted masks is reported. The downstream proxy is only informative if the causal link between masking quality and trajectory accuracy is established, which the paper does not do. Direct evaluation on a dataset with moving-object masks (e.g., using the MPI Sintel ground-truth segmentation or a synthetic moving-camera set) or at least mask/box metrics on a suitable benchmark is needed to support the central claim.
- [Table 1 and Ablation Study] The claimed >60% improvement over LEAP-VO is attributed to MONA's moving-object detection and segmentation, but no control experiment isolates this mechanism. An ablation replacing MONA's masks with simpler filters, such as removing all points whose optical flow magnitude exceeds a multiple of the mean, removing points inside all YOLO boxes, or randomly removing a matched number of points, would test whether the gain requires the dynamic/static distinction. The existing ablation in Fig. 3 only compares segmentation mask appearance for three prompt types; it does not quantify the effect of these masks on trajectory estimation. Without such a control, the improvement in Table 1 could be due to generic outlier rejection or dataset-specific point selection rather than accurate moving-object segmentation.
- [Methods, Eq. (5) and Eq. (6)] The adaptive threshold θt is never defined. Eq. (5) defines the mean optical flow magnitude m̄_t, and the text says the threshold is 'dynamically scaled' from m̄_t, but the scaling rule is not specified. Similarly, Eq. (6) introduces τ0 as a threshold for dynamic points inside a bounding box, but the paper does not state its value, how it is selected, or how sensitive results are to it. The parameters n, m, k (number of detection points, anchor points, grid size) and λ in Eq. (4) are also left unspecified. These free parameters are load-bearing for reproducibility and for assessing whether the thresholding is genuinely adaptive, so they need to be concretely defined.
- [Table 1] Table 1 reports single-point estimates without error bars, confidence intervals, or repeated-run statistics. The pipeline includes stochastic components (random selection of detection points, CoTracker, RAFT, YOLO, SAM, and random point filtering in LEAP-VO), so run-to-run variability is likely non-negligible. The claim of state-of-the-art performance and the quantitative 'over 60% improvement' should be supported by variance information or multiple runs.
minor comments (5)
- [Methods, Eq. (2)] Eq. (2) uses p(x|V, xq) on both sides of the product; the second factor should be p(y|V, xq), and the product structure should be stated clearly to avoid confusion between the coordinate symbol and the conditioning variables.
- [Introduction] There is a typo in 'UA V' (should be 'UAV'), and the phrase 'In Dynamic Points Extraction , random points are selected' contains an extra space before the comma. A final proofreading pass is needed.
- [Experiments and Results] The qualitative trajectory comparison in Fig. 2 would be more informative with quantitative per-sequence results and a description of how many Sintel sequences show improvement or regression. A per-sequence breakdown of ATE/RPE would strengthen the claim beyond aggregate numbers.
- [References] Several references in the introduction, such as many smart-city and urban-planning citations, are only loosely connected to the technical content. The authors should consider focusing the related-work discussion on moving object detection, dynamic SLAM/VO, and segmentation methods.
- [Moving Object Segmentation, Eq. (6)] The notation u= arg min_{b in B} {Area(b) | D_u ≥ τ0} is ambiguous because the subscript u appears in the condition and in the set being minimized; this should be rewritten for clarity, for example by defining the set of candidate boxes first.
Circularity Check
No significant circularity: MONA is an independent detection/segmentation pipeline; the LEAP-VO integration is an external evaluation, and the acknowledged lack of direct mask metrics is a validation limitation, not circular reasoning.
full rationale
The paper's derivation chain is not circular. MONA's dynamic points are produced by an explicit, independent pipeline: CoTracker trajectories (Eq. 1), a Cauchy dynamic-probability model (Eqs. 2-4) taken from LEAP-VO, and an optical-flow adaptive threshold (Eq. 5). Its moving-object segmentation uses YOLO boxes filtered by dynamic-point counts with a size-scaled threshold (Eq. 6) and SAM masks. None of these quantities is defined in terms of the reported evaluation metrics (ATE, RPE trans., RPE rot.) on MPI Sintel, and no parameter is fitted to those outputs. The reported over-60% improvement of MONA+LEAP-VO over LEAP-VO is an empirical downstream measurement, not an identity or a fitted prediction. LEAP-VO serves both as the source of one component (the dynamic-probability model) and as the evaluation harness, but this reuse is integration, not circularity: the improvement could have failed, and the paper would still not be circular. The paper's own caveat that 'no public dataset is available to directly evaluate detection accuracy and mask quality' is an honest limitation of the proxy validation, and the concern that gains might come from generic point filtering is a correctness or confounding concern outside the circularity definition. No load-bearing self-citations appear; the cited prior work by the current authors (e.g., Wu, Zhao, and He 2024) is not used to justify the central claim. Score 0.
Assumptions & free parameters
free parameters (4)
- τ0
- θt scaling rule
- n, m, k
- λ
assumptions (3)
- domain assumption LEAP-VO's dynamic probability model is reliable on unseen videos
- ad hoc to paper Optical flow magnitude relative to the frame mean separates dynamic from static points
- domain assumption CoTracker, RAFT, YOLO, and SAM perform well enough to support the pipeline
Cite this review
Pith. "Pith review of MONA: Moving Object Detection from Videos Shot by Dynamic Camera." pith.science (2026). https://pith.science/paper/MZQTNSXX
@misc{pith2026250113183,
author = {Pith},
title = {Pith review of: MONA: Moving Object Detection from Videos Shot by Dynamic Camera},
year = {2026},
howpublished = {\url{https://pith.science/paper/MZQTNSXX}},
note = {Machine review of arXiv:2501.13183}
}
read the original abstract
Dynamic urban environments, characterized by moving cameras and objects, pose significant challenges for camera trajectory estimation by complicating the distinction between camera-induced and object motion. We introduce MONA, a novel framework designed for robust moving object detection and segmentation from videos shot by dynamic cameras. MONA comprises two key modules: Dynamic Points Extraction, which leverages optical flow and tracking any point to identify dynamic points, and Moving Object Segmentation, which employs adaptive bounding box filtering, and the Segment Anything for precise moving object segmentation. We validate MONA by integrating with the camera trajectory estimation method LEAP-VO, and it achieves state-of-the-art results on the MPI Sintel dataset comparing to existing methods. These results demonstrate MONA's effectiveness for moving object detection and its potential in many other applications in the urban planning field.
Figures
Reference graph
Works this paper leans on
-
[1]
M.; Civera, J.; and Neira, J
Bescos, B.; Fácil, J. M.; Civera, J.; and Neira, J. 2018. DynaSLAM: Tracking, Mapping, and Inpainting in Dynamic Scenes. IEEE Robotics and Automation Letters, 3(4): 4076--4083
2018
-
[2]
Butler, D. J.; Wulff, J.; Stanley, G. B.; and Black, M. J. 2012. A naturalistic open source movie for optical flow evaluation. In A. Fitzgibbon et al. (Eds.) , ed., European Conf. on Computer Vision (ECCV), Part IV, LNCS 7577, 611--625. Springer-Verlag
work page 2012
-
[3]
R.; Setiaji; Hodge, G.; Pramestri, Z
Caldeira, J.; Fout, A.; Kesari, A.; Sefala, R.; Walsh, J.; Dupre, K.; Khaefi, M. R.; Setiaji; Hodge, G.; Pramestri, Z. A.; et al. 2020. Improving traffic safety through video analysis in Jakarta, Indonesia. In Intelligent Systems and Applications: Proceedings of the 2019 Intelligent Systems Conference (IntelliSys) Volume 2, 642--649. Springer
work page 2020
-
[4]
Chapel, M.-N.; and Bouwmans, T. 2020. Moving objects detection with a moving camera: A comprehensive review. Computer science review, 38: 100310
work page 2020
-
[5]
Chen, L.-H.; Lu, S.; Zeng, A.; Zhang, H.; Wang, B.; Zhang, R.; and Zhang, L. 2024 a . MotionLLM: Understanding Human Behaviors from Human Motions and Videos. arXiv preprint arXiv:2405.20340
arXiv 2024
-
[6]
Chen, W.; Chen, L.; Wang, R.; and Pollefeys, M. 2024 b . LEAP-VO: Long-term Effective Any Point Tracking for Visual Odometry. In CVPR
work page 2024
-
[7]
Chen, X.; Ma, H.; Wan, J.; Li, B.; and Xia, T. 2017. Multi-view 3D Object Detection Network for Autonomous Driving. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6526--6534
work page 2017
-
[8]
Chen, Y.; Liu, S.; Shen, X.; and Jia, J. 2019. Fast Point R-CNN. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 9774--9783
work page 2019
Show all 46 references
-
[9]
Ellenfeld, M.; Moosbauer, S.; Cardenes, R.; Klauck, U.; and Teutsch, M. 2021. Deep fusion of appearance and frame differencing for motion segmentation. In proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4339--4349
2021
-
[10]
Foth, M.; Heikkinen, T.; Ylipulli, J.; Luusua, A.; Satchell, C.; and Ojala, T. 2014. UbiOpticon: participatory sousveillance with urban screens and mobile phone cameras. In Proceedings of The International Symposium on Pervasive Displays, 56--61
2014
-
[11]
Guo, C.; Zuo, X.; Wang, S.; and Cheng, L. 2022. Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. In ECCV, 580--597
2022
-
[12]
Huang, C.; Fang, S.; Wu, H.; Wang, Y.; and Yang, Y. 2024. Low-altitude intelligent transportation: System architecture, infrastructure, and key technologies. Journal of Industrial Information Integration, 42: 100694
2024
-
[13]
Jiang, B.; Chen, X.; Liu, W.; Yu, J.; Yu, G.; and Chen, T. 2024. Motiongpt: Human motion as a foreign language. NeurIPS
2024
-
[14]
Karaev, N.; Rocco, I.; Graham, B.; Neverova, N.; Vedaldi, A.; and Rupprecht, C. 2024. CoTracker: It is Better to Track Together. In Proc. ECCV
2024
-
[15]
Karagulian, F.; Liberto, C.; Corazza, M.; Valenti, G.; Dumitru, A.; and Nigro, M. 2023. Pedestrian flows characterization and estimation with computer vision techniques. Urban Science, 7(2): 65
2023
-
[16]
Khanam, R.; and Hussain, M. 2024. Yolov11: An overview of the key architectural enhancements. arXiv preprint arXiv:2410.17725
2024 arXiv
-
[17]
C.; Lo, W.-Y.; et al
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4015--4026
2023
-
[18]
Ku, J.; Mozifian, M.; Lee, J.; Harakeh, A.; and Waslander, S. L. 2018. Joint 3D Proposal Generation and Object Detection from View Aggregation. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1--8
2018
-
[19]
H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O
Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O. 2019. PointPillars: Fast Encoders for Object Detection From Point Clouds. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12689--12697
2019
-
[20]
Li, B. 2017. 3D fully convolutional network for vehicle detection in point cloud. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1513--1518
2017
-
[21]
Li, B.; Ouyang, W.; Sheng, L.; Zeng, X.; and Wang, X. 2019. GS3D: An Efficient 3D Object Detection Framework for Autonomous Driving. 1019--1028
2019
-
[22]
Li, Z.; Gao, Z.; Wang, K.; Mei, Y.; Zhu, C.; Chen, L.; Wu, X.; and Niyato, D. 2024. Unauthorized UAV Countermeasure for Low-Altitude Economy: Joint Communications and Jamming Based on MIMO Cellular Systems. IEEE Internet of Things Journal, 1--1
2024
-
[23]
Lin, J.; Zeng, A.; Lu, S.; Cai, Y.; Zhang, R.; Wang, H.; and Zhang, L. 2023. Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset. Advances in Neural Information Processing Systems
2023
-
[24]
Moon, G.; Choi, H.; and Lee, K. M. 2022. NeuralAnnot: Neural Annotator for 3D Human Mesh Training Sets. In Computer Vision and Pattern Recognition Workshop (CVPRW)
2022
-
[25]
Mur-Artal, M. J. M. M., Ra\'ul; and Tard\'os, J. D. 2015. ORB-SLAM : a Versatile and Accurate Monocular SLAM System. IEEE Transactions on Robotics, 31(5): 1147--1163
2015
-
[26]
Myagmar-Ochir, Y.; and Kim, W. 2023. A survey of video surveillance systems in smart city. Electronics, 12(17): 3567
2023
-
[27]
Nikolakis, N.; Maratos, V.; and Makris, S. 2019. A cyber physical system (CPS) approach for safe human-robot collaboration in a shared workplace. Robotics and Computer-Integrated Manufacturing, 56: 233--243
2019
-
[28]
E.; Cai, Z.; Yang, L.; Zhang, T.; and Liu, Z
Pang, H. E.; Cai, Z.; Yang, L.; Zhang, T.; and Liu, Z. 2024. Benchmarking and analyzing 3D human pose and shape estimation beyond algorithms. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22. Red Hook, NY, USA: Curran Assoc...
2024
-
[29]
Shen, S.; Cai, Y.; Wang, W.; and Scherer, S. 2023. DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic Environments. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 4048--4055
2023
-
[30]
Shen, Z.; Pi, H.; Xia, Y.; Cen, Z.; Peng, S.; Hu, Z.; Bao, H.; Hu, R.; and Zhou, X. 2024. World-Grounded Human Motion Recovery via Gravity-View Coordinates. In ACM SIGGRAPH Asia
2024
-
[31]
Shi, S.; Wang, X.; and Li, H. 2019. PointRCNN: 3D Object Proposal Generation and Detection From Point Cloud. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[32]
Shi, S.; Wang, Z.; Shi, J.; Wang, X.; and Li, H. 2019. From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network. arXiv preprint arXiv:1907.03670
2019 arXiv
-
[33]
Shin, S.; Kim, J.; Halilaj, E.; and Black, M. J. 2024. WHAM : Reconstructing World-grounded Humans with Accurate 3D Motion. In CVPR
2024
-
[34]
Teed, Z.; and Deng, J. 2020. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16, 402--419. Springer
2020
-
[35]
Teed, Z.; and Deng, J. 2021. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. In Ranzato, M.; Beygelzimer, A.; Dauphin, Y.; Liang, P.; and Vaughan, J. W., eds., NeurIPS, volume 34, 16558--16569. Curran Associates, Inc
2021
-
[36]
Teed, Z.; Lipson, L.; and Deng, J. 2023. Deep Patch Visual Odometry. NeurIPS
2023
-
[37]
Tordesillas, J.; and How, J. P. 2023. Deep-PANTHER: Learning-Based Perception-Aware Trajectory Planner in Dynamic Environments. IEEE Robotics and Automation Letters, 8(3): 1399--1406
2023
-
[38]
Wang, J.; Yuan, Y.; Luo, Z.; Xie, K.; Lin, D.; Iqbal, U.; Fidler, S.; and Khamis, S. 2023. Learning Human Dynamics in Autonomous Driving Scenarios. In ICCV, 20739--20749
2023
-
[39]
Wang, W.; Hu, Y.; and Scherer, S. 2020. TartanVO: A Generalizable Learning-based VO
2020
-
[40]
M.; and Jia, Y
Wang, W.; Li, R.; Chen, Y.; Diekel, Z. M.; and Jia, Y. 2019. Facilitating Human–Robot Collaborative Tasks by Teaching-Learning-Collaboration From Human Demonstrations. IEEE Transactions on Automation Science and Engineering, 16(2): 640--653
2019
-
[41]
Wu, G.; Zhao, Z.; and He, Y. 2024. RELAX: Reinforcement Learning Enabled 2D-LiDAR Autonomous System for Parsimonious UAVs. arXiv:2309.08095
2024 arXiv
-
[42]
Xiong, K.; Xie, J.; and Leng, S. 2024. eVTOL Communication and Trajectory Optimization in Low-Altitude Economy. In 2024 IEEE/CIC International Conference on Communications in China (ICCC Workshops), 845--850
2024
-
[43]
Yazdi, M.; and Bouwmans, T. 2018. New trends on moving object detection in video images captured by a moving camera: A survey. Computer science review, 28: 157--177
2018
-
[44]
Ye, V.; Pavlakos, G.; Malik, J.; and Kanazawa, A. 2023. Decoupling Human and Camera Motion from Videos in the Wild. In CVPR
2023
-
[45]
Yuan, Y.; Iqbal, U.; Molchanov, P.; Kitani, K.; and Kautz, J. 2022. Glamr: Global occlusion-aware human mesh recovery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 11038--11049
2022
-
[46]
Zhou, B.; Pan, J.; Gao, F.; and Shen, S. 2021. RAPTOR: Robust and Perception-Aware Trajectory Replanning for Quadrotor Fast Flight. IEEE Transactions on Robotics, 37(6): 1992--2009
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.