REVIEW 4 major objections 6 minor 1 cited by
On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A drone that keeps learning depth while flying avoids obstacles 30 percent better than one flown on pre-training alone.
desk verdict Real on-board online learning for event-based depth, but the metric MAE claim needs scale-calibration documentation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the contrast-maximization loss, which warps accumulated events along a predicted optical flow and measures how sharp the resulting image of warped events is; sharper means the motion estimate is more consistent with the event stream. The paper constructs optical flow from depth and ego-motion through the projective equation $x' \sim K P D(x) K^{-1} x$, turning depth estimation into a self-supervised problem that needs no ground truth. The load-bearing mechanism is a per-event parallel CUDA implementation of the warp, splat, and backward gradient computations: one thread per event, no zero-padding, no processing of events that warp outside the image, and analytical gradients instead of autograd. This drops the cost of the loss and its gradient from over 10 ms to well under a millisecond on a desktop GPU, which is what makes on-device learning possible. A geometry-consistency loss on depth across consecutive frames stabilizes the scale of predictions.
What would settle it
Run the same online-learning flight experiment in a second indoor environment without artificial texture and with different obstacle placements; if the distance between pilot interventions and the held-out depth MAE do not improve over the pre-trained-only model, the reported benefit is specific to the first test setup rather than a general property of the method.
Extended reading notes
Core claim
The paper's central claim is that online, on-device self-supervised learning from event data improves monocular depth estimation and obstacle-avoidance behavior on a small flying drone compared with pre-training alone. It supports this claim by reimplementing the contrast-maximization loss as per-event parallel CUDA kernels that warp and splat events independently, skip events that leave the image, and supply analytical gradients, reducing runtime by about 100x and memory use by 2-5x relative to a batched PyTorch baseline. With this efficiency, a 430k-parameter recurrent network can run forward and backward passes at about 30 Hz on the drone's embedded GPU while learning. After roughly two minutes of flight, the network's depth error on a held-out test sequence decreased and its RSAT (deblurring-quality) metric improved; in repeated flights, the distance between human-pilot interventions grew by about 30% when online learning was added to pre-training. Training from scratch did not produce meaningful depth within the flight time, indicating that pre-training provides the representation that online learning then adapts to the operational environment.
Load-bearing premise
The training signal assumes the scene is completely static with no occlusions or disocclusions, because optical flow is constructed purely from depth and camera pose; any independently moving object produces a wrong warp and a corrupt gradient.
Editorial extensions
If this is right
- Online learning converges within about two minutes of flight, so a drone can adapt its depth perception in the field without ground truth.
- Pre-training on a diverse dataset remains necessary; training from scratch on the same flight data does not produce meaningful depth in the same time.
- The efficiency gains are not limited to depth: any pipeline involving warping and splatting of events (or images) could use the same per-event parallel CUDA approach.
- The depth model achieves state-of-the-art among self-supervised event-only methods on MVSEC and runs at higher frequency than ground truth (100 Hz vs 10 Hz), avoiding boundary artifacts.
- Obstacle avoidance can be driven directly by relative (scale-free) depth via binning inverse depth into yaw commands, so metric scale is not required for control.
Reading between the lines
- Because the per-event parallel pattern removes padding and avoids autograd overhead, the same CUDA kernels could be dropped into other event-based losses (for example photometric or motion-segmentation losses), potentially extending online learning to dynamic scenes and to even smaller drones.
- The 30% intervention-distance improvement conflates depth quality with the control policy; an ablation that holds the control law fixed and measures only depth error on the obstacle region would isolate how much of the gain comes from perception.
- The static-scene assumption is the obvious next target: adding a flow-confidence mask or motion segmentation to the contrast-maximization loss could let online learning keep adapting in environments with people or vehicles, which the paper's own qualitative results show are underestimated.
- If the scale-free depth is consistent over time, the same yaw-rate control scheme could be reused for other tasks like following a corridor or landing, without metric calibration; this is untested.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a CUDA-accelerated, per-event parallel implementation of the contrast-maximization loss for self-supervised monocular depth and ego-motion estimation from event cameras. It reports roughly 100x runtime and 2-5x memory improvements over batched PyTorch processing, enabling online, on-board learning on a small quadrotor. The authors present benchmark results on MVSEC and DSEC and a drone experiment in which pre-training plus online learning is claimed to improve both metric depth accuracy (MAE) and obstacle-avoidance behavior (distance between pilot interventions) relative to pre-training only.
Significance. If substantiated, this is a timely and significant robotics contribution: it demonstrates a feasible path to on-device, self-supervised continual learning for event-based depth, with a real flight experiment and external behavioral metrics rather than only loss-based evaluation. The paper's strengths include the concrete runtime/memory characterization of the custom CUDA kernels, the use of independent ground-truth depth and pilot-intervention distances for the online-learning claim, and unusually candid supplementary material about what the contrast-maximization loss can and cannot improve. The main weaknesses are in the evaluation protocol: metric-scale calibration for scale-ambiguous monocular depth is undocumented for the central MAE curves and for the MVSEC benchmark, one DSEC row uses a test-set-fitted scale, and the flight-statistics reporting is incomplete. These are addressable reporting and analysis gaps rather than fundamental contradictions, so the core idea and system demonstration remain credible.
major comments (4)
- [Sec. 4.2, Fig. 5 (right), Eq. (2)] The MAE-vs-ground-truth curve is the main quantitative evidence that online learning improves depth accuracy, but the network is scale-free: Eq. (2) depends on depth only through the ratio t/D, so a global depth scale is unobservable. The supplementary material states in Sec. 6 that the first phase of online learning improves contrast 'mostly due to just learning the correct magnitude of the optical flow ... through scaling depth and ego-motion.' The manuscript does not state how metric scale was assigned to each checkpoint evaluated in Fig. 5 (right), nor whether a per-checkpoint scale was fitted on the test sequence. If such fitting was done, the reported MAE decrease could reflect scale alignment rather than improved depth structure; the obstacle-avoidance controller in Sec. 3.4 uses inverse-depth differences, so it too can improve from a global scale change alone. Please document the calibration protocol and report scale factors per checkpoint, or replace metric MAE with scale-invariant errors such as scale-invariant log error, delta-1 accuracy, or median absolute relative error.
- [Sec. 4.1, Table 1] The MVSEC MAE values are reported in meters, yet no text describes how the monocular, scale-ambiguous depth output was brought to metric scale for this benchmark. This makes the MVSEC results non-reproducible and leaves the state-of-the-art claim among self-supervised methods unsupported. Please specify the calibration procedure (e.g., training-set median ratio, per-sequence fit, or stereo-based scale) or report scale-invariant metrics alongside the metric numbers.
- [Sec. 4.1, Table 2] The row labeled 'Ours (best scale)' is produced by a grid search on the test set and is therefore an oracle-scale upper bound rather than a valid evaluation of the method's metric accuracy. The 'approx. scale' row is the defensible comparison and should be the headline result; 'best scale' should be moved to the supplementary material or explicitly labeled as an oracle. As currently presented, the table invites an overly favorable reading of the method's disparity accuracy.
- [Sec. 4.2, Fig. 5 (left)] The headline claim of a roughly 30% improvement in distance between pilot interventions is presented without the number of flights per condition, the intervention protocol, or any variance/statistical summary beyond boxplots. Please report the sample sizes, define the intervention criterion, and state whether the PT and PT+OL conditions use the same pretrained weights and the same obstacle layout. Without these details, the improvement cannot be distinguished from run-to-run variation.
minor comments (6)
- [Table 1] There is a stray '2' before the first Zhu et al. [46] row; this looks like an orphaned footnote marker and should be cleaned up.
- [References] Reference [8] contains a trailing '3' in the reference list that appears to be a page or citation artifact; please remove it.
- [Fig. 5 caption] The caption quotes percentages of ~65% and ~30% but does not state whether these are computed from means or medians of the intervention distances; please specify.
- [Sec. 4.2] RSAT is defined only in the Fig. 5 caption; please define it in the main text at first use so the reader can interpret the right-hand panel without looking at the caption.
- [Sec. 4.1, Table 1] The meaning of 'Ours (dense)' is explained only in passing in the text; please add a sentence explicitly defining dense depth as depth that is not masked by events.
- [Sec. 5 / Limitations] The Limitations paragraphs note that dynamic objects are underestimated and that the yaw controller is attracted to corners; these are important scope restrictions and should be reflected in the conclusion's claims, which currently read as if the obstacle-avoidance benefit is shown for general environments.
Circularity Check
Main claim is externally benchmarked; two minor self-referential evaluation choices (RSAT and DSEC best-scale) are contained and non-load-bearing.
-
other
[Section 4.2, Drone experiments, Fig. 5 right (page 7)]
"The model shows significant improvement not only in the RSAT (ratio of squared average timestamps) metric [19], which is strongly correlated with the contrast maximization loss used to optimize the network, but also in MAE (mean absolute error) when compared against ground truth depth."
RSAT is a deblurring statistic over the warped-event timestamps that the contrast-maximization loss in Eq. (1) directly rewards; reporting RSAT improvement is thus a near-restatement of the training objective under another name. The paper itself concedes the correlation. This metric is only illustrative, however: the depth-accuracy claim is carried by the external MAE evaluation and by the pilot-intervention distances, so the self-referential metric is not load-bearing.
-
fitted input called prediction
[Section 4.1, DSEC evaluation after Table 2 (page 6)]
"Additionally, we conduct a grid search on the scaling factor to achieve the highest accuracy on the test set, reported as 'best scale'."
The 'best scale' rows in Table 2 report MAE/RMSE after fitting the monocular-depth scale factor to the test-set ground truth, i.e., after minimizing the exact metric being reported. Those numbers are therefore test-set-optimized lower bounds rather than predictions. The paper does label this row and also reports an 'approx. scale' row, and the abstract's state-of-the-art claim rests on the MVSEC results, not on this DSEC table, so the practice is contained rather than central.
full rationale
The central claim is that online learning on the drone improves depth accuracy and obstacle avoidance relative to pre-training only. The supporting evidence is external to the training loss: Fig. 5 reports MAE against ground-truth depth on an unseen held-out sequence, and the flight experiments report distances between human-pilot interventions. Neither quantity is a function of the contrast-maximization objective, so the main derivation chain is not circular. The MVSEC state-of-the-art claim is likewise a comparison against external self-supervised baselines using ground-truth MAE. Two contained self-referential aspects are present. First, the RSAT curve in Fig. 5 is acknowledged by the authors to be strongly correlated with the loss being optimized, making it a near-duplicate diagnostic rather than independent evidence. Second, the DSEC 'best scale' row is obtained by grid-searching the scale factor on the test set, so those numbers are test-set-fitted rather than predictions; the adjacent 'approx. scale' row is the honest predictive comparison, and the paper does not rest its headline claim on this table. The reliance on the authors' own earlier works [19], [28], and [40] for the contrast-maximization pipeline, RSAT, and architecture is not load-bearing circularity: these are prior peer-reviewed results, and the on-device fine-tuning contribution is validated by external metrics. A separate reporting gap, not circularity, is that the per-checkpoint scale alignment behind the Fig. 5 MAE values is not documented, and the supplementary material concedes that early loss improvement is mostly scale adjustment; this is a correctness and evidence concern, not a by-construction reduction.
Assumptions & free parameters
free parameters (3)
- Monocular-to-metric depth scale factor =
Not documented for MVSEC; DSEC uses a median-ratio 'approx scale' and a test-set grid-searched 'best scale'
- Geometry consistency loss weight lambda =
0.05
- Yaw-control gains (lambda_goal, lambda_avoid, alpha, sigma) =
0.2, 1.0, 0.5, 12.0
assumptions (3)
- domain assumption Static scene with no occlusion or disocclusion; optical flow follows purely from depth and ego-motion.
- domain assumption Events are triggered by scene or camera motion, and a correct flow warp aligns events from the same edge into a high-contrast image of warped events.
- domain assumption Depth is scale-free, but the scale is consistent enough across consecutive predictions for the geometry-consistency loss and for metric evaluation.
Cite this review
Pith. "Pith review of On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events." pith.science (2026). https://pith.science/paper/KS3TJBXH
@misc{pith2026241206359,
author = {Pith},
title = {Pith review of: On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events},
year = {2026},
howpublished = {\url{https://pith.science/paper/KS3TJBXH}},
note = {Machine review of arXiv:2412.06359}
}
read the original abstract
Event cameras provide low-latency perception for only milliwatts of power. This makes them highly suitable for resource-restricted, agile robots such as small flying drones. Self-supervised learning based on contrast maximization holds great potential for event-based robot vision, as it foregoes the need for high-frequency ground truth and allows for online learning in the robot's operational environment. However, online, on-board learning raises the major challenge of achieving sufficient computational efficiency for real-time learning, while maintaining competitive visual perception performance. In this work, we improve the time and memory efficiency of the contrast maximization pipeline, making on-device learning of low-latency monocular depth possible. We demonstrate that online learning on board a small drone yields more accurate depth estimates and more successful obstacle avoidance behavior compared to only pre-training. Benchmarking experiments show that the proposed pipeline is not only efficient, but also achieves state-of-the-art depth estimation performance among self-supervised approaches. Our work taps into the unused potential of online, on-device robot learning, promising smaller reality gaps and better performance.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues
Harris eigenvalues and spatiotemporal density values from event cameras encode motion direction and, when added to an optical flow network, improve accuracy in data-scarce settings.
Reference graph
Works this paper leans on
-
[28]
Federico Paredes-Vallés, Kirk Y . W. Scheper, Christophe De Wagter, and Guido C. H. E. de Croon. Taming Contrast Maximization for Learning Sequential, Low-latency, Event- based Optical Flow. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 9695–9705,
-
[2]
Unsuper- vised Scale-consistent Depth and Ego-motion Learning from Monocular Video
Jiawang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid. Unsuper- vised Scale-consistent Depth and Ego-motion Learning from Monocular Video. In Advances in Neural Information Pro- cessing Systems. Curran Associates, Inc., 2019. 2, 3
work page 2019
-
[40]
Yilun Wu, Federico Paredes-Vallés, and Guido C. H. E. de Croon. Lightweight Event-based Optical Flow Estimation via Iterative Deblurring. In 2024 IEEE International Con- ference on Robotics and Automation (ICRA) , pages 14708– 14715, 2024. 3
work page 2024
-
[1]
Monocular Event-Based Vision for Dodging Static Obstacles with a Quadrotor
Anish Bhattacharya, Marco Cannici, Nishanth Rao, Yuezhan Tao, Vijay Kumar, Nikolai Matni, and Davide Scaramuzza. Monocular Event-Based Vision for Dodging Static Obstacles with a Quadrotor. In8th Annual Conference on Robot Learn- ing, 2024. 3
work page 2024
-
[3]
Levi Burner, Anton Mitrokhin, Cornelia Fermüller, and Yiannis Aloimonos. EVIMO2: An Event Camera Dataset for Motion Segmentation, Optical Flow, Structure from Motion, and Visual Inertial Odometry in Indoor Scenes with Monoc- ular or Stereo Algorithms. arXiv:2205.03467 [cs], 2022. 2
arXiv 2022
-
[4]
CNN- based single image obstacle avoidance on a quadrotor
Punarjay Chakravarty, Klaas Kelchtermans, Tom Roussel, Stijn Wellens, Tinne Tuytelaars, and Luc Van Eycken. CNN- based single image obstacle avoidance on a quadrotor. In 2017 IEEE International Conference on Robotics and Au- tomation (ICRA), pages 6369–6374, 2017. 5
work page 2017
-
[5]
Ani Hsieh, Christopher Korpela, Vijay Kumar, Camillo J
Kenneth Chaney, Fernando Cladera, Ziyun Wang, Anthony Bisulco, M. Ani Hsieh, Christopher Korpela, Vijay Kumar, Camillo J. Taylor, and Kostas Daniilidis. M3ED: Multi- Robot, Multi-Sensor, Multi-Environment Event Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 4016–4023, 2023. 1
work page 2023
-
[6]
Yuhua Chen, Cordelia Schmid, and Cristian Sminchisescu. Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7063–7072, 2019. 2
work page 2019
Show all 50 references
-
[7]
Selection and Cross Sim- ilarity for Event-Image Deep Stereo
Hoonhee Cho and Kuk-Jin Yoon. Selection and Cross Sim- ilarity for Event-Image Deep Stereo. In Computer Vision – ECCV 2022, pages 470–486. Springer, Cham, 2022. 6
2022
-
[8]
Tempo- ral Event Stereo via Joint Learning with Stereoscopic Flow
Hoonhee Cho, Jae-Young Kang, and Kuk-Jin Yoon. Tempo- ral Event Stereo via Joint Learning with Stereoscopic Flow. In Computer Vision – ECCV 2024, pages 294–314. Springer, Cham, 2025. 5, 6, 3
2024
-
[9]
Are We Ready for Au- tonomous Drone Racing? The UZH-FPV Drone Racing Dataset
Jeffrey Delmerico, Titus Cieslewski, Henri Rebecq, Matthias Faessler, and Davide Scaramuzza. Are We Ready for Au- tonomous Drone Racing? The UZH-FPV Drone Racing Dataset. In 2019 International Conference on Robotics and Automation (ICRA), pages 6713–6719, 2019. 4, 7
2019
-
[10]
Dy- namic obstacle avoidance for quadrotors with event cameras
Davide Falanga, Kevin Kleber, and Davide Scaramuzza. Dy- namic obstacle avoidance for quadrotors with event cameras. Science Robotics, 5:eaaz9712, 2020. 8
2020
-
[11]
DecTrain: Deciding When to Train a DNN Online,
Zih-Sing Fu, Soumya Sudhakar, Sertac Karaman, and Vivi- enne Sze. DecTrain: Deciding When to Train a DNN Online,
-
[12]
A Unifying Contrast Maximization Framework for Event Cameras, With Applications to Motion, Depth, and Optical Flow Estimation
Guillermo Gallego, Henri Rebecq, and Davide Scaramuzza. A Unifying Contrast Maximization Framework for Event Cameras, With Applications to Motion, Depth, and Optical Flow Estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 3867–...
2018
-
[13]
Focus Is All You Need: Loss Functions for Event- Based Vision
Guillermo Gallego, Mathias Gehrig, and Davide Scara- muzza. Focus Is All You Need: Loss Functions for Event- Based Vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12280– 12289, 2019. 2, 3
2019
-
[14]
Event-based Vision: A Survey
Guillermo Gallego, Tobi Delbruck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew Davison, Joerg Conradt, Kostas Daniilidis, and Da- vide Scaramuzza. Event-based Vision: A Survey. IEEE Transactions on Pattern Analysis and Machine Intell...
2020
-
[15]
DSEC: A Stereo Event Camera Dataset for Driving Scenarios
Mathias Gehrig, Willem Aarents, Daniel Gehrig, and Da- vide Scaramuzza. DSEC: A Stereo Event Camera Dataset for Driving Scenarios. IEEE Robotics and Automation Let- ters, 6:4947–4954, 2021. 1, 5, 6
2021
-
[16]
Out of the Room: Generalizing Event-Based Dynamic Motion Seg- mentation for Complex Scenes
Stamatios Georgoulis, Weining Ren, Alfredo Bochicchio, Daniel Eckert, Yuanyou Li, and Abel Gawel. Out of the Room: Generalizing Event-Based Dynamic Motion Seg- mentation for Complex Scenes. In 2024 International Con- ference on 3D Vision (3DV), pages 442–452, 2024. 2
2024
-
[17]
Clement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. Digging Into Self-Supervised Monocu- lar Depth Estimation. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 3828–3838,
-
[18]
Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown Cameras
Ariel Gordon, Hanhan Li, Rico Jonschkowski, and Anelia Angelova. Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown Cameras. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8977–8986, 2019. 2, 3
2019
-
[19]
Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural Networks
Jesse Hagenaars, Federico Paredes-Vallés, and Guido de Croon. Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural Networks. In Advances in Neu- ral Information Processing Systems, 2021. 2, 3, 7, 8
2021
-
[20]
Motion- prior Contrast Maximization for Dense Continuous-Time Motion Estimation, 2024
Friedhelm Hamann, Ziyun Wang, Ioannis Asmanis, Kenneth Chaney, Guillermo Gallego, and Kostas Daniilidis. Motion- prior Contrast Maximization for Dense Continuous-Time Motion Estimation, 2024. 2
2024
-
[21]
Self-supervised monocular distance learn- ing on a lightweight micro air vehicle
Kevin Lamers, Sjoerd Tijmons, Christophe De Wagter, and Guido de Croon. Self-supervised monocular distance learn- ing on a lightweight micro air vehicle. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1779–1784, 2016. 3
2016
-
[22]
Unsupervised Monocular Depth Learning in Dynamic Scenes
Hanhan Li, Ariel Gordon, Hang Zhao, Vincent Casser, and Anelia Angelova. Unsupervised Monocular Depth Learning in Dynamic Scenes. In Proceedings of the 2020 Conference on Robot Learning, pages 1908–1917. PMLR, 2021. 2
2020
-
[23]
UFO Depth: Unsupervised learning with flow-based odom- etry optimization for metric depth estimation
Vlad Lic ˘aret, Victor Robu, Alina Marcu, Drago¸ s Costea, Emil Slu¸ sanschi, Rahul Sukthankar, and Marius Leordeanu. UFO Depth: Unsupervised learning with flow-based odom- etry optimization for metric depth estimation. In 2022 In- ternational Conference on Robotics and Automa...
2022
-
[24]
Nano Quadcopter Obstacle Avoidance with a Lightweight Monocular Depth Network
Cheng Liu, Yingfu Xu, Erik-Jan van Kampen, and Guido de Croon. Nano Quadcopter Obstacle Avoidance with a Lightweight Monocular Depth Network. IFAC- PapersOnLine, 56:9312–9317, 2023. 3, 5 9
2023
-
[25]
Robot Operating Sys- tem 2: Design, architecture, and uses in the wild
Steven Macenski, Tully Foote, Brian Gerkey, Chris Lalancette, and William Woodall. Robot Operating Sys- tem 2: Design, architecture, and uses in the wild. Science Robotics, 7:eabm6074, 2022. 5, 1
2022
-
[26]
Splatting-Based Synthesis for Video Frame Interpolation
Simon Niklaus, Ping Hu, and Jiawen Chen. Splatting-Based Synthesis for Video Frame Interpolation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, pages 713–723, 2023. 8
2023
-
[27]
Federico Paredes-Vallés and Guido C. H. E. de Croon. Back to Event Basics: Self-Supervised Learning of Image Recon- struction for Event Cameras via Photometric Constancy. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 3446–3455, ...
2021
-
[29]
Paredes-Vallés, J
F. Paredes-Vallés, J. J. Hagenaars, J. Dupeyroux, S. Stroobants, Y . Xu, and G. C. H. E. de Croon. Fully neu- romorphic vision and control for autonomous drone flight. Science Robotics, 9:eadi0591, 2024. 2
2024
-
[30]
Depth Distillation: Unsupervised Metric Depth Estimation for UA Vs by Finding Consensus Between Kinematics, Optical Flow and Deep Learning
Mihai Pirvu, Victor Robu, Vlad Licaret, Dragos Costea, Alina Marcu, Emil Slusanschi, Rahul Sukthankar, and Mar- ius Leordeanu. Depth Distillation: Unsupervised Metric Depth Estimation for UA Vs by Finding Consensus Between Kinematics, Optical Flow and Deep Learning. In Proceed...
2021
-
[31]
Anurag Ranjan, Varun Jampani, Lukas Balles, Kihwan Kim, Deqing Sun, Jonas Wulff, and Michael J. Black. Compet- itive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation. In Proceedings of the IEEE/CVF Conference on Computer ...
2019
-
[32]
Secrets of Event-based Optical Flow, Depth and Ego-motion Estimation by Contrast Maximiza- tion
Shintaro Shiba, Yannick Klose, Yoshimitsu Aoki, and Guillermo Gallego. Secrets of Event-based Optical Flow, Depth and Ego-motion Estimation by Contrast Maximiza- tion. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–18, 2024. 2, 8
2024
-
[33]
Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes
Yihong Sun and Bharath Hariharan. Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes. In Thirty-Seventh Conference on Neural Information Process- ing Systems, 2023. 2
2023
-
[34]
Persistent self-supervised learning: From stereo to monocular vision for obstacle avoidance
Kevin van Hecke, Guido de Croon, Laurens van der Maaten, Daniel Hennes, and Dario Izzo. Persistent self-supervised learning: From stereo to monocular vision for obstacle avoidance. International Journal of Micro Air Vehicles, 10: 186–206, 2018. 2
2018
-
[35]
SfM- Net: Learning of Structure and Motion from Video, 2017
Sudheendra Vijayanarasimhan, Susanna Ricco, Cordelia Schmid, Rahul Sukthankar, and Katerina Fragkiadaki. SfM- Net: Learning of Structure and Motion from Video, 2017. 2
2017
-
[36]
CoVIO: Online Continual Learning for Visual-Inertial Odometry
Niclas Vödisch, Daniele Cattaneo, Wolfram Burgard, and Abhinav Valada. CoVIO: Online Continual Learning for Visual-Inertial Odometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2464–2473, 2023. 8
2023
-
[37]
CoDEPS: Online Continual Learning for Depth Estimation and Panoptic Segmentation
Niclas Vödisch, Kürsat Petek, Wolfram Burgard, and Abhi- nav Valada. CoDEPS: Online Continual Learning for Depth Estimation and Panoptic Segmentation. InRobotics: Science and Systems XIX, 2023. 2
2023
-
[38]
Learning Depth From Monocular Videos Us- ing Direct Methods
Chaoyang Wang, José Miguel Buenaposada, Rui Zhu, and Simon Lucey. Learning Depth From Monocular Videos Us- ing Direct Methods. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 2022– 2030, 2018. 3
2022
-
[39]
Pizer, and Jan-Michael Frahm
Rui Wang, Stephen M. Pizer, and Jan-Michael Frahm. Re- current Neural Network for (Un-)Supervised Learning of Monocular Video Visual Odometry and Depth. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5555–5564, 2019. 2
2019
-
[41]
C. Ye, A. Mitrokhin, C. Fermüller, J. A. Yorke, and Y . Aloimonos. Unsupervised Learning of Dense Optical Flow, Depth and Egomotion with Event-Based Sensors. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5831–5838, 2020. 2
2020
-
[42]
GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
Zhichao Yin and Jianping Shi. GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1983–1992, 2018. 2, 8
1983
-
[43]
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. Unsupervised Learning of Depth and Ego-Motion from Video. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6612–6619, 2017. 2, 3
2017
-
[44]
The Mul- tivehicle Stereo Event Camera Dataset: An Event Camera Dataset for 3D Perception
Alex Zihao Zhu, Dinesh Thakur, Tolga Özaslan, Bernd Pfrommer, Vijay Kumar, and Kostas Daniilidis. The Mul- tivehicle Stereo Event Camera Dataset: An Event Camera Dataset for 3D Perception. IEEE Robotics and Automation Letters, 3:2032–2039, 2018. 1, 2, 4
2018
-
[45]
Unsupervised Event-Based Learning of Optical Flow, Depth, and Egomotion
Alex Zihao Zhu, Liangzhe Yuan, Kenneth Chaney, and Kostas Daniilidis. Unsupervised Event-Based Learning of Optical Flow, Depth, and Egomotion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 989–997, 2019. 2, 5
2019
-
[46]
Self-Supervised Event- Based Monocular Depth Estimation Using Cross-Modal Consistency
Junyu Zhu, Lina Liu, Bofeng Jiang, Feng Wen, Hongbo Zhang, Wanlong Li, and Yong Liu. Self-Supervised Event- Based Monocular Depth Estimation Using Cross-Modal Consistency. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7704–7710,
2023
-
[48]
step 0” (only pre-training) to “step 3k
Extra qualitative results DSEC. We present additional qualitative results from the DSEC test set in Fig. 8. While the self-supervised results exhibit less sharp boundaries, they are free from artifacts commonly introduced by supervised learning, such as dis- continuities at im...
-
[49]
For offline training on datasets, we train for 50 epochs with the Adam optimizer and a learning rate of 1e-
Implementation details Training. For offline training on datasets, we train for 50 epochs with the Adam optimizer and a learning rate of 1e-
-
[50]
bilinear
For contrast maximization, we accumulate 10 bins of events, and warp all events to all bin edges. Furthermore, we set the weight for the geometric consistency loss λ = 0.05. For on-device learning, we lower the learning rate to 1e-5. Specifics per dataset are mentioned below. ...
1980
-
[2023]
2, 5 10 On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events Supplementary Material https://mavlab.tudelft.nl/depth_from_events
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.