REVIEW 3 cited by
NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural implicit representations have recently become popular in simultaneous localization and mapping (SLAM), especially in dense visual SLAM. However, previous works in this direction either rely on RGB-D sensors, or require a separate monocular SLAM approach for camera tracking and do not produce high-fidelity dense 3D scene reconstruction. In this paper, we present NICER-SLAM, a dense RGB SLAM system that simultaneously optimizes for camera poses and a hierarchical neural implicit map representation, which also allows for high-quality novel view synthesis. To facilitate the optimization process for mapping, we integrate additional supervision signals including easy-to-obtain monocular geometric cues and optical flow, and also introduce a simple warping loss to further enforce geometry consistency. Moreover, to further boost performance in complicated indoor scenes, we also propose a local adaptive transformation from signed distance functions (SDFs) to density in the volume rendering equation. On both synthetic and real-world datasets we demonstrate strong performance in dense mapping, tracking, and novel view synthesis, even competitive with recent RGB-D SLAM systems.
Forward citations
Cited by 3 Pith papers
-
Princeton365: A Diverse Dataset with Accurate Camera Pose
Princeton365 is a 365-video SLAM/NVS benchmark with board-calibrated millimeter-accurate 6-DoF poses, a new scale-aware optical-flow error metric, and an NVS benchmark of fully non-Lambertian 360-degree scans.
-
Enhanced Velocity Field Modeling for Gaussian Video Reconstruction
Velocity field rendering with flow-based losses and flow-assisted densification lifts dynamic Gaussian novel-view PSNR by about 2.5 dB on Nvidia-long and Neu3D.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
Discussion (0). Continue with ORCID to comment.