Pith. sign in

REVIEW 2 major objections 4 minor 240 references

Object Learning and Robust 3D Reconstruction

T0 review · 2 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Unsupervised object representations—learned from motion in 2D and geometric consistency in 3D—can segment objects of interest and ignore dynamic distractors, yielding cleaner 3D reconstruction from casual captures.

desk verdict A solid thesis compiling three peer-reviewed papers; the robust-3D methods are genuinely useful, but the robustness claim has a hidden high-occupancy ceiling that the thesis never acknowledges. read the letter →

arxiv 2504.17812 v1 pith:WCNGDDZC submitted 2025-04-22 cs.CV eess.IV

classification cs.CVeess.IV
keywords unsupervisedobjectsegmentationcapsulenetworksself-supervisedmotionlearningrobustestimationneuralradiancefields3DGaussiansplattingdistractorremovalcasualcapturereconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis argues that explicit object representations learned without supervision—by motion in 2D and geometric consistency in 3D—are the key to robust computer vision. It shows that a capsule network can parse a single image into movable parts using video pairs as training data, and that the same parts can be used for unsupervised classification and segmentation. It then shows that treating transient objects as outliers in a robust optimization makes neural radiance fields and 3D Gaussian splatting reconstruct clean static scenes from casual, distractor-filled captures. If these claims hold, 3D reconstruction from everyday photographs no longer requires clean data, manual masks, or class-specific detectors, and image understanding gains viewpoint invariance and compositional generality.

What carries the argument

The machinery that carries the argument is the object hypothesis itself: a capsule is a part descriptor $c_k = (s_k, \theta_k, d_k)$ holding a canonical shape vector, a pose transformation, and a depth scalar; combined with a neural implicit decoder $D_\omega$, it yields a visible mask $\Lambda_k^+$ that explains a portion of the flow field $\Phi(u)=\sum_k \Lambda_k^+(u)[T_k(u)-u]$. In 3D, the corresponding mechanism is a trimmed estimator with a spatial-coherence prior: residuals $\epsilon(r)$ are thresholded at their median, blurred with a $3\times3$ kernel, and aggregated over $8\times8$ patches to set binary inlier/outlier weights $W(r)$, which are then used in an iteratively reweighted least-squares loss. SpotLessSplats replaces the RGB-residual classifier with a semantic-feature clustering or an MLP $H(F;\theta)$ trained on weak labels derived from lower/upper residual thresholds, and adds utilization-based pruning to stabilize Gaussian splatting. These mechanisms turn 'objectness' into a computable, self-supervising signal.

What would settle it

Capture a scene with a soft shadow cast by a moving person plus a fine static texture of similar residual magnitude; if the robust loss either leaves the shadow in or removes the texture, the spatial-coherence assumption fails.

Watch

Extended reading notes

Core claim

The thesis's central claim is that a network with an explicit object-based representation—parts that carry shape, pose, and depth in 2D; photometrically consistent regions in 3D—can learn to segment scenes without labels and thereby become more robust. In FlowCapsules, motion between video frames serves as the training signal: an encoder parses a single image into primary capsules, and a decoder renders their shapes; the capsules' poses and visibility masks produce a flow field that warps one frame toward the next, and optimizing this flow teaches the network which image regions are movable objects. In RobustNeRF, distractors are treated as outliers in a trimmed least-squares objective: pixels with high residual, spatially smoothed and patch-aggregated, are excluded from the NeRF loss. In SpotLessSplats, the same idea is transferred to 3D Gaussian splatting, but outlier detection uses semantic features from a text-to-image diffusion model rather than raw color, so distractors that share the background color are still masked.

Load-bearing premise

The load-bearing premise is that the objects worth keeping are exactly the regions that move together in 2D and look inconsistent across views in 3D, so motion and residual-based trimming can separate them from static background detail without supervision.

Editorial extensions

If this is right

  • Casual captures—with pedestrians, pets, or moving shadows—can be reconstructed into clean 3D models without per-image labeling, because the optimizer itself learns which pixels to ignore.
  • A single robust-loss recipe transfers from NeRF-style volumetric rendering to Gaussian splatting, so distractor handling need not be re-engineered for each new scene representation.
  • Unsupervised part representations give viewpoint invariance and shape completion under occlusion, improving classification and segmentation on cluttered data.
  • Reconstruction quality degrades gracefully as the fraction of cluttered training images grows: on one scene, RobustNeRF stays above 31 dB PSNR while the base mip-NeRF 360 drops from 33 to 25 dB.
  • Utilization-based pruning removes floaters and cuts the number of Gaussians by a factor of 2 to 4.5 with little quality loss, even on clean scenes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's demonstrations, the learned-outlier view could be pointed at transient phenomena the thesis does not test—specular highlights, rain, lens flare—as long as they are spatially coherent and photometrically inconsistent.
  • A testable extension is to use the FlowCapsules motion cue to generate pseudo-masks for video frames and feed those masks to a 3D reconstructor, closing the loop between 2D object learning and 3D robustness without any labels.
  • Because the thesis shows residual trimming is statistically inefficient on clean data, one could switch between robust and non-robust losses based on an online estimate of the clutter fraction, preserving quality when no distractors are present.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This PhD thesis bundles three published projects (FlowCapsules, RobustNeRF, SpotLessSplats) under the thesis that unsupervised, object-based representations improve robustness in image understanding and 3D reconstruction. Chapter 3 learns part capsules from motion via a self-supervised flow-rendering loss; Chapter 4 removes transient distractors in NeRF training through a trimmed, smoothed, patch-aggregated robust loss; Chapter 5 adapts this idea to 3D Gaussian Splatting using semantic features from a pre-trained diffusion model, with clustering or an MLP for outlier masks, plus utilization-based pruning. The empirical evaluation covers synthetic and real-world datasets (Geo, Exercise, Kubric, D2NeRF, RobustNeRF's own captures, NeRF On-the-go) with held-out clean test views, and consistently reports gains over mip-NeRF 360, D2NeRF, NeRF On-the-go, NeRF-HuGS, and vanilla 3DGS.

Significance. If the results hold, the thesis makes a solid contribution: RobustNeRF and SpotLessSplats provide practical, unsupervised ways to handle transient distractors in casual captures, and FlowCapsules demonstrates motion-supervised part discovery with shape completion under occlusion. The empirical work is extensive and mostly well-controlled, including paired distractor/distractor-free captures and held-out test views, and each chapter is an already-peer-reviewed publication. The methods are described in enough detail to reimplement, and several limitations (small distractors, soft shadows, semantically similar instances) are acknowledged explicitly. The reuse of published material makes the thesis a coherent anthology rather than a single new derivation, but its central claims are empirical and evaluated on external benchmarks as well as new datasets.

major comments (2)
  1. [Sec. 4.4.5 / Fig. 4.17] The central robustness claim of Chapter 4 has a hidden outlier-occupancy ceiling that is visible in the paper's own ablation. The trimming rule (4.8) uses a per-image median residual, the patch rule (4.10) requires at least 60% of a 16x16 neighborhood to be inliers, and the text around Fig. 4.17 states that in the hard Kubric setting (44% outlier pixels) any TR above 50% gives worse results. With the default TR=0.6, the method is already past its operating point on that scene, and the reconstruction degrades substantially. The limitation paragraphs in Sec. 4.5 mention small distractors and statistical inefficiency but do not acknowledge this high-occupancy ceiling, which is directly load-bearing for the claim of robust 3D modelling in casual captures. I recommend either adding a precise condition on distractor pixel fraction, or reporting a robustness curve (as in Fig. 4.17) for the natural scenes and stating the operating range explicitly in the conclusions.
  2. [Sec. 5.4.1, Eqs. (5.2)-(5.4) and Fig. 5.8] The SpotLessSplats masking pipeline inherits the same median-based ceiling from RobustNeRF. The weak labels U and L in Eqs. (5.8)-(5.9) are quantile masks from Eq. (5.2), which relies on a per-image/global median of residual magnitudes. If more than half the pixels in the histogram are distractors, the median falls inside the distractor residual distribution and the bootstrap labels are contaminated; the semantic clustering or MLP can then propagate the error rather than correct it. The thesis reports results on NeRF On-the-go 'high' occlusion scenes, but it does not quantify the outlier pixel fraction in those scenes, nor does it test SLS on the 44%-occupancy hard Kubric setup. Please add an explicit statement of the required inlier-majority condition, and include an experiment with controlled outlier fractions (e.g., the Kubric easy/medium/hard settings) for both SLS-agg and SLS-mlp.
minor comments (4)
  1. [Sec. 3.5.4, Table 3.3] The row 'No 6-Layer' in the occlusion-inductive-bias ablation is ambiguous: 'No' appears to mean 'without depth ordering', but the column header does not say so, and the reader must infer this from the surrounding text. Please rename the row to 'No depth ordering'.
  2. [Sec. 5.3] There is a typo '3DSG' in the sentence 'renders the entire image in a single forward pass'; it should read '3DGS'.
  3. [Abstract and Sec. 1.1] The phrase 'detecting and removal of the object of interest from the input images' in the abstract is confusing, because the 3D chapters remove distractors, not objects of interest. Please rephrase to 'detecting and removing transient distractors'.
  4. [Sec. 4.4.2 / Fig. 4.14] The D2NeRF hyperparameter tuning is reported as 'Config 1' being best, but the tuning is only done on two of the four datasets (Statue and Crab). This is acknowledged in the text, but it would help to state explicitly that the comparison to D2NeRF is therefore not a fully tuned baseline on all scenes.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: central derivations are self-contained and validated on held-out benchmarks.

full rationale

The thesis's three projects are each evaluated against external ground truth or held-out clean views, so their central claims do not reduce to their inputs by construction. FlowCapsules learns part masks from a motion-warping proxy loss (Eqs. 3.5-3.7) and is scored by IoU against ground-truth masks on Geo/Exercise (Table 3.1); motion is an inductive bias, not a fitted label. RobustNeRF's trimming weights (Eqs. 4.8-4.10) are IRLS weights computed from current residuals, and the method is scored on distractor-free held-out frames (Fig. 4.11), with hyperparameters hand-chosen rather than fitted to the test set. SpotLessSplats' MLP masks (Eqs. 5.5-5.9) are weakly supervised by residual-derived labels, but the final reconstruction is evaluated on clean held-out views (Figs. 5.3, 5.5, 5.6), and ablations show that the semantic-feature component contributes over the residual-only RobustFilter. The frequent self-citations (e.g., Sabour et al. 2023 as the starting point for Chapter 5) are explanatory and are backed by experiments reported within the thesis itself; none functions as an unverified uniqueness theorem or imported ansatz. The acknowledged limitations—high-occupancy distractors in Fig. 4.17, small distractors in Sec. 4.5, and semantically similar instances/soft shadows in Sec. 5.5.4—are empirical boundary conditions, not circular reasoning.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central claims are empirical and depend on several hand-chosen hyperparameters and domain assumptions. The hyperparameters are not fitted to a specific target result; they are constants used across datasets, with sensitivity analyses. The axioms are standard in the field, though the distractor-coherence and feature-space assumptions are specific to the proposed methods and are explicitly discussed as limitations.

free parameters (7)
  • RobustNeRF trim threshold TR = 0.6
    Used in Eq. (4.10) to label 8x8 patches as inliers or outliers; sensitivity analysis shows performance depends on TR relative to outlier proportion.
  • RobustNeRF patch and neighborhood size = 16x16 patch, 16 neighborhood
    Used in Eq. (4.10); larger neighborhoods act as stronger regularizers, bounded by memory.
  • SpotLessSplats pruning threshold kappa = 1e-8
    Utilization-based pruning in Eq. (5.12) prunes Gaussians with utilization below this threshold.
  • SpotLessSplats warm-up scheduler params = beta1=3e-4, beta2=1.5
    Controls the staircase exponential decay in Eq. (5.11) for mask warm-up; varied for high-occlusion scenes.
  • SpotLessSplats clustering count C = 100
    Number of agglomerative clusters in Eq. (5.3); performance plateaus for C>=100.
  • SpotLessSplats feature config = Stable Diffusion v2.1 layer 2, time step 261, empty prompt
    Feature extractor chosen from Tang et al.; affects mask quality.
  • FlowCapsules capsule count K and encoding C = K=8, C=32 (Geo); K=16, C=16 (Exercise)
    Ablations show IoU is robust to these choices, but they affect part granularity.
assumptions (5)
  • standard math Backpropagation and stochastic gradient descent can train the proposed networks.
    All methods rely on neural network optimization; this is standard in deep learning.
  • domain assumption COLMAP provides sufficiently accurate camera poses for 3D reconstruction.
    Used to train NeRF and 3DGS models; slight pose errors affect apartment scenes, as noted in Section 4.4.4.
  • domain assumption Motion is a valid cue for object definition, based on Gestalt psychology and infant studies.
    Foundational for FlowCapsules; if this assumption fails, learned part decomposition loses its meaning.
  • ad hoc to paper Distractors appear as spatially coherent photometric outliers in casual captures.
    Core to the robust trimming in RobustNeRF and SpotLessSplats; the thesis admits limitations for small and shadow distractors.
  • ad hoc to paper Pre-trained Stable Diffusion features encode semantic structure sufficient to distinguish distractors from static objects.
    Load-bearing for SpotLessSplats; fails when similar-class instances are near each other (Section 5.5.4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Object Learning and Robust 3D Reconstruction." pith.science (2026). https://pith.science/paper/WCNGDDZC

@misc{pith2026250417812,
  author       = {Pith},
  title        = {Pith review of: Object Learning and Robust 3D Reconstruction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WCNGDDZC}},
  note         = {Machine review of arXiv:2504.17812}
}
read the original abstract

In this thesis we discuss architectural designs and training methods for a neural network to have the ability of dissecting an image into objects of interest without supervision. The main challenge in 2D unsupervised object segmentation is distinguishing between foreground objects of interest and background. FlowCapsules uses motion as a cue for the objects of interest in 2D scenarios. The last part of this thesis focuses on 3D applications where the goal is detecting and removal of the object of interest from the input images. In these tasks, we leverage the geometric consistency of scenes in 3D to detect the inconsistent dynamic objects. Our transient object masks are then used for designing robust optimization kernels to improve 3D modelling in a casual capture setup. One of our goals in this thesis is to show the merits of unsupervised object based approaches in computer vision. Furthermore, we suggest possible directions for defining objects of interest or foreground objects without requiring supervision. Our hope is to motivate and excite the community into further exploring explicit object representations in image understanding tasks.

Figures

Figures reproduced from arXiv: 2504.17812 by the authors.

Figure 4
Figure 4. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 4
Figure 4. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 1.1
Figure 1.1. Examples of human visual perception properties based on gestalt psychology. [PITH_FULL_IMAGE:figures/full_fig_p014_1_1.png] view at source ↗
Figures from the paper (65 more)
Figure 1.2
Figure 1.2. Figure 1.2: Deep Learning has different principals for image understanding than humans. [PITH_FULL_IMAGE:figures/full_fig_p015_1_2.png]
Figure 3.1
Figure 3.1. Figure 3.1: Self-supervised training for learning primary capsules: An image encoder is trained to decompose the scene into a collection of primary capsules. Learning is accomplished in an unsupervised manner, using flow estimation from capsule shapes and poses as a proxy task. …
Figure 3.2
Figure 3.2. Figure 3.2: Inference architecture. (left) The encoder Eω parses an image into part capsules, each comprising a shape vector sk, a pose θk, and a scalar depth value dk. (right) The shape decoder Dω is an implicit function. It takes as input a shape vector, sk, and a location in …
Figure 3.3
Figure 3.3. Figure 3.3: Encoder architecture. The encoder comprises convolution layers with ReLU activation, followed by down-sampling via 2×2 AveragePooling. Following the last convolution layer is a tanh fully connected layer, and a fully connected layer grouped into K, C-dimensional caps…
Figure 3.4
Figure 3.4. Figure 3.4: Self-supervised training – Training uses a proxy motion task in which the capsule encoder is applied to a pair of successive video frames, providing K primary capsule encodings from each frame. Visible part masks, Λ+ k , and their corresponding poses, Pθ, determine a…
Figure 3.5
Figure 3.5. Figure 3.5: Decoder architecture. A neural implicit function (Chen and Zhang, 2019) is used to represent part masks. An MLP with SELU activations (Klambauer et al., 2017) takes as input a shape vector s and a pixel position u. Applied to a pixel grid, it produces a logit grid fo…
Figure 3.6
Figure 3.6. Figure 3.6: Estimated flows and predicted next frames on training data from [PITH_FULL_IMAGE:figures/full_fig_p031_3_6.png]
Figure 3.7
Figure 3.7. Figure 3.7: Estimated flows and predicted frames on randomly selected images from the [PITH_FULL_IMAGE:figures/full_fig_p032_3_7.png]
Figure 3.8
Figure 3.8. Figure 3.8: Inferred FlowCapsule shapes and corresponding visibility masks on Geo (rows 1–3), and Geo [PITH_FULL_IMAGE:figures/full_fig_p033_3_8.png]
Figure 3.9
Figure 3.9. Figure 3.9: The ground truth segment masks along with sample FlowCapsule masks Λ [PITH_FULL_IMAGE:figures/full_fig_p034_3_9.png]
Figure 3.10
Figure 3.10. Figure 3.10: (left) SCAE reconstructions after training on Geo and Geo [PITH_FULL_IMAGE:figures/full_fig_p035_3_10.png]
Figure 3.11
Figure 3.11. Figure 3.11: The ground truth segment masks along with sample FlowCapsule masks Λ [PITH_FULL_IMAGE:figures/full_fig_p037_3_11.png]
Figure 4.1
Figure 4.1. Figure 4.1: NeRF assumes photometric consistency in the observed images of a scene. Violations of this [PITH_FULL_IMAGE:figures/full_fig_p041_4_1.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p041_5.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p042_5.png]
Figure 4.2
Figure 4.2. Figure 4.2: Ambiguity – A simple 2D scene where a static object (blue) is captured by three cameras. During the first and third capture the scene is not photo-consistent as a distractor was within the field of view. Not photo-consistent portions of the scene can end up being enc…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p044_4.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p044_5.png]
Figure 4.3
Figure 4.3. Figure 4.3: Histograms – Robust estimators perform well when the distribution of residuals agrees with the one implied by the estimator (e.g., Gaussian for L2, Laplacian for L1). Here we visualize the ground-truth distribution of residuals (bottom-left), which is hardly a good m…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p045_4.png]
Figure 4.4
Figure 4.4. Figure 4.4: Kernels – (top-left) Family of robust kernels (Barron, 2019), including L2 (α=2), Charbon￾nier (α=1) and Geman-McClure (α=−2). (top-right) Mid-training, residual magnitudes are similar for distractors and fine-grained details, and pixels with large residuals are lear…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p046_4.png]
Figure 4.5
Figure 4.5. Figure 4.5: Algorithm – We visualize our weight function computed by residuals on two examples: (top) the residuals of a (mid-training) NeRF rendered from a training viewpoint, (bottom) a toy residual image containing residual of small spatial extent (dot, line) and residuals of…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p047_4.png]
Figure 4.6
Figure 4.6. Figure 4.6: Residuals – For the dataset shown in the top row, we visualize the dynamics of the Ro￾bustNeRF training residuals, which show how over time the estimated distractor weights go from being random ((t/T)=0.5%) to identify distractor pixels ((t/T)=100%) without any expli…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p048_4.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p049_4.png]
Figure 4.7
Figure 4.7. Figure 4.7: Dataset – Sample training images showing the distractors in each scene. Statue and Android were acquired manually, and the others with a robotic arm. In the robotic setting we have pixel-perfect alignment of distractor vs. distractor-free images. Images are kept in t…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p050_4.png]
Figure 4.8
Figure 4.8. Figure 4.8: Challenges in Apartment Scenes – Each row, from left to right, shows a ground truth photo, a RobustNeRF render, and the difference between the two. Best viewed in PDF. (Top) Note the fold in the table cloth in ground truth image and the lack of fine-grained detail on…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p051_4.png]
Figure 4.9
Figure 4.9. Figure 4.9: Natural Scenes – Key facts about natural scenes introduced in this work. Includes number of paired photos with (# Clut.) and without (# Clean) distractors. Extra photos (# Extra) do not contain distractors and are taken from unpaired camera poses. naturally changes i…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p052_4.png]
Figure 4.10
Figure 4.10. Figure 4.10: Synthetic Kubric Scenes – Example Kubric synthetic images for three datasets with different ratio of outlier pixels. The sofa, lamp, and bookcase are static objects in all three setups. The easy setup has 1 small distractor, the medium setup has 3 medium distractors…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p053_4.png]
Figure 4.11
Figure 4.11. Figure 4.11: Evaluation on Natural Scenes – RobustNeRF outperforms baselines and D2NeRF (Wu et al., 2022a) on novel view synthesis with real-world captures. The table provides a quantitative comparison of RobustNeRF, D2NeRF and mip-NeRF 360 using different reconstruction losses.…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p054_4.png]
Figure 4.12
Figure 4.12. Figure 4.12: Evaluations on D2NeRF Synthetic Scenes – Quantitative and qualitative evaluations on the Kubric synthetic dataset introduced by D2NeRF, consisting of 200 training frames (with distractor) and 100 novel views for evaluation (without distractor). Crab BabyYoda LPIPS↓ …
Figure 4.13
Figure 4.13. Figure 4.13: Effect of Image Order on D2NeRF – As this model is based on space-time NeRFs (Park et al., 2021b), to make it compatible with our setting we create a ’temporal’ indexing of the photos. Here, we visualize: (left) with our heuristic ordering; (right) with another rand…
Figure 4.14
Figure 4.14. Figure 4.14: D2NeRF HParam Tuning – The performance of D2NeRF is heavily influenced by the choice of hyperparameters. In particular, optimal choices of hyperparameters are noted to be strongly influenced by the amount of object and camera motion, as well as video length. We tune…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p056_4.png]
Figure 4.15
Figure 4.15. Figure 4.15: Ablations – Blindly trimming the loss causes details to be lost. Smoothing recovers fine-grained detail, while patch-based evaluation speeds up training and adds more detail. Patching enables the model to reach PSNR of 30, almost 4× faster [PITH_FULL_IMAGE:figures/…
Figure 4.16
Figure 4.16. Figure 4.16: Sensitivity and Limitations – (left) Reconstruction accuracy for BabyYoda as we increase the fraction of train images with distractors. (right) Accuracy vs training time on clean BabyYoda im￾ages (distractor-free). Distribution Sensitivity – [PITH_FULL_IMAGE:figure…
Figure 4.17
Figure 4.17. Figure 4.17: Sensitivity to TR – RobustNeRF’s reconstruction quality as a function of TR on scenes with different inlier/outlier proportions. Overestimating TR increases training time without affecting final reconstruction accuracy. Neigh./Patch 4/2 8/4 16/2 16/4 16/8 TR = 0.6 1…
Figure 4.18
Figure 4.18. Figure 4.18: Sensitivity to hyper-parameters. PSNR on distractor-free frames on the Crab dataset as a function of RobustNeRF’s neighborhood size, patch size, and TR. we observe that after the 250k iterations the model has not converged yet. On average training with 30% of loss r…
Figure 4.19
Figure 4.19. Figure 4.19: Qualitative results on scenes with view-dependent effects. RobustNeRF naturally captures view-dependent effects in scenes with (3rd-6th rows) and without (1st and 2nd row) distractors. three wooden robots with articulated joints as distractors, and even in this setu…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p059_4.png]
Figure 4.20
Figure 4.20. Figure 4.20: Qualitative results on D2NeRF Pick scene. Renders of static model components. Results for NeRF-W and D2NeRF are provided by Wu et al. (2022a). Note how RobustNeRF naturally captures specular reflections and shadows (green, right) [PITH_FULL_IMAGE:figures/full_fig_p…
Figure 4.21
Figure 4.21. Figure 4.21: Statue – Qualitative results on Statue. It is helpful to zoom in to see details [PITH_FULL_IMAGE:figures/full_fig_p061_4_21.png]
Figure 4.22
Figure 4.22. Figure 4.22: Android – Qualitative results on Android. It is helpful to zoom in to see details [PITH_FULL_IMAGE:figures/full_fig_p061_4_22.png]
Figure 4.23
Figure 4.23. Figure 4.23: Crab – Qualitative results on Crab. It is helpful to zoom in to see details. 4.5 Conclusions We address a central problem in training NeRF models, namely, optimization in the presence of distractors, such as transient or moving objects and photometric phenomena that…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p062_4.png]
Figure 4.24
Figure 4.24. Figure 4.24: BabyYoda – Qualitative results on BabyYoda. It is helpful to zoom in to see details. 51 [PITH_FULL_IMAGE:figures/full_fig_p063_4_24.png]
Figure 5.1
Figure 5.1. Figure 5.1: SpotLessSplats cleanly reconstructs a scene with many transient occluders ( [PITH_FULL_IMAGE:figures/full_fig_p065_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Our outlier classification using clustered semantic features covers the distractor balloon fully, but [PITH_FULL_IMAGE:figures/full_fig_p066_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: Our method accurately reconstructs scenes with different levels of transient occlusion, avoiding [PITH_FULL_IMAGE:figures/full_fig_p068_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Lower and upper error residual labels provide a weak supervision for training an MLP classifier [PITH_FULL_IMAGE:figures/full_fig_p071_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Quantitative and qualitative evaluation on RobustNeRF ( [PITH_FULL_IMAGE:figures/full_fig_p073_5_5.png]
Figure 5.6
Figure 5.6. Figure 5.6: SLS reconstructs scenes from NeRF On-the-go ( [PITH_FULL_IMAGE:figures/full_fig_p075_5_6.png]
Figure 5.7
Figure 5.7. Figure 5.7: Quantitative and qualitative results on MipNeRF360 ( [PITH_FULL_IMAGE:figures/full_fig_p076_5_7.png]
Figure 5.8
Figure 5.8. Figure 5.8: We ablate our different robust masking methods on [PITH_FULL_IMAGE:figures/full_fig_p077_5_8.png]
Figure 5.9
Figure 5.9. Figure 5.9: Qualitative results on scenes from NeRF On-the-go ( [PITH_FULL_IMAGE:figures/full_fig_p079_5_9.png]
Figure 5.10
Figure 5.10. Figure 5.10: Ablations on variants from Section 5.4.1 show replacing the MLP eq. (5.5) in SLS-mlp with a CNN reduces quality. Varying its regularization coefficient λ in eq. (5.6) shows minimal impact. More agglomerative clusters in SLS-agg eq. (5.3) improve performance, plateau…
Figure 5.11
Figure 5.11. Figure 5.11: Ablation on adaptations from Section 5.4.2 show disabling UBP (section 5.4.2) may produce higher reconstruction metrics but leaks transients as seen in the lower-left corner of the image; replacing it with “Opacity Reset” as originally introduced in 3DGS is also ine…
Figure 5.12
Figure 5.12. Figure 5.12: SLS-MLP can correctly distinguish between similar-looking oranges when the non-dsitractor [PITH_FULL_IMAGE:figures/full_fig_p082_5_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

240 extracted references · 57 canonical work pages

  1. [1]

    Adamkiewicz, M., Chen, T., Caccavale, A., Gardner, R., Culbertson, P., Bohg, J., and Schwager, M. (2022). Vision-only robot navigation in a neural radiance world. IEEE Robotics and Automation Letters

  2. [2]

    Afshar, P., Mohammadi, A., and Plataniotis, K. N. (2018). Brain tumor type classification via capsule networks. In IEEE International Conference on Image Processing (ICIP) , pages 3129--3133. IEEE

  3. [3]

    and Torresani, L

    Ahmed, K. and Torresani, L. (2019). Star-caps: Capsule networks with straight-through attentive routing. In NeurIPS , pages 9098--9107

  4. [4]

    Amir, S., Gandelsman, Y., Bagon, S., and Dekel, T. (2022). Deep ViT Features as Dense Visual Descriptors . What is Motion For Workshop in ECCV

  5. [5]

    Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu c i \'c , M., and Schmid, C. (2021). Vivit: A video vision transformer. In Proceedings of the IEEE/CVF international conference on computer vision , pages 6836--6846

  6. [6]

    and Lipman, Y

    Atzmon, M. and Lipman, Y. (2020). Sal: Sign agnostic learning of shapes from raw data. In IEEE CVPR , pages 2565--2574

  7. [7]

    Bai, S., Torr, P., et al. (2021). Visual parser: Representing part-whole hierarchies with transformers. arXiv preprint arXiv:2107.05790

  8. [8]

    Baker, N., Lu, H., Erlikhman, G., and Kellman, P. J. (2018). Deep convolutional networks do not classify based on global object shape. PLoS computational biology , 14(12):e1006613

Show all 240 references
  1. [9]

    Bao, Z., Tokmakov, P., Jabri, A., Wang, Y.-X., Gaidon, A., and Hebert, M. (2022). Discovering objects that can move. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 11789--11798

  2. [10]

    Barron, J., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., and Srinivasan, P. (2021a). Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields . In ICCV

  3. [11]

    Barron, J. T. (2019). A general and adaptive robust loss function. CVPR

  4. [12]

    T., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., and Srinivasan, P

    Barron, J. T., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., and Srinivasan, P. P. (2021b). Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. ICCV

  5. [13]

    T., Mildenhall, B., Verbin, D., Srinivasan, P

    Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., and Hedman, P. (2022a). Mip- NeRF 360: Unbounded Anti - Aliased Neural Radiance Fields . In CVPR

  6. [14]

    T., Mildenhall, B., Verbin, D., Srinivasan, P

    Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., and Hedman, P. (2022b). Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In CVPR

  7. [15]

    Basri, R., Galun, M., Geifman, A., Jacobs, D., Kasten, Y., and Kritchman, S. (2020). Frequency bias in neural networks for input of non-uniform density. In ICML , pages 685--694

  8. [16]

    F., Wu, J., Tenenbaum, J., and Yamins, D

    Bear, D., Fan, C., Mrowca, D., Li, Y., Alter, S., Nayebi, A., Schwartz, J., Fei-Fei, L. F., Wu, J., Tenenbaum, J., and Yamins, D. L. (2020). Learning physical graph representations from visual scenes. NeurIPS

  9. [17]

    Biederman, I. (1987). Recognition-by-components: a theory of human image understanding. Psychological review , 94(2):115

  10. [18]

    T., Lensch, H

    Boss, M., Engelhardt, A., Kar, A., Li, Y., Sun, D., Barron, J. T., Lensch, H. P., and Jampani, V. (2022). SAMURAI : S hape A nd M aterial from U nconstrained R eal-world A rbitrary I mage collections. In NeurIPS

  11. [19]

    and Bethge, M

    Brendel, W. and Bethge, M. (2019). Approximating cnns with bag-of-local-features models works surprisingly well on imagenet. arXiv preprint arXiv:1904.00760

  12. [20]

    P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A

    Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A. (2019). Monet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390

  13. [21]

    and Fox, D

    Byravan, A. and Fox, D. (2017). Se3-nets: Learning rigid body motion using deep neural networks. IEEE ICRA

  14. [22]

    Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021). Emerging Properties in Self - Supervised Vision Transformers . In ICCV

  15. [23]

    Chabra, R., Lenssen, J., Ilg, E., Schmidt, T., Straub, J., Lovegrove, S., and Newcombe, R. (2020). Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction . In ECCV

  16. [24]

    X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., Xiao, J., Yi, L., and Yu, F

    Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., Xiao, J., Yi, L., and Yu, F. (2015). ShapeNet : An Information-Rich 3D model repository

  17. [25]

    L., Tagliasacchi, A., and Sitzmann, V

    Charatan, D., Li, S. L., Tagliasacchi, A., and Sitzmann, V. (2024). pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In CVPR

  18. [26]

    B., Yamins, D

    Chen, H., Venkatesh, R., Friedman, Y., Wu, J., Tenenbaum, J. B., Yamins, D. L., and Bear, D. M. (2022a). Unsupervised segmentation in real-world images via spelke object inference. arXiv preprint arXiv:2205.08515

  19. [27]

    Chen, J., Qin, Y., Liu, L., Lu, J., and Li, G. (2024a). Nerf-hugs: Improved neural radiance fields in non-static scenes using heuristics-guided segmentation. CVPR

  20. [28]

    Chen, X., Zhang, Q., Li, X., Chen, Y., Feng, Y., Wang, X., and Wang, J. (2022b). Hallucinated neural radiance fields in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12943--12952

  21. [29]

    Chen, X., Zhang, Q., Li, X., Chen, Y., Feng, Y., Wang, X., and Wang, J. (2022c). Hallucinated Neural Radiance Fields in the Wild . In CVPR

  22. [30]

    Chen, Y., Xu, H., Zheng, C., Zhuang, B., Pollefeys, M., Geiger, A., Cham, T.-J., and Cai, J. (2024b). Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv preprint arXiv:2403.14627

  23. [31]

    Chen, Z., Funkhouser, T., Hedman, P., and Tagliasacchi, A. (2023). Mobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. In CVPR

  24. [32]

    and Zhang, H

    Chen, Z. and Zhang, H. (2019). Learning Implicit Fields for Generative Shape Modeling . In CVPR

  25. [33]

    Chetverikov, D., Svirko, D., Stepanov, D., and Krsek, P. (2002). The trimmed iterative closest point algorithm. In ICPR , volume 3

  26. [34]

    Chibane, J., Bansal, A., Lazova, V., and Pons-Moll, G. (2021). Stereo radiance fields (srf): Learning view synthesis for sparse views of novel scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 7911--7920

  27. [35]

    Chng, S.-F., Ramasinghe, S., Sherrah, J., and Lucey, S. (2022). Garf: Gaussian activated radiance fields for high fidelity reconstruction and pose estimation. arXiv e-prints

  28. [36]

    Dahmani, H., Bennehar, M., Piasco, N., Roldao, L., and Tsishkou, D. (2024). SWAG : Splatting in the Wild images with Appearance -conditioned Gaussians

  29. [37]

    Deng, F., Zhi, Z., Lee, D., and Ahn, S. (2021). Generative scene graph networks. In ICLR

  30. [38]

    Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In IEEE CVPR , pages 248--255

  31. [39]

    Dong, Y., Ruan, S., Su, H., Kang, C., Wei, X., and Zhu, J. (2022). Viewfool: Evaluating the robustness of visual recognition to adversarial viewpoints. Advances in Neural Information Processing Systems , 35:36789--36803

  32. [40]

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929

  33. [41]

    B., and Wu, J

    Du, Y., Zhang, Y., Yu, H.-X., Tenenbaum, J. B., and Wu, J. (2021). Neural radiance flow for 4d view synthesis and video processing. In ICCV . IEEE Computer Society

  34. [42]

    Duarte, K., Rawat, Y., and Shah, M. (2018). Videocapsulenet: A simplified network for action detection. In NeurIPS

  35. [43]

    El Banani, M., Raj, A., Maninis, K.-K., Kar, A., Li, Y., Rubinstein, M., Sun, D., Guibas, L., Johnson, J., and Jampani, V. (2024). Probing the 3d awareness of visual foundation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) ...

  36. [44]

    F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M

    Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T. (2022). Savi++: Towards end-to-end object-centric learning from real-world videos. arXiv preprint arXiv:2206.07764

  37. [45]

    R., Jones, O

    Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I. (2020). GENESIS : Generative scene inference and sampling with object-centric latent representations. In International Conference on Learning Representations

  38. [46]

    Fang, J., Yi, T., Wang, X., Xie, L., Zhang, X., Liu, W., Nie ner, M., and Tian, Q. (2022). Fast dynamic radiance fields with time-aware neural voxels. arXiv preprint arXiv:2205.15285

  39. [47]

    Fridovich-Keil, S., Yu, A., Tancik, M., Chen, Q., Recht, B., and Kanazawa, A. (2022). Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5501--5510

  40. [48]

    Gadelha, M., Wang, R., and Maji, S. (2019). Shape reconstruction using differentiable projections and deep priors. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 22--30

  41. [49]

    Gao, C., Saraf, A., Kopf, J., and Huang, J.-B. (2021). Dynamic view synthesis from dynamic monocular video. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 5712--5721

  42. [50]

    Geiger, A., Lenz, P., Stiller, C., and Urtasun, R. (2013). Vision meets robotics: The kitti dataset. The International Journal of Robotics Research , 32(11):1231--1237

  43. [51]

    Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence , 2(11):665--673

  44. [52]

    Goli, L., Reading, C., Sellán, S., Jacobson, A., and Tagliasacchi, A. (2024). Bayes' Rays : Uncertainty quantification in neural radiance fields. In CVPR

  45. [53]

    Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Sch \"o lkopf, B. (2019). Recurrent independent mechanisms. arXiv preprint arXiv:1909.10893

  46. [54]

    J., Gnanapragasam, D., Golemo, F., Herrmann, C., Kipf, T., Kundu, A., Lagun, D., Laradji, I., Liu, H.-T

    Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D. J., Gnanapragasam, D., Golemo, F., Herrmann, C., Kipf, T., Kundu, A., Lagun, D., Laradji, I., Liu, H.-T. D., Meyer, H., Miao, Y., Nowrouzezahrai, D., Oztireli, C., Pot, E., Radwan, N., Rebain, D....

  47. [55]

    L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M., and Lerchner, A

    Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M., and Lerchner, A. (2019). Multi-object representation learning with iterative variational inference. In International Conference on Machine Learning

  48. [56]

    H., Valpola, H., and Schmidhuber, J

    Greff, K., Rasmus, A., Berglund, M., Hao, T. H., Valpola, H., and Schmidhuber, J. (2016). Tagger: Deep unsupervised perceptual grouping. In Advances in Neural Information Processing Systems

  49. [57]

    Greff, K., Van Steenkiste, S., and Schmidhuber, J. (2020). On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208

  50. [58]

    Hahn, T., Pyeon, M., and Kim, G. (2019). Self-routing capsule networks. In NeurIPS

  51. [59]

    He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 770--778

  52. [60]

    Hedlin, E., Sharma, G., Mahajan, S., He, X., Isack, H., Rhodin, A. K. H., Tagliasacchi, A., and Yi, K. M. (2024). Unsupervised Keypoints from Pretrained Diffusion Models . In CVPR

  53. [61]

    P., Mildenhall, B., Barron, J

    Hedman, P., Srinivasan, P. P., Mildenhall, B., Barron, J. T., and Debevec, P. (2021). Baking neural radiance fields for real-time view synthesis. ICCV

  54. [62]

    Hinton, G. (2021). How to represent part-whole hierarchies in a neural network. arXiv preprint arXiv:2102.12627

  55. [63]

    E., Krizhevsky, A., and Wang, S

    Hinton, G. E., Krizhevsky, A., and Wang, S. D. (2011a). Transforming auto-encoders. In International conference on artificial neural networks , pages 44--51. Springer

  56. [64]

    E., Krizhevsky, A., and Wang, S

    Hinton, G. E., Krizhevsky, A., and Wang, S. D. (2011b). Transforming auto-encoders. In International Conference on Artifical Neural Networks

  57. [65]

    E., Sabour, S., and Frosst, N

    Hinton, G. E., Sabour, S., and Frosst, N. (2018a). Matrix capsules with em routing. In International conference on learning representations

  58. [66]

    E., Sabour, S., and Frosst, N

    Hinton, G. E., Sabour, S., and Frosst, N. (2018b). Matrix capsules with EM routing. In ICLR

  59. [67]

    Ho, J., Jain, A., and Abbeel, P. (2020). Denoising diffusion probabilistic models. NeurIPS

  60. [68]

    and Murphy, K

    Huang, J. and Murphy, K. (2016). Efficient inference in occlusion-aware generative models of images. ICLR Workshop (arXiv:1511.06362)

  61. [69]

    and El-kenawy, E.-S

    Ibrahim, A. and El-kenawy, E.-S. M. (2020). Image segmentation methods based on superpixel techniques: A survey. Journal of Computer Science and Information Systems

  62. [70]

    Jain, A., Tancik, M., and Abbeel, P. (2021). Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis . In ICCV

  63. [71]

    Jason, J., Harley, A., and Derpanis, K. (2016). Back to basics: U nsupervised learning of optical flow via brightness constancy and motion smoothnes. ECCV , pages 3--10

  64. [72]

    Jeong, Y., Ahn, S., Choy, C., Anandkumar, A., Cho, M., and Park, J. (2021). Self-calibrating neural radiance fields. In ICCV

  65. [73]

    J., and Black, M

    Jepson, A., Fleet, D. J., and Black, M. J. (2002). A layered motion representation with occlusion and compact spatial support. ECCV , pages 692--706

  66. [74]

    Johnson-Laird, P. N. (2010). Mental models and human reasoning. Proceedings of the National Academy of Sciences , 107(43):18243--18250

  67. [75]

    and Frey, B

    Jojic, N. and Frey, B. (2001). Learning flexible sprites in video layers. CVPR

  68. [76]

    Kajiya, J. T. and Von Herzen, B. P. (1984). Ray tracing volume densities. In ACM TOG

  69. [77]

    Kanizsa, G. (1976). Subjective contours. Scientific American , 234(4):48--53

  70. [78]

    A., Schieber, H., Schischka, N., Görgülü, M., Grötzner, F., Ladikos, A., Roth, D., Navab, N., and Busam, B

    Karaoglu, M. A., Schieber, H., Schischka, N., Görgülü, M., Grötzner, F., Ladikos, A., Roth, D., Navab, N., and Busam, B. (2023). Dynamon: Motion-aware fast and robust camera localization for dynamic nerf

  71. [79]

    Karnewar, A., Ritschel, T., Wang, O., and Mitra, N. (2022). Relu fields: The little non-linearity that could. ACM TOG (Proc. SIGGRAPH)

  72. [80]

    Kerbl, B., Kopanas, G., Leimkuehler, T., and Drettakis, G. (2023). 3D Gaussian Splatting for Real - Time Radiance Field Rendering . ACM TOG (Proc. SIGGRAPH)

  73. [81]

    Kerbl, B., Meuleman, A., Kopanas, G., Wimmer, M., Lanvin, A., and Drettakis, G. (2024). A hierarchical 3d gaussian representation for real-time rendering of very large datasets. ACM TOG (Proc. SIGGRAPH)

  74. [82]

    Kim, I., Choi, M., and Kim, H. J. (2023). Up-nerf: Unconstrained pose-prior-free neural radiance fields. arXiv:2311.03784

  75. [83]

    and Ba, J

    Kingma, D. and Ba, J. (2014). Adam: A method for stochastic optimization. CoRR

  76. [84]

    Kingma, D., Salimans, T., Poole, B., and Ho, J. (2021). Variational diffusion models. Advances in neural information processing systems , 34:21696--21707

  77. [85]

    F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K

    Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K. (2021). Conditional object-centric learning from video. Under submission

  78. [86]

    Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S. (2017). Self-normalizing neural networks. arXiv preprint arXiv:1706.02515

  79. [87]

    W., and Hinton, G

    Kosiorek, A., Sabour, S., Teh, Y. W., and Hinton, G. E. (2019). Stacked capsule autoencoders. NeurIPS , pages 15486--15496

  80. [88]

    Kulhanek, J., Peng, S., Kukelova, Z., Pollefeys, M., and Sattler, T. (2024). WildGaussians : 3D Gaussian Splatting In the Wild . arXiv

  81. [89]

    and Seitz, S

    Kutulakos, K. and Seitz, S. (2000). A theory of shape by space carving. IJCV

  82. [90]

    and Bagci, U

    LaLonde, R. and Bagci, U. (2018). Capsules for object segmentation. arXiv preprint arXiv:1804.04241

  83. [91]

    Lee, J., Kim, I., Heo, H., and Kim, H. J. (2023). Semantic-aware occlusion filtering neural radiance fields in the wild. arXiv preprint arXiv:2303.03966

  84. [92]

    Lee, S., Hwang, I., Kang, G.-C., and Zhang, B.-T. (2022). Improving robustness to texture bias via shape-focused augmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4323--4331

  85. [93]

    Levy, A., Matthews, M., Sela, M., Wetzstein, G., and Lagun, D. (2024). MELON : NeRF with Unposed Images in SO (3). In 3DV

  86. [94]

    Li, P., Wang, S., Yang, C., Liu, B., Qiu, W., and Wang, H. (2023). Nerf-ms: Neural radiance fields with multi-sequence. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 18591--18600

  87. [95]

    Li, T., Slavcheva, M., Zollhoefer, M., Green, S., Lassner, C., Kim, C., Schmidt, T., Lovegrove, S., Goesele, M., Newcombe, R., et al. (2022). Neural 3d video synthesis from multi-view video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ,...

  88. [96]

    and Chen, J

    Li, Z. and Chen, J. (2015). Superpixel segmentation using linear spectral clustering. In CVPR

  89. [97]

    Li, Z., Niklaus, S., Snavely, N., and Wang, O. (2021a). Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 6498--6508

  90. [98]

    Li, Z., Niklaus, S., Snavely, N., and Wang, O. (2021b). Neural scene flow fields for space-time view synthesis of dynamic scenes. In CVPR

  91. [99]

    Lin, C.-H., Ma, W.-C., Torralba, A., and Lucey, S. (2021a). Barf: Bundle-adjusting neural radiance fields. In CVPR

  92. [100]

    Lin, C.-H., Ma, W.-C., Torralba, A., and Lucey, S. (2021b). BARF: Bundle-Adjusting Neural Radiance Fields . In ICCV

  93. [101]

    D., Williams, F., Jacobson, A., Fidler, S., and Litany, O

    Liu, H.-T. D., Williams, F., Jacobson, A., Fidler, S., and Litany, O. (2022a). Learning smooth neural functions via lipschitz regularization. ACM TOG

  94. [102]

    J., Keppo, J., Shan, Y., Qie, X., and Shou, M

    Liu, J.-W., Cao, Y.-P., Mao, W., Zhang, W., Zhang, D. J., Keppo, J., Shan, Y., Qie, X., and Shou, M. Z. (2022b). Devrf: Fast deformable voxel radiance fields for dynamic scenes. arXiv preprint arXiv:2205.15723

  95. [103]

    Liu, L., Gu, J., Zaw Lin, K., Chua, T.-S., and Theobalt, C. (2020). Neural sparse voxel fields. Advances in Neural Information Processing Systems , 33:15651--15663

  96. [104]

    Liu, Y.-L., Gao, C., Meuleman, A., Tseng, H.-Y., Saraf, A., Kim, C., Chuang, Y.-Y., Kopf, J., and Huang, J.-B. (2023). Robust dynamic radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 13--23

  97. [105]

    Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T. (2020). Object-centric learning with slot attention. In Advances in Neural Information Processing Systems

  98. [106]

    Lombardi, S., Simon, T., Saragih, J., Schwartz, G., Lehrmann, A., and Sheikh, Y. (2019). Neural volumes: Learning dynamic renderable volumes from images. ACM TOG

  99. [107]

    Luiten, J., Kopanas, G., Leibe, B., and Ramanan, D. (2023). Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713

  100. [108]

    H., Holynski, A., and Darrell, T

    Luo, G., Dunlap, L., Park, D. H., Holynski, A., and Darrell, T. (2023). Diffusion Hyperfeatures : Searching Through Time and Space for Semantic Correspondence . In NeurIPS

  101. [109]

    Ma, L., Li, X., Liao, J., Zhang, Q., Wang, X., Wang, J., and Sander, P. V. (2022). Deblur-nerf: Neural radiance fields from blurry images. In CVPR

  102. [110]

    Mahendran, A., Thewlis, J., and Vedaldi, A. (2018). Self-supervised segmentation by grouping optical-flow. ECCV Workshop

  103. [111]

    Marchisio, A., De Marco, A., Colucci, A., Martina, M., and Shafique, M. (2023). Robcaps: evaluating the robustness of capsule networks against affine transformations and adversarial attacks. In 2023 International Joint Conference on Neural Networks (IJCNN) , pages 1--9. IEEE

  104. [112]

    S., Barron, J

    Martin-Brualla, R., Radwan, N., Sajjadi, M. S., Barron, J. T., Dosovitskiy, A., and Duckworth, D. (2021a). Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages...

  105. [113]

    S., Barron, J

    Martin-Brualla, R., Radwan, N., Sajjadi, M. S., Barron, J. T., Dosovitskiy, A., and Duckworth, D. (2021b). Nerf in the wild: Neural radiance fields for unconstrained photo collections. In CVPR

  106. [114]

    Martin-Brualla, R., Radwan, N., Sajjadi, M. S. M., Barron, J., Dosovitskiy, A., and Duckworth, D. (2021c). NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections . In CVPR

  107. [115]

    Mather, G. (2006). Foundations of perception . Psychology Press

  108. [116]

    Mescheder, L., Oechsle, M., Niemeyer, M., Nowozin, S., and Geiger, A. (2019). Occupancy networks: Learning 3d reconstruction in function space. CVPR

  109. [117]

    P., and Barron, J

    Mildenhall, B., Hedman, P., Martin-Brualla, R., Srinivasan, P. P., and Barron, J. T. (2021). NeRF in the dark: High dynamic range view synthesis from noisy raw images. In CVPR

  110. [118]

    Mildenhall, B., Srinivasan, P., Tancik, M., Barron, J., Ramamoorthi, R., and Ng, R. (2020a). NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis . In ECCV

  111. [119]

    P., Tancik, M., Ramamoorthi, R., and Ng, R

    Mildenhall, B., Srinivasan, P. P., Tancik, M., Ramamoorthi, R., and Ng, R. (2020b). Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV

  112. [120]

    P., Hedman, P., Martin-Brualla, R., and Barron, J

    Mildenhall, B., Verbin, D., Srinivasan, P. P., Hedman, P., Martin-Brualla, R., and Barron, J. T. (2022). MultiNeRF : A Code Release for Mip-NeRF 360, Ref-NeRF , and RawNeRF

  113. [121]

    and Van Nguyen, H

    Mobiny, A. and Van Nguyen, H. (2018). Fast capsnet for lung cancer screening. In Conference on Medical Image Computing and Computer-Assisted Intervention , pages 741--749. Springer

  114. [122]

    M\"uller, T., Evans, A., Schied, C., and Keller, A. (2022). Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG (Proc. SIGGRAPH)

  115. [123]

    M \"u ller, T., Evans, A., Schied, C., and Keller, A. (2022). Instant neural graphics primitives with a multiresolution hash encoding. TOG

  116. [124]

    M \"u llner, D. (2011). Modern hierarchical, agglomerative clustering algorithms. arXiv

  117. [125]

    T., Mildenhall, B., Sajjadi, M

    Niemeyer, M., Barron, J. T., Mildenhall, B., Sajjadi, M. S. M., Geiger, A., and Radwan, N. (2022). Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In CVPR

  118. [126]

    Oechsle, M., Peng, S., and Geiger, A. (2021). Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In ICCV

  119. [127]

    Olah, C., Mordvintsev, A., and Schubert, L. (2017). Feature visualization. Distill , 2(11):e7

  120. [128]

    Oquab, M., Darcet, T., Moutakanni, T., Vo, H. V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Maira...

  121. [129]

    J., Florence, P., Straub, J., Newcombe, R., and Lovegrove, S

    Park, J. J., Florence, P., Straub, J., Newcombe, R., and Lovegrove, S. (2019). DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation . In CVPR

  122. [130]

    T., Bouaziz, S., Goldman, D

    Park, K., Sinha, U., Barron, J. T., Bouaziz, S., Goldman, D. B., Seitz, S. M., and Martin-Brualla, R. (2021a). Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 5865--5874

  123. [131]

    T., Bouaziz, S., Goldman, D

    Park, K., Sinha, U., Hedman, P., Barron, J. T., Bouaziz, S., Goldman, D. B., Martin-Brualla, R., and Seitz, S. M. (2021b). Hypernerf: a higher-dimensional representation for topologically varying neural radiance fields. ACM Transactions on Graphics (TOG)

  124. [132]

    C., Kim, J.-Y., and Kang, N

    Park, S., Son, M., Jang, S., Ahn, Y. C., Kim, J.-Y., and Kang, N. (2023). Temporal interpolation is all you need for dynamic neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4212--4221

  125. [133]

    Peng, S., Niemeyer, M., Mescheder, L., Pollefeys, M., and Geiger, A. (2020). Convolutional occupancy networks. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part III 16 , pages 523--540. Springer

  126. [134]

    and Hebert, P

    Perreault, S. and Hebert, P. (2007). Median filtering in constant time. In IEEE Transactions on Image Processing

  127. [135]

    W., Kabir, T., and Liu, F

    Picard, R. W., Kabir, T., and Liu, F. (1993). Real-time recognition with the entire brodatz texture database. In CVPR

  128. [136]

    T., and Mildenhall, B

    Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. (2022). Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint 2209.14988

  129. [137]

    Pumarola, A., Corona, E., Pons-Moll, G., and Moreno-Noguer, F. (2021). D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10318--10327

  130. [138]

    R., Su, H., Mo, K., and Guibas, L

    Qi, C. R., Su, H., Mo, K., and Guibas, L. J. (2017). Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 652--660

  131. [139]

    Qian, Z., Wang, S., Mihajlovic, M., Geiger, A., and Tang, S. (2024). 3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5020--5030

  132. [140]

    Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A. (2019). On the spectral bias of neural networks. In ICML , pages 5301--5310. PMLR

  133. [141]

    and Hanrahan, P

    Ramamoorthi, R. and Hanrahan, P. (2001). An efficient representation for irradiance environment maps. In ACM TOG

  134. [142]

    Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv

  135. [143]

    Rawlinson, D., Ahmed, A., and Kowadlo, G. (2018). Sparse unsupervised capsules generalize better. CoRR

  136. [144]

    M., Lagun, D., and Tagliasacchi, A

    Rebain, D., Matthews, M., Yi, K. M., Lagun, D., and Tagliasacchi, A. (2022a). LOLNeRF: Learn from One Look . In CVPR

  137. [145]

    M., Lagun, D., and Tagliasacchi, A

    Rebain, D., Matthews, M., Yi, K. M., Lagun, D., and Tagliasacchi, A. (2022b). LOLNerf : Learn from one look. In CVPR

  138. [146]

    Reizenstein, J., Shapovalov, R., Henzler, P., Sbordone, L., Labatut, P., and Novotny, D. (2021). Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In ICCV

  139. [147]

    P., Barron, J

    Rematas, K., Liu, A., Srinivasan, P. P., Barron, J. T., Tagliasacchi, A., Funkhouser, T., and Ferrari, V. (2022a). Urban radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12932--12942

  140. [148]

    P., Barron, J

    Rematas, K., Liu, A., Srinivasan, P. P., Barron, J. T., Tagliasacchi, A., Funkhouser, T., and Ferrari, V. (2022b). Urban radiance fields. CVPR

  141. [149]

    P., Barron, J

    Rematas, K., Liu, A., Srinivasan, P. P., Barron, J. T., Tagliasacchi, A., Funkhouser, T., and Ferrari, V. (2022c). Urban Radiance Fields . In CVPR

  142. [150]

    Ren, W., Zhu, Z., Sun, B., Chen, J., Pollefeys, M., and Peng, S. (2024). NeRF On -the-go: Exploiting Uncertainty for Distractor -free NeRFs in the Wild . In CVPR

  143. [151]

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2021). High-resolution image synthesis with latent diffusion models. 2022 ieee. In CVPR

  144. [152]

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022). High- Resolution Image Synthesis With Latent Diffusion Models . In CVPR

  145. [153]

    J., and Carlone, L

    Rosinol, A., Leonard, J. J., and Carlone, L. (2022). Nerf-slam: Real-time dense monocular slam with neural radiance fields. arXiv preprint arXiv:2210.13641

  146. [154]

    Sabour, S., Frosst, N., and Hinton, G. E. (2017a). Dynamic routing between capsules. Advances in neural information processing systems , 30

  147. [155]

    Sabour, S., Frosst, N., and Hinton, G. E. (2017b). Dynamic routing between capsules. In NeurIPS

  148. [156]

    J., and Tagliasacchi, A

    Sabour, S., Goli, L., Kopanas, G., Matthews, M., Lagun, D., Guibas, L., Jacobson, A., Fleet, D. J., and Tagliasacchi, A. (2024). Spotlesssplats: Ignoring distractors in 3d gaussian splatting. arXiv preprint arXiv:2406.20055

  149. [157]

    Sabour, S., Tagliasacchi, A., Yazdani, S., Hinton, G., and Fleet, D. J. (2021). Unsupervised part representation by flow capsules. In International Conference on Machine Learning , pages 9213--9223. PMLR

  150. [158]

    J., and Tagliasacchi, A

    Sabour, S., Vora, S., Duckworth, D., Krasin, I., Fleet, D. J., and Tagliasacchi, A. (2023). Robustnerf: Ignoring distractors with robust losses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 20626--20636

  151. [159]

    L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al

    Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. (2022). Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS

  152. [160]

    S., Meyer, H., Pot, E., Bergmann, U., Greff, K., Radwan, N., Vora, S., Lu c i \'c , M., Duckworth, D., Dosovitskiy, A., et al

    Sajjadi, M. S., Meyer, H., Pot, E., Bergmann, U., Greff, K., Radwan, N., Vora, S., Lu c i \'c , M., Duckworth, D., Dosovitskiy, A., et al. (2022). Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. In Proceedings of t...

  153. [161]

    Santoro, A., Faulkner, R., Raposo, D., Rae, J., Chrzanowski, M., Weber, T., Wierstra, D., Vinyals, O., Pascanu, R., and Lillicrap, T. (2018). Relational recurrent neural networks. Advances in neural information processing systems , 31

  154. [162]

    Schonberger, J. L. and Frahm, J.-M. (2016). Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 4104--4113

  155. [163]

    Sch\" o nberger, J. L. and Frahm, J.-M. (2016). Structure-from-motion revisited. In CVPR

  156. [164]

    Schonberger, J. L. and Frahm, J.-M. (2016). Structure-from-motion revisited. In CVPR

  157. [165]

    L., Zheng, E., Pollefeys, M., and Frahm, J.-M

    Sch\" o nberger, J. L., Zheng, E., Pollefeys, M., and Frahm, J.-M. (2016). Pixelwise view selection for unstructured multi-view stereo. In ECCV

  158. [166]

    D., Patel, V., and Patel, Y

    Shah, M., Gandhi, K., Joshi, S., Nagar, M. D., Patel, V., and Patel, Y. (2023). Adversarial attacks and defenses in capsule networks: A critical review of robustness challenges and mitigation strategies. In International Conference on Advanced Computing Techniques in Engineeri...

  159. [167]

    and Hoffman, D

    Singh, M. and Hoffman, D. (2001). Part-based representations of visual shape and implications for visual cognition. In Kellman, P. and Shipley, T., editors, From fragments to objects: Segmentation and grouping in vision , chapter 9, pages 401--459. Elsevier Science

  160. [168]

    Sitzmann, V., Zollh \"o fer, M., and Wetzstein, G. (2019a). Scene representation networks: Continuous 3d-structure-aware neural scene representations. In Proc. NeurIPs

  161. [169]

    Sitzmann, V., Zollhöfer, M., and Wetzstein, G. (2019b). Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations . In NeurIPS

  162. [170]

    and Hauberg, S

    Skafte, N. and Hauberg, S. (2019). Explicit disentanglement of appearance and perspective in generative models. NeurIPS , pages 1018--1028

  163. [171]

    M., and Szeliski, R

    Snavely, N., Seitz, S. M., and Szeliski, R. (2006). Photo tourism: Exploring photo collections in 3d. In ACM TOG

  164. [172]

    Sohn, K. (2016). Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems , 29

  165. [173]

    and Ermon, S

    Song, Y. and Ermon, S. (2019). Generative modeling by estimating gradients of the data distribution. NeurIPS

  166. [174]

    Spelke, E. S. (1990). Principles of object perception. Cognitive science , 14(1):29--56

  167. [175]

    Spelke, E. S. and Kinzler, K. D. (2007). Core knowledge. Developmental science , 10(1):89--96

  168. [176]

    Srivastava, N., Goh, H., and Salakhutdinov, R. (2019). Geometric capsule autoencoders for 3d point clouds. arXiv preprint arXiv:1912.03310

  169. [177]

    Sun, C., Sun, M., and Chen, H.-T. (2022a). Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR

  170. [178]

    Sun, J., Chen, X., Wang, Q., Li, Z., Averbuch-Elor, H., Zhou, X., and Snavely, N. (2022b). Neural 3d reconstruction in the wild. In ACM SIGGRAPH 2022 Conference Proceedings , pages 1--9

  171. [179]

    Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al. (2020a). Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern r...

  172. [180]

    Sun, W., Jiang, W., Tagliasacchi, A., Trulls, E., and Yi, K. M. (2020b). ACNe: Attentive Context Normalization for Robust Permutation-Equivariant Learning . CVPR

  173. [181]

    E., and Yi, K

    Sun, W., Tagliasacchi, A., Deng, B., Sabour, S., Yazdani, S., Hinton, G. E., and Yi, K. M. (2020c). Canonical capsules: Unsupervised capsules in canonical pose. arXiv preprint

  174. [182]

    E., and Yi, K

    Sun, W., Tagliasacchi, A., Deng, B., Sabour, S., Yazdani, S., Hinton, G. E., and Yi, K. M. (2021). Canonical capsules: Self-supervised capsules in canonical pose. Advances in Neural Information Processing Systems , 34:24993--25005

  175. [183]

    Suwajanakorn, S., Snavely, N., Tompson, J., and Norouzi, M. (2019). Discovery of latent 3d keypoints via end-to-end geometric reasoning. NeurIPS

  176. [184]

    and Li, H

    Tagliasacchi, A. and Li, H. (2016). Modern techniques and applications for real-time non-rigid registration. In Proc. SIGGRAPH Asia (Technical Course Notes)

  177. [185]

    and Mildenhall, B

    Tagliasacchi, A. and Mildenhall, B. (2022). Volume Rendering Digest (for NeRF)

  178. [186]

    Takikawa, T. (2021). Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In CVPR

  179. [187]

    P., Barron, J

    Tancik, M., Casser, V., Yan, X., Pradhan, S., Mildenhall, B., Srinivasan, P. P., Barron, J. T., and Kretzschmar, H. (2022a). Block-nerf: Scalable large scene neural view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8248--8258

  180. [188]

    P., Barron, J

    Tancik, M., Casser, V., Yan, X., Pradhan, S., Mildenhall, B., Srinivasan, P. P., Barron, J. T., and Kretzschmar, H. (2022b). Block- NeRF : Scalable Large Scene Neural View Synthesis . In CVPR

  181. [189]

    P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J

    Tancik, M., Srinivasan, P. P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J. T., and Ng, R. (2020). Fourier features let networks learn high frequency functions in low dimensional domains. arXiv preprint arXiv:2006.10739

  182. [190]

    P., and Hariharan, B

    Tang, L., Jia, M., Wang, Q., Phoo, C. P., and Hariharan, B. (2023). Emergent correspondence from image diffusion. In NeurIPS

  183. [191]

    and Deng, J

    Teed, Z. and Deng, J. (2020). RAFT : Recurrent all-pairs field transforms for optical flow. In ECCV

  184. [192]

    Tewari, A., Thies, J., Mildenhall, B., Srinivasan, P., Tretschk, E., Yifan, W., Lassner, C., Sitzmann, V., Martin-Brualla, R., Lombardi, S., et al. (2022). Advances in neural rendering. In Computer Graphics Forum

  185. [193]

    Thies, J., Zollh \"o fer, M., and Nie ner, M. (2019). Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG) , 38(4):1--12

  186. [194]

    Tretschk, E., Tewari, A., Golyanik, V., Zollh \"o fer, M., Lassner, C., and Theobalt, C. (2021). Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video. In ICCV

  187. [195]

    Tschernezki, V., Larlus, D., and Vedaldi, A. (2021). NeuralDiff : Segmenting 3D objects that move in egocentric videos. In Proc. 3DV

  188. [196]

    Turki, H., Ramanan, D., and Satyanarayanan, M. (2022). Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12922--12931

  189. [197]

    Van Steenkiste, S., Chang, M., Greff, K., and Schmidhuber, J. (2018). Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. arXiv preprint arXiv:1802.10353

  190. [198]

    van Steenkiste, S., Zoran, D., Yang, Y., Rubanova, Y., Kabra, R., Doersch, C., Gokay, D., Pot, E., Greff, K., Hudson, D., et al. (2025). Moving off-the-grid: Scene-grounded video representations. Advances in Neural Information Processing Systems , 37:124319--124346

  191. [199]

    N., Kaiser, ., and Polosukhin, I

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems , 30

  192. [200]

    D., Chang, M., Janner, M., Finn, C., Wu, J., Tenenbaum, J., and Levine, S

    Veerapaneni, R., Co-Reyes, J. D., Chang, M., Janner, M., Finn, C., Wu, J., Tenenbaum, J., and Levine, S. (2020). Entity abstraction in visual model-based reinforcement learning. In Conference on Robot Learning , pages 1439--1456. PMLR

  193. [201]

    T., and Srinivasan, P

    Verbin, D., Hedman, P., Mildenhall, B., Zickler, T., Barron, J. T., and Srinivasan, P. P. (2022). Ref-NeRF : Structured view-dependent appearance for neural radiance fields. CVPR

  194. [202]

    Vijayanarasimhan, S., Ricco, S., Schmid, C., Sukthankar, R., and Fragkiadaki, K. (2017). Sfm-net: Learning of structure and motion from video. arXiv preprint arXiv:1704.07804

  195. [203]

    S., Pot, E., Tagliasacchi, A., and Duckworth, D

    Vora, S., Radwan, N., Greff, K., Meyer, H., Genova, K., Sajjadi, M. S., Pot, E., Tagliasacchi, A., and Duckworth, D. (2021). Nesf: Neural semantic fields for generalizable semantic segmentation of 3d scenes. TMLR

  196. [204]

    H., Kubovy, M., E., P

    Wagemans, J., Elder, J. H., Kubovy, M., E., P. S., A., P. M., M., S., and von der Heydt, R. (2012). A century of gestalt psychology in visual perception: I. perceptual grouping and figure–ground organization. Psychological Bulletin , 138(6):1172–1217

  197. [205]

    Wang, C., Eckart, B., Lucey, S., and Gallo, O. (2021a). Neural trajectory fields for dynamic novel view synthesis. arXiv preprint arXiv:2105.05994

  198. [206]

    and Adelson , E

    Wang , J. and Adelson , E. H. (1994). Representing moving images with layers. IEEE Trans Image Processing , 3(5):625--638

  199. [207]

    Wang, J., Wang, P., Long, X., Theobalt, C., Komura, T., Liu, L., and Wang, W. (2022). Neuris: Neural reconstruction of indoor scenes using normal priors. arXiv preprint

  200. [208]

    Wang, P., Liu, L., Liu, Y., Theobalt, C., Komura, T., and Wang, W. (2021b). Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS

  201. [209]

    Wang, Y., Wang, J., and Qi, Y. (2024). WE - GS : An In -the-wild Efficient 3D Gaussian Representation for Unconstrained Photo Collections

  202. [210]

    Wang, Z., Bovik, A., Sheikh, H., and Simoncelli, E. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing

  203. [211]

    Wang, Z., Wu, S., Xie, W., Chen, M., and Prisacariu, V. A. (2021c). Nerf--: Neural radiance fields without known camera parameters. arXiv preprint arXiv:2102.07064

  204. [212]

    Wang, Z., Wu, S., Xie, W., Chen, M., and Prisacariu, V. A. (2021d). NeRF--: Neural Radiance Fields Without Known Camera Parameters

  205. [213]

    Weiss, Y. (1997). Smoothness in layers: M otion segmentation using nonparametric mixture estima-tion. IEEE CVPR

  206. [214]

    Wu, G., Yi, T., Fang, J., Xie, L., Zhang, X., Wei, W., Liu, W., Tian, Q., and Wang, X. (2024). 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 20310--20320

  207. [215]

    Wu, T., Zhong, F., Cole, F., Tagliasacchi, A., and Oztireli, C. (2022a). D2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video. In NeurIPS

  208. [216]

    Wu, T., Zhong, F., Tagliasacchi, A., Cole, F., and Oztireli, C. (2022b). D2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video. Advances in Neural Information Processing Systems , 35:32653--32666

  209. [217]

    W., Zhong, F., Tagliasacchi, A., Cole, F., and Oztireli, C

    Wu, T. W., Zhong, F., Tagliasacchi, A., Cole, F., and Oztireli, C. (2022c). D 2NeRF : Self - Supervised Decoupling of Dynamic and Static Objects from a Monocular Video . In NeurIPS

  210. [218]

    Wu, Z., Rubanova, Y., Kabra, R., Hudson, D., Gilitschenski, I., Aytar, Y., van Steenkiste, S., Allen, K., and Kipf, T. (2025). Neural assets: 3d-aware multi-object scene synthesis with image diffusion models. Advances in Neural Information Processing Systems , 37:76289--76318

  211. [219]

    Xian, W., Huang, J.-B., Kopf, J., and Kim, C. (2021). Space-time neural irradiance fields for free-viewpoint video. In CVPR

  212. [220]

    Xie, Y., Takikawa, T., Saito, S., Litany, O., Yan, S., Khan, N., Tombari, F., Tompkin, J., Sitzmann, V., and Sridhar, S. (2022). Neural fields in visual computing and beyond. Comput. Graph. Forum

  213. [221]

    Xu, J., Mei, Y., and Patel, V. M. (2024). Wild-gs: Real-time novel view synthesis from unconstrained photo collections

  214. [222]

    T., Tenenbaum, J

    Xu, Z., Liu, Z., Sun, C., Murphy, K., Freeman, W. T., Tenenbaum, J. B., and Wu, J. (2019). Unsupervised discovery of parts, structure, and dynamics. In ICLR

  215. [223]

    Yan, C., Qu, D., Xu, D., Zhao, B., Wang, Z., Wang, D., and Li, X. (2024). Gs-slam: Dense visual slam with 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 19595--19604

  216. [224]

    Yang, Y., Zhang, S., Huang, Z., Zhang, Y., and Tan, M. (2023). Cross- Ray Neural Radiance Fields for Novel - View Synthesis from Unconstrained Image Collections . In ICCV

  217. [225]

    Yi, T., Fang, J., Wang, J., Wu, G., Xie, L., Zhang, X., Liu, W., Tian, Q., and Wang, X. (2024). Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  218. [226]

    Yu, A., Li, R., Tancik, M., Li, H., Ng, R., and Kanazawa, A. (2021a). PlenOctrees for Real-Time Rendering of Neural Radiance Fields . In ICCV

  219. [227]

    Yu, A., Ye, V., Tancik, M., and Kanazawa, A. (2021b). pixelnerf: Neural radiance fields from one or few images. In Proc. CVPR

  220. [228]

    Yu, A., Ye, V., Tancik, M., and Kanazawa, A. (2021c). pixelNeRF: Neural Radiance Fields from One or Few Images . In CVPR

  221. [229]

    Yu, Z., Chen, A., Huang, B., Sattler, T., and Geiger, A. (2024). Mip-splatting: Alias-free 3d gaussian splatting. In CVPR

  222. [230]

    Yu, Z., Peng, S., Niemeyer, M., Sattler, T., and Geiger, A. (2022). Monosdf: Exploring monocular geometric cues for neural implicit surface reconstruction. In NeurIPS

  223. [231]

    Zhang, D., Wang, C., Wang, W., Li, P., Qin, M., and Wang, H. (2024). Gaussian in the Wild : 3D Gaussian Splatting for Unconstrained Image Collections

  224. [232]

    P., Jampani, V., Sun, D., and Yang, M.-H

    Zhang, J., Herrmann, C., Hur, J., Cabrera, L. P., Jampani, V., Sun, D., and Yang, M.-H. (2023). A Tale of Two Features : Stable Diffusion Complements DINO for Zero - Shot Semantic Correspondence . In NeurIPS

  225. [233]

    Y., Ramanan, D., and Tulsiani, S

    Zhang, J. Y., Ramanan, D., and Tulsiani, S. (2022). Relpose: Predicting probabilistic relative rotation for single objects in the wild. In ECCV

  226. [234]

    Zhang, K., Riegler, G., Snavely, N., and Koltun, V. (2020). Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492

  227. [235]

    Zhang, Q., Nian Wu, Y., and Zhu, S.-C. (2018a). Interpretable convolutional neural networks. IEEE Conference on Computer Vision and Pattern Recognition , pages 8827--8836

  228. [236]

    A., Shechtman, E., and Wang, O

    Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. (2018b). The unreasonable effectiveness of deep features as a perceptual metric. In CVPR

  229. [237]

    and Huang, L

    Zhao, L. and Huang, L. (2021). An investigation on sparsity of capsnets for adversarial robustness. In Proceedings of the 1st International Workshop on Adversarial Learning for Multimedia , pages 55--61

  230. [238]

    Zhao, Y., Birdal, T., Deng, H., and Tombari, F. (2019). 3d point capsule networks. IEEE CVPR , pages 1009--1018

  231. [239]

    R., and Pollefeys, M

    Zhu, Z., Peng, S., Larsson, V., Xu, W., Bao, H., Cui, Z., Oswald, M. R., and Pollefeys, M. (2022). Nice-slam: Neural implicit scalable encoding for slam. In CVPR

  232. [240]

    Zwicker, M., Pfister, H., Van Baar, J., and Gross, M. (2001). Surface splatting. In ACM TOG

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.