Pith. sign in

REVIEW 3 major objections 5 minor 60 references

ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read ROSA reconstructs real objects from flashlight photos as compact adaptive meshes with arbitrarily high-resolution SVBRDF textures.

desk verdict Solid inverse rendering paper with a useful new combination of adaptive mesh refinement and tile-based texture generation; the central claim holds broadly, but the evaluation needs error bars, quantitative ablations, and a stress test on albedo-to-normal leakage. read the letter →

arxiv 2501.18595 v1 pith:B3UILJVR submitted 2025-01-30 cs.CV

classification cs.CV
keywords inverserenderingSVBRDFadaptivemeshrefinementdetailtransfernormalmaptile-basedtexturesynthesiscollocatedlightdifferentiable
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ROSA is an inverse rendering method that reconstructs a relightable 3D object from a limited set of collocated-light photographs, producing a triangle mesh and a spatially varying BRDF texture. Its central claim is that geometric detail should be allocated dynamically: details larger than roughly one pixel are transferred from the estimated normal map into the mesh itself, while the normal map keeps only the rest. The mesh is refined adaptively where texture-space curvature and mesh curvature agree, with per-vertex smoothness relaxed in those regions, so the final model stays compact without a separate simplification pass. Textures are generated tile-by-tile by a single pre-trained decoder, so their resolution is bounded by the number of tiles rather than by the network. If correct, this removes two known bottlenecks, oversized meshes and fixed low-resolution textures, in practical object capture for AR/VR and games.

What carries the argument

The load-bearing object is the pair formed by the normal loss and the adaptive control signals. The normal loss, Eq. (16), is $$L_{\text{normal}} = \|\mathbf{m} \odot ((\operatorname{sg}[\hat{\mathbf{N}}_V] + \hat{\mathbf{N}}) - \hat{\mathbf{N}}_V)\|_1,$$ where the stop-gradient on $\hat{\mathbf{N}}_V$ turns the rendered geometric normal into a fixed target derived from the texture normal, pulling vertices toward the shape implied by the normal map. The control signals are the texture-space curvature $c_n(v_i)$, a Scharr-convolution measure of normal-map variation per vertex, and the mesh curvature $c_V(v_i)$ from the Laplacian; these set per-vertex smoothness weights $\lambda_i(t)$ and edge-length thresholds $e_i(t)$ that drive local refining. This pair transfers detail from texture to geometry while keeping resolution where it is needed. The tile-based texture decoder, a pre-trained deconvolution network that maps fixed-size latent codes to 128x128 SVBRDF tiles, carries the appearance side and removes the network-resolution limit.

What would settle it

Render a flat plane with a sinusoidal normal perturbation of wavelength about two to three pixels under collocated light and run the pipeline with a known camera trajectory. If the central claim holds, the reconstructed mesh should develop corrugations matching the perturbation and the silhouette of the object should no longer be perfectly flat; if the mesh stays flat and the detail remains only in the normal texture, the transfer has failed. A quantitative version is to compare the Chamfer distance to a ground-truth mesh for this scene versus a physically flat plane with the same printed normal map.

Watch

Extended reading notes

Core claim

The paper's discovery is that the normal map can act as a hinge for geometry optimization: a stop-gradient normal loss makes the rendered geometric normal chase the rendered perturbed normal, and a curvature criterion computed from both the mesh Laplacian and the normal map decides where to subdivide. In ROSA's terms, the optimization jointly updates vertex positions and decoder weights while periodically remeshing locally (Loop subdivision with edge flips) and preconditioning smoothness non-uniformly. This transfers all visible surface detail into real 3D shape, avoiding normal-map artifacts such as wrong silhouettes and missing shadows, while texture tiles from a single pre-trained decoder allow arbitrarily high-resolution SVBRDF atlases. The paper demonstrates this on synthetic ground-truth objects and on real smartphone flashlight captures, with meshes of 30k to 71k vertices for complex objects.

Load-bearing premise

The adaptive geometry transfer assumes that the estimated texture normal is a faithful record of true geometric surface variation; if the normal map encodes illumination or albedo effects instead, as the paper concedes can happen under environment lighting, the refined mesh and relaxed smoothness will bake those artifacts into the shape.

Editorial extensions

If this is right

  • Reconstructed models are compact by construction: synthetic objects use 30k to 71k vertices, and the real-world superman capture uses 28k vertices versus 139k for USAR, with comparable or better image metrics and no post-hoc simplification.
  • Texture resolution is decoupled from the decoder's output size; total atlas resolution is controlled by the number of blended 128x128 tiles, so arbitrarily large SVBRDF textures are possible without retraining.
  • Visible geometric details (larger than one pixel in the recorded views) migrate from the normal map to the mesh, fixing silhouette and shadow errors that pure normal-map representations suffer from.
  • The method is not restricted to collocated light: replacing the rasterizer with a differentiable Monte Carlo renderer and estimating environment lighting yields results similar to existing environment-light inverse rendering methods, per the paper's experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable corollary of the transfer criterion is that if the input normal map carries no geometric signal (for example, a flat albedo shading variation), the adaptive refinement should stop and the mesh should stay coarse; running the pipeline on such a case would isolate the normal loss's role.
  • The subdivision threshold, encoded in $e_{\min}=0.01875$ and $e_{\max}=0.375$ in normalized space, implies an explicit detail-size cutoff; one could derive a closed-form relation between that threshold, the camera footprint, and the pixel size to predict which details stay in the normal map.
  • The tile mosaic formulation suggests natural extensions to streaming or out-of-core texture generation for gigapixel atlases, since each tile is decoded independently and blended only at overlaps, a consequence the authors do not explore.
  • Because the paper concedes that environment-light estimation can bake material color into normals, a promising next step would couple the normal loss with an explicit albedo-lighting regularizer to make the detail transfer robust outside collocated setups.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents ROSA, an inverse rendering method that jointly reconstructs a triangle mesh and an SVBRDF texture atlas from a limited set of collocated-light images. The mesh is initialized from a visual hull and then adaptively subdivided: the subdivision criterion and the per-vertex smoothness preconditioning are driven by a combination of texture-space normal curvature and mesh curvature. The appearance is produced by fine-tuning a pre-trained decoder network on fixed-size texture tiles, which allows arbitrarily large atlas resolutions. The method is evaluated on 14 synthetic objects and two real-world captures, with quantitative image metrics and geometry metrics, and is compared against IRON and USAR.

Significance. If the technical concerns below are resolved, ROSA would be a practically valuable contribution: it produces compact, relightable meshes with SVBRDF textures, and the tile-based decoder is a sensible way to decouple texture resolution from the network output resolution. The paper also ships source code, includes quantitative comparisons on real data, and states its limitations explicitly. The central claims, however, rest on two load-bearing assumptions: that the estimated normal map is a reliable geometric signal, and that the normal loss in Eq. (16) performs the intended detail transfer. Neither assumption is currently verified to the standard the paper needs, so the recommendation is major revision.

major comments (3)
  1. [Sec. 4.4, Eq. (16)] The normal loss as written is Lnormal = || m ⊙ ( (sg[cNV] + cN) - cNV ) ||_1. If gradients are propagated to the decoder weights Θ through cN, as stated in the text, then the term +cN in the residual drives cN toward zero during gradient descent. That would suppress the texture normals instead of transferring them to geometry, which contradicts the method's core mechanism. Please clarify the intended sign or specify that gradients to Θ are stopped for this loss; if the implemented loss matches the text, provide an analysis or experiment showing that cN does not collapse.
  2. [Sec. 4.3 and Sec. 5.2, Eqs. (5), (10), (11)] The adaptive refinement pipeline treats the estimated normal map as an encoding of true surface orientation. Under collocated light the equation I = ρ (n·l) is ambiguous between albedo and normal, so a dark albedo patch can be explained by a tilted normal and then amplified by the refinement loop: high cn lowers both the edge-length target and the smoothness weight, triggering subdivision exactly where the false signal appears. The authors concede in Sec. 5.2 that under environment lighting the estimated normals 'tend to bake the colors of the object material', but the collocated-light case, which is the paper's focus, is not tested for this failure mode. Please add a controlled experiment, e.g. a flat object with high-contrast albedo, and report whether the adaptive mesh embeds the albedo pattern as geometry; alternatively, provide a quantitative ablation that isolates albedo-to-normal leakage.
  3. [Sec. 5, Table 1 and Fig. 7] The quantitative evaluation consists of point estimates without error bars or repeated-run variance, and the ablation study in Fig. 7 is purely qualitative. Since the paper's central claim is that adaptive geometry transfer improves compactness and fidelity, the ablations should be accompanied by quantitative geometry and image metrics (e.g., CD/HD and PSNR/LPIPS) for configurations with and without local remeshing and local smoothness adaptation. Reporting mean and standard deviation over the 14 synthetic objects would also make the comparison to prior work more convincing.
minor comments (5)
  1. [Sec. 5.2 / Sec. 4.2] The solver is called BICGSTAB in Sec. 4.2 but 'BIGCGSTAB' in Sec. 5.2; please make the spelling consistent.
  2. [Sec. 6] In the conclusion, 'SVBBRDF' should be 'SVBRDF'.
  3. [Table 1] The table header is confusing: 'PSNR ↑ CD ↓ HD ↓' with subcolumns 'Image κd κs σ' does not make clear which metrics are computed on rendered images and which are computed on reference appearance features. Please restructure the table so the reader can see the per-object and per-material-type breakdown at a glance.
  4. [Sec. 4.3, Eqs. (11)-(12)] The notation for the edge-length indicator is inconsistent: Eq. (12) defines e_i(t), but the text later refers to 'ev' when deciding whether to split an edge. Please unify the notation.
  5. [Sec. 4.2, Eq. (3)] The preconditioned update uses the matrix (I + ΛV L), which is stated to be non-symmetric. It would help to specify the exact linear system solved by BICGSTAB and to state the convergence tolerance, as this affects the reproducibility of the optimization.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular reduction: outputs are driven by input images and intermediate normal/geometry estimates; only non-load-bearing self-citations.

full rationale

The central derivation is not circular. Geometry is optimized against input images through the masked image loss (Eq. 14) and silhouette loss (Eq. 15); the adaptive resolution is driven by normal-map curvature c_n (Eq. 5) and mesh curvature c_V (Eq. 6), which are intermediate quantities of the same inverse-rendering optimization rather than pre-existing target values. Eq. (16) is a consistency regularizer between texture normals and rendered surface normals, not a fitted input renamed as a prediction. The decoder pre-training on a material database (Sec. 4.1, refs. [14,17]) provides a prior for SVBRDF plausibility but does not determine the adaptive mesh output; the method is additionally compared against independent baselines (IRON [56], nvdiffrec [30], nvdiffrecmc [15]). The only self-citations are the authors' own material database [17] and their earlier USAR method [18] used as a baseline, and neither is load-bearing. The Sec. 5.2 admission that environment-light estimation can bake material colors into normals is a limitation of the assumed normal-as-geometry signal, not a circular reduction of prediction to input. No equation in the paper makes a claimed result equivalent to its inputs by construction.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on hand-set hyperparameters (smoothing bounds, edge bounds, curvature weights, loss weights, damping schedules) and several domain assumptions, the most load-bearing being that the estimated normal map faithfully encodes geometric surface variation. No new physical entities are introduced.

free parameters (5)
  • Smoothing bounds lambda_min, lambda_max = 16, 64
    Hand-chosen vertex smoothness bounds in Eq. (10); they control how strongly planar regions are regularized versus curved regions. No sensitivity analysis is provided.
  • Edge length bounds e_min, e_max = 0.01875, 0.375
    Hand-chosen minimum and maximum edge lengths in normalized space (Sec. 4.3) that determine when mesh edges are split. These directly affect how compact or dense the final mesh is.
  • Curvature scaling weights w_n, w_V = 3 and 1/16
    Weights in Eqs. (9) and (11) that map texture-space and geometry-space curvature into smoothness and edge-length indicators. Chosen without reported sensitivity analysis.
  • Loss weights w_img and w_normal = 0.05 and 0.01 for synthetic; 1e-3 and 1e-4 for real data
    Set in Sec. 4.4; the authors lower them for real-world data because the material model is limited and the losses are higher. This is per-domain tuning that affects the reported trade-offs.
  • Damping schedule constants = sigmoid(20(t-0.2)) and sigmoid(20(t-0.3))
    Eqs. (7)-(8) define when loss weighting and resolution control activate during optimization. The steepness and delays are hand-set and are not evaluated for sensitivity.
assumptions (5)
  • domain assumption Input images are captured with a collocated camera and point light, so shading is primarily a function of local surface orientation and reflectance.
    Section 3 defines inputs as images under collocated light; the whole decomposition into SVBRDF plus normals relies on this restricted illumination. The environment-light extension in Sec. 5.2 is acknowledged to be less stable.
  • domain assumption The Cook-Torrance microfacet model with GGX distribution and a 10-channel SVBRDF is sufficient to explain the observed appearance.
    Material model in Sec. 3. The authors lower image and normal loss weights for real data because the material model is limited, which shows the assumption is only approximately met.
  • domain assumption The normal map estimated by the pretrained decoder is a reliable proxy for true surface orientation, so normal-map curvature indicates where mesh refinement is needed.
    Used in Eqs. (5), (9)-(11) and (16). If the normal map bakes in lighting or albedo variation, the geometry transfer will embed non-geometric detail. The authors note this exact failure mode for environment-light reconstructions in Sec. 5.2.
  • ad hoc to paper Geometric details larger than one image pixel can be separated from albedo variation and represented as mesh geometry.
    Stated in Sec. 1 as the goal of detail transfer, but never proven or bounded. Separating geometry from albedo under collocated light is ill-posed, as the paper itself notes.
  • domain assumption The material database used to pretrain the decoder yields plausible SVBRDFs and stable refinement for arbitrary random latent codes.
    Sec. 4.1 relies on a cited database [14,17] and the pretrained autoencoder. Pretraining data and exact training protocol are not fully specified in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer." pith.science (2026). https://pith.science/paper/B3UILJVR

@misc{pith2026250118595,
  author       = {Pith},
  title        = {Pith review of: ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B3UILJVR}},
  note         = {Machine review of arXiv:2501.18595}
}
read the original abstract

Reconstructing an object's shape and appearance in terms of a mesh textured by a spatially-varying bidirectional reflectance distribution function (SVBRDF) from a limited set of images captured under collocated light is an ill-posed problem. Previous state-of-the-art approaches either aim to reconstruct the appearance directly on the geometry or additionally use texture normals as part of the appearance features. However, this requires detailed but inefficiently large meshes, that would have to be simplified in a post-processing step, or suffers from well-known limitations of normal maps such as missing shadows or incorrect silhouettes. Another limiting factor is the fixed and typically low resolution of the texture estimation resulting in loss of important surface details. To overcome these problems, we present ROSA, an inverse rendering method that directly optimizes mesh geometry with spatially adaptive mesh resolution solely based on the image data. In particular, we refine the mesh and locally condition the surface smoothness based on the estimated normal texture and mesh curvature. In addition, we enable the reconstruction of fine appearance details in high-resolution textures through a pioneering tile-based method that operates on a single pre-trained decoder network but is not limited by the network output resolution.

Figures

Figures reproduced from arXiv: 2501.18595 by the authors.

Figure 1
Figure 1. Exemplary renderings in Unity emphasizing artifacts [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of our inverse rendering framework, which performs decoder-based texture estimation and triangle mesh optimization [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Architecture of our decoder for texture estimation which [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 5
Figure 5. Figure 5: Damping functions used to increase the stability of the [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Our normal loss supports the transfer of information [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Our adaptive geometry reconstruction approach, with the [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Reconstructions of synthetic input data based on 50-70 [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Comparison of reconstruction results on real-world objects for collocated capture scenarios based on 100 (superman) and 170 [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Comparison of reconstruction results on NeRF datasets [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 59 canonical work pages

  1. [1]

    Neural Re- flectance Fields for Appearance Acquisition

    Sai Bi, Zexiang Xu, Pratul Srinivasan, Ben Mildenhall, Kalyan Sunkavalli, Milo ˇs Ha ˇsan, Yannick Hold-Geoffroy, David Kriegman, and Ravi Ramamoorthi. Neural Re- flectance Fields for Appearance Acquisition. arXiv preprint arXiv:2008.03824, 2020. 2

  2. [2]

    Deep 3D Capture: Geometry and Reflectance from Sparse Multi-View Images

    Sai Bi, Zexiang Xu, Kalyan Sunkavalli, David Kriegman, and Ravi Ramamoorthi. Deep 3D Capture: Geometry and Reflectance from Sparse Multi-View Images. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2

  3. [3]

    Bar- ron, Ce Liu, and Hendrik Lensch

    Mark Boss, Raphael Braun, Varun Jampani, Jonathan T. Bar- ron, Ce Liu, and Hendrik Lensch. NeRD: Neural Reflectance Decomposition from Image Collections. In IEEE Interna- tional Conference on Computer Vision (ICCV), 2021. 1, 2

  4. [4]

    SAMURAI: Shape And Material from Un- constrained Real-world Arbitrary Image collection

    Mark Boss, Andreas Engelhardt, Abhishek Kar, Yuanzhen Li, Deqing Sun, Jonathan Barron, Hendrik Lensch, and Varun Jampani. SAMURAI: Shape And Material from Un- constrained Real-world Arbitrary Image collection. Ad- vances in Neural Information Processing Systems (NeurIPS), 35, 2022. 2

  5. [5]

    Two-shot Spatially-varying BRDF and Shape Estimation

    Mark Boss, Varun Jampani, Kihwan Kim, Hendrik Lensch, and Jan Kautz. Two-shot Spatially-varying BRDF and Shape Estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2

  6. [6]

    A remeshing approach to multiresolution modeling

    Mario Botsch and Leif Kobbelt. A remeshing approach to multiresolution modeling. In Proceedings of the 2004 Eu- rographics/ACM SIGGRAPH Symposium on Geometry Pro- cessing, SGP ’04, New York, NY , USA, 2004. Association for Computing Machinery. 6

  7. [7]

    SupeRV ol: Super- Resolution Shape and Reflectance Estimation in Inverse V ol- ume Rendering

    Mohammed Brahimi, Bjoern Haefner, Tarun Yenamandra, Bastian Goldluecke, and Daniel Cremers. SupeRV ol: Super- Resolution Shape and Reflectance Estimation in Inverse V ol- ume Rendering. In IEEE/CVF Winter Conference on Appli- cations of Computer Vision, pages 3139–3149, 2024. 2

  8. [8]

    MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures

    Zhiqin Chen, Thomas Funkhouser, Peter Hedman, and An- drea Tagliasacchi. MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 2

Show all 60 references
  1. [9]

    Factored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy Objects

    Yue Fan, Ivan Skorokhodov, Oleg V oynov, Savva Ig- natyev, Evgeny Burnaev, Peter Wonka, and Yiqun Wang. Factored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy Objects. arXiv preprint arXiv:2305.17929, 2023. 2

  2. [10]

    Deferred Neural Lighting: Free-viewpoint Relighting from Unstructured Photographs

    Duan Gao, Guojun Chen, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. Deferred Neural Lighting: Free-viewpoint Relighting from Unstructured Photographs. ACM Transac- tions on Graphics (TOG), 39(6), 2020. 2

  3. [11]

    Surface Simplifica- tion Using Quadric Error Metrics

    Michael Garland and Paul S Heckbert. Surface Simplifica- tion Using Quadric Error Metrics. In Annual Conference on Computer Graphics and Interactive Techniques (SIG- GRAPH), 1997. 1

  4. [12]

    Ref-NeuS: Ambiguity-Reduced Neural Implicit Sur- face Learning for Multi-View Reconstruction with Reflec- tion

    Wenhang Ge, Tao Hu, Haoyu Zhao, Shu Liu, and Ying-Cong Chen. Ref-NeuS: Ambiguity-Reduced Neural Implicit Sur- face Learning for Multi-View Reconstruction with Reflec- tion. 2023. 2

  5. [13]

    Shape from Tracing: Towards Reconstructing 3D Object Geometry and SVBRDF Material from Images via Differentiable Path Tracing

    Purvi Goel, Loudon Cohen, James Guesman, Vikas Thamizharasan, James Tompkin, and Daniel Ritchie. Shape from Tracing: Towards Reconstructing 3D Object Geometry and SVBRDF Material from Images via Differentiable Path Tracing. In International Conference on 3D Vision (3DV) . IEEE...

  6. [14]

    MaterialGAN: Reflectance Capture us- ing a Generative SVBRDF Model

    Yu Guo, Cameron Smith, Milo ˇs Haˇsan, Kalyan Sunkavalli, and Shuang Zhao. MaterialGAN: Reflectance Capture us- ing a Generative SVBRDF Model. ACM Transactions on Graphics (TOG), 39(6), 2020. 2, 3

  7. [15]

    Shape, Light and Material Decomposition from Images us- ing Monte Carlo Rendering and Denoising

    Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. Shape, Light and Material Decomposition from Images us- ing Monte Carlo Rendering and Denoising. Advanced Neu- ral Information Processing Systems (NeurIPS), 35, 2022. 2, 3, 7, 8

  8. [16]

    TensoIR: Tensorial Inverse Rendering

    Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Song- fang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. TensoIR: Tensorial Inverse Rendering. InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  9. [17]

    Captur- ing Anisotropic SVBRDFs

    Julian Kaltheuner, Lukas Bode, and Reinhard Klein. Captur- ing Anisotropic SVBRDFs. In International Symposium on Vision, Modeling, and Visualization (VMV), 2021. 2, 3

  10. [18]

    Uni- fied shape and appearance reconstruction with joint camera parameter refinement

    Julian Kaltheuner, Patrick Stotko, and Reinhard Klein. Uni- fied shape and appearance reconstruction with joint camera parameter refinement. Graphical Models, 129, 2023. 2, 6, 7

  11. [19]

    Learning Efficient Illumination Multiplexing for Joint Capture of Re- flectance and Shape

    Kaizhang Kang, Cihui Xie, Chengan He, Mingqi Yi, Minyi Gu, Zimin Chen, Kun Zhou, and Hongzhi Wu. Learning Efficient Illumination Multiplexing for Joint Capture of Re- flectance and Shape. ACM Transactions on Graphics (TOG), 38(6), 2019. 3

  12. [20]

    Modular Primitives for High-Performance Differentiable Rendering

    Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular Primitives for High-Performance Differentiable Rendering. ACM Trans- actions on Graphics (TOG), 39(6), 2020. 3

  13. [21]

    Learning to Recon- struct Shape and Spatially-Varying Reflectance from a Single Image

    Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Learning to Recon- struct Shape and Spatially-Varying Reflectance from a Single Image. ACM Transactions on Graphics (TOG), 37(6), 2018. 2

  14. [22]

    ENVIDR: Implicit Differentiable Renderer with Neural Environment Lighting

    Ruofan Liang, Huiting Chen, Chunlin Li, Fan Chen, Sel- vakumar Panneer, and Nandita Vijaykumar. ENVIDR: Implicit Differentiable Renderer with Neural Environment Lighting. In IEEE International Conference on Computer Vision (ICCV), 2023. 2

  15. [23]

    Neural V olumes: Learning Dynamic Renderable V olumes from Im- ages

    Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural V olumes: Learning Dynamic Renderable V olumes from Im- ages. ACM Transactions on Graphics (TOG) , 38(4), 2019. 2

  16. [24]

    Smooth Subdivision Surfaces Based on Tri- angles

    Charles Loop. Smooth Subdivision Surfaces Based on Tri- angles. 1987. 4

  17. [25]

    Unified Shape and SVBRDF Recovery using Differentiable Monte Carlo Rendering

    Fujun Luan, Shuang Zhao, Kavita Bala, and Zhao Dong. Unified Shape and SVBRDF Recovery using Differentiable Monte Carlo Rendering. Computer Graphics Forum (CGF), 40(4), 2021. 2, 6

  18. [26]

    Neural Microfacet Fields for Inverse Ren- dering

    Alexander Mai, Dor Verbin, Falko Kuester, and Sara Fridovich-Keil. Neural Microfacet Fields for Inverse Ren- dering. In IEEE International Conference on Computer Vi- sion (ICCV), 2023. 2

  19. [27]

    Practical Physically-Based Shading in Film and Game Pro- duction

    Stephen McAuley, Stephen Hill, Naty Hoffman, Yoshiharu Gotanda, Brian Smits, Brent Burley, and Adam Martinez. Practical Physically-Based Shading in Film and Game Pro- duction. In ACM SIGGRAPH 2012 Courses. 2012. 8

  20. [28]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In European Conference on Computer Vision (ECCV), 2020. 1, 2, 8

  21. [29]

    Extracting Triangular 3D Models, Materials, and Light- ing From Images

    Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas M¨uller, and Sanja Fi- dler. Extracting Triangular 3D Models, Materials, and Light- ing From Images. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3

  22. [30]

    Extracting Triangular 3D Models, Materials, and Light- ing From Images

    Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas M¨uller, and Sanja Fi- dler. Extracting Triangular 3D Models, Materials, and Light- ing From Images. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognitio...

  23. [31]

    Giljoo Nam, Joo Ho Lee, Diego Gutierrez, and Min H. Kim. Practical SVBRDF Acquisition of 3D Objects with Unstruc- tured Flash Photography. ACM Transactions on Graphics (TOG), 37(6), 2018. 2

  24. [32]

    Large Steps in Inverse Rendering of Geometry

    Baptiste Nicolet, Alec Jacobson, and Wenzel Jakob. Large Steps in Inverse Rendering of Geometry. ACM Transactions on Graphics (TOG), 40(6), 2021. 4

  25. [33]

    Differentiable V olumetric Rendering: Learning Implicit 3D Representations without 3D Supervi- sion

    Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable V olumetric Rendering: Learning Implicit 3D Representations without 3D Supervi- sion. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2020. 2

  26. [34]

    UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction

    Michael Oechsle, Songyou Peng, and Andreas Geiger. UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction. InIEEE International Conference on Computer Vision (ICCV), 2021. 2

  27. [35]

    Continuous remeshing for inverse render- ing

    Werner Palfinger. Continuous remeshing for inverse render- ing. Computer Animation and Virtual Worlds, 33(5), 2022. 6

  28. [36]

    Revisiting Point Cloud Simplification: A Learnable Feature Preserving Approach

    Rolandos Alexandros Potamias, Giorgos Bouritsas, and Ste- fanos Zafeiriou. Revisiting Point Cloud Simplification: A Learnable Feature Preserving Approach. In European Con- ference on Computer Vision (ECCV). Springer, 2022. 1

  29. [37]

    Neural Mesh Simplification

    Rolandos Alexandros Potamias, Stylianos Ploumpis, and Stefanos Zafeiriou. Neural Mesh Simplification. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1

  30. [38]

    Multi-resolution 3D ap- proximations for rendering complex scenes

    Jarek Rossignac and Paul Borrel. Multi-resolution 3D ap- proximations for rendering complex scenes. In Modeling in Computer Graphics: Methods and Applications . Springer,

  31. [39]

    Towards Scalable Multi-View Recon- struction of Geometry and Materials

    Carolin Schmitt, Bo ˇzidar Anti´c, Andrei Neculai, Joo Ho Lee, and Andreas Geiger. Towards Scalable Multi-View Recon- struction of Geometry and Materials. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023. 3

  32. [40]

    On Joint Estimation of Pose, Geometry and svBRDF from a Handheld Scanner

    Carolin Schmitt, Simon Donne, Gernot Riegler, Vladlen Koltun, and Andreas Geiger. On Joint Estimation of Pose, Geometry and svBRDF from a Handheld Scanner. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3

  33. [41]

    Scene Representation Networks: Continuous 3D- Structure-Aware Neural Scene Representations

    Vincent Sitzmann, Michael Zollh ¨ofer, and Gordon Wet- zstein. Scene Representation Networks: Continuous 3D- Structure-Aware Neural Scene Representations. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019. 2

  34. [42]

    NeRV: Neural Reflectance and Visibility Fields for Relight- ing and View Synthesis

    Pratul P Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T Barron. NeRV: Neural Reflectance and Visibility Fields for Relight- ing and View Synthesis. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2021. 1, 2

  35. [43]

    Neural-PBIR Reconstruction of Shape, Mate- rial, and Illumination

    Cheng Sun, Guangyan Cai, Zhengqin Li, Kai Yan, Cheng Zhang, Carl Marshall, Jia-Bin Huang, Shuang Zhao, and Zhao Dong. Neural-PBIR Reconstruction of Shape, Mate- rial, and Illumination. In IEEE International Conference on Computer Vision (ICCV), 2023. 2

  36. [44]

    Delicate Tex- tured Mesh Recovery from NeRF via Adaptive Surface Re- finement

    Jiaxiang Tang, Hang Zhou, Xiaokang Chen, Tianshu Hu, Er- rui Ding, Jingdong Wang, and Gang Zeng. Delicate Tex- tured Mesh Recovery from NeRF via Adaptive Surface Re- finement. In IEEE International Conference on Computer Vision (ICCV), 2023. 2

  37. [45]

    De- ferred Neural Rendering: Image Synthesis Using Neural Textures

    Justus Thies, Michael Zollh ¨ofer, and Matthias Nießner. De- ferred Neural Rendering: Image Synthesis Using Neural Textures. ACM Transactions on Graphics (TOG) , 38(4),

  38. [46]

    Differ- entiable Signed Distance Function Rendering

    Delio Vicini, S ´ebastien Speierer, and Wenzel Jakob. Differ- entiable Signed Distance Function Rendering. ACM Trans- actions on Graphics (TOG), 41(4), 2022. 2

  39. [47]

    Marschner, Hongsong Li, and Ken- neth E

    Bruce Walter, Stephen R. Marschner, Hongsong Li, and Ken- neth E. Torrance. Microfacet Models for Refraction through Rough Surfaces. Rendering Techniques, 2007, 2007. 3

  40. [48]

    NeuS: Learning Neural Im- plicit Surfaces by V olume Rendering for Multi-view Recon- struction

    Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. NeuS: Learning Neural Im- plicit Surfaces by V olume Rendering for Multi-view Recon- struction. Advanced Neural Information Processing Systems (NeurIPS), 34, 2021. 2

  41. [49]

    Image quality assessment: from error visibility to structural similarity

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Process- ing (TIP), 13(4), 2004. 6

  42. [50]

    NeFII: Inverse Rendering for Reflectance Decomposition with Near-Field Indirect Illumi- nation

    Haoqian Wu, Zhipeng Hu, Lincheng Li, Yongqiang Zhang, Changjie Fan, and Xin Yu. NeFII: Inverse Rendering for Reflectance Decomposition with Near-Field Indirect Illumi- nation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2

  43. [51]

    NeuTex: Neu- ral Texture Mapping for V olumetric Neural Rendering

    Fanbo Xiang, Zexiang Xu, Milos Hasan, Yannick Hold- Geoffroy, Kalyan Sunkavalli, and Hao Su. NeuTex: Neu- ral Texture Mapping for V olumetric Neural Rendering. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2

  44. [52]

    NeuMesh: Learning Disentangled Neural Mesh-based Implicit Field for Geometry and Texture Editing

    Bangbang Yang, Chong Bao, Junyi Zeng, Hujun Bao, Yinda Zhang, Zhaopeng Cui, and Guofeng Zhang. NeuMesh: Learning Disentangled Neural Mesh-based Implicit Field for Geometry and Texture Editing. In European Conference on Computer Vision (ECCV). Springer, 2022. 2

  45. [53]

    Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, and Kwan-Yee K. Wong. PS-NeRF: Neural Inverse Rendering for Multi-view Photometric Stereo. In European Conference on Computer Vision (ECCV), 2022. 2

  46. [54]

    V ol- ume Rendering of Neural Implicit Surfaces

    Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. V ol- ume Rendering of Neural Implicit Surfaces. Advanced Neu- ral Information Processing Systems (NeurIPS), 34, 2021. 2

  47. [55]

    Multiview Neural Surface Reconstruction by Disentangling Geometry and Ap- pearance

    Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview Neural Surface Reconstruction by Disentangling Geometry and Ap- pearance. Advanced Neural Information Processing Systems (NeurIPS), 33, 2020. 1, 2

  48. [56]

    IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images

    Kai Zhang, Fujun Luan, Zhengqi Li, and Noah Snavely. IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images. In IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  49. [57]

    PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relight- ing

    Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relight- ing. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2021. 3

  50. [58]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 6

  51. [59]

    Freeman, and Jonathan T

    Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul De- bevec, William T. Freeman, and Jonathan T. Barron. NeR- Factor: Neural Factorization of Shape and Reflectance Under an Unknown Illumination. ACM Transactions on Graphics (TOG), 40(6), 2021. 1, 2

  52. [60]

    NeMF: Inverse V olume Rendering with Neural Microflake Field

    Youjia Zhang, Teng Xu, Junqing Yu, Yuteng Ye, Yanqing Jing, Junle Wang, Jingyi Yu, and Wei Yang. NeMF: Inverse V olume Rendering with Neural Microflake Field. In IEEE International Conference on Computer Vision (ICCV), 2023. 2

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.