REVIEW 4 major objections 6 minor 2 cited by
Benchmarking Image Similarity Metrics for Novel View Synthesis Applications
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that DreamSim, a learned perceptual similarity metric, should replace SSIM, PSNR, and LPIPS for scoring novel-view-synthesis renders because it degrades gradually with corruption severity and ignores pixel changes humans…
desk verdict Real benchmark results with a load-bearing external-validity gap; the DreamSim-for-NVS recommendation outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a controlled corruption benchmark, called NVS Corruptions, that applies twelve artifact types (blur, brightness, color shift, contrast, floaters, grayscale, pixelation, rotation, saturation, shadows, splats, warp) at twenty severity levels to building-scene images, together with a 750k-image ImageNet-C subset and localized foreground/background and pixel-crop corruptions. All metric outputs are normalized to a common 0-to-1 scale, and the paper reads the resulting median score-versus-severity curves for monotonic decline and discriminative power. DreamSim, a learned perceptual similarity model trained on human-judged similarity pairs, is the object whose behavior the curves are meant to validate.
What would settle it
Run a study in which humans rate real NVS renders from a NeRF or 3D Gaussian Splatting model for utility, then compare how well DreamSim, SSIM, PSNR, and LPIPS correlate with those ratings; the paper's recommendation fails if a classical metric correlates as well as or better than DreamSim, or if DreamSim's score-severity curves on the synthetic corruption suite diverge from its curves on the real artifact set.
Extended reading notes
Core claim
The paper claims that DreamSim is robust in scoring NVS renders in a predictable and differentiable manner and is less sensitive to corruptions that are imperceptible to humans. The experimental evidence is a set of normalized score-severity curves: on the NVS Scenes building imagery, SSIM, PSNR, and LPIPS collapse after minimal blur or other corruption and then stay flat, while DreamSim's median score descends steadily across severity levels and separates major intensity increments. On 750k ImageNet-C images, DreamSim keeps this monotone response on ten of twelve corruption types, with the two exceptions (fog and elastic transform) explained by the corruptions themselves having little perceptual change or introducing rotation. In the local-corruption experiments, DreamSim alone is nearly unaffected by cropping of 1 to 10 pixels and shows stronger penalties for foreground than background corruption, which the paper interprets as agreement with human visual priorities. The stated conclusion is that DreamSim or a future NVS-trained version should be used for general NVS evaluation because it quantifies render alignment with human judgment.
Load-bearing premise
The load-bearing premise is that the twelve artificial corruptions faithfully stand in for artifacts real NVS models produce and that score changes under those corruptions represent how useful humans would find the rendered image.
Editorial extensions
If this is right
- NVS model rankings produced with SSIM, PSNR, or LPIPS could change substantially when scored with DreamSim, since the classical metrics conflate imperceptible artifacts with real quality loss.
- Evaluation of real-world NVS deployments, such as site modeling or first-responder imagery, should weight scene-level content preservation rather than per-pixel fidelity.
- Small camera pose errors that shift edges by a few pixels will no longer dominate quality scores, so pose-estimation failures will be judged by their actual impact on perceived content.
- A DreamSim variant fine-tuned on NVS-specific render artifacts would be the natural next evaluation tool and would likely sharpen the already monotone response to corruption.
Reading between the lines
- The paper treats its synthetic corruptions as stand-ins for real NVS artifacts, but it does not measure metric behavior on actual NeRF or 3D Gaussian Splatting renders; that comparison is the direct test the study leaves open.
- Because the foreground/background asymmetry reflects DreamSim's training data, applications where background content matters (e.g., site modeling) may need a differently calibrated metric rather than a universally preferred one.
- The human-utility interpretation is inferred from perceptual plausibility rather than directly measured; a human-rating study on the same corrupted corpus would settle whether DreamSim's monotonic curves correspond to task usefulness.
- If DreamSim responds monotonically to synthetic severity while classical metrics do not, then using DreamSim could also serve as a cheap automated corruption-robustness monitor for NVS systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper benchmarks four image similarity metrics (SSIM, PSNR, LPIPS, and DreamSim) for use in novel view synthesis (NVS) evaluation. The authors construct two corrupted-image corpora: a new 'NVS Scenes' set with 12 custom corruption types at 20 severity levels, and a 750k-image subset of ImageNet-C with its standard 15 corruptions and 5 severities. They also apply localized foreground/background corruptions and small pixel-level cropping to ImageNet images, then plot median normalized metric scores across severity levels. Based on the resulting score curves, the paper concludes that DreamSim is more robust to minor corruptions, more discriminative across severity levels, and better aligned with human judgment than SSIM, PSNR, and LPIPS, and recommends DreamSim for general NVS evaluation.
Significance. If fully validated, the paper would provide a large-scale empirical characterization of how four image similarity metrics respond to synthetic corruptions, and would support a specific recommendation for NVS evaluation. The scale of the ImageNet-C experiment (750k variants) and the use of segmentation to separate foreground and background corruptions are notable strengths. However, the paper's central recommendation rests on two externally oriented claims that are not directly tested: that the synthetic corruptions faithfully represent artifacts produced by real NVS models, and that the observed score trends correspond to human perceptual judgment. As presented, the work is a useful internal comparison of metric behaviors under controlled distortions, but the human-alignment and real-NVS applicability conclusions are unsubstantiated.
major comments (4)
- [Section 3.1 and Section 5] The paper's central recommendation—that 'DreamSim or a future NVS-specifically trained version of DreamSim should be used for general evaluation of NVS models'—rests on the assertion (Section 3.1) that the 12 NVS Corruptions 'simulate common corruptions typically seen in NVS model outputs.' The text provides example real artifacts in Figure 2 but never quantitatively compares the distribution of those artifacts with the synthetic corruption suite (e.g., type frequencies, severity ranges, spatial extent). Without such validation, the benchmark measures metric behavior under hand-chosen distortions, not specifically under NVS render artifacts. Please either add a validation experiment on real NeRF/3DGS renders (e.g., apply the metrics to paired real renders and categorize their artifacts) or considerably temper the claim that the results transfer to NVS outputs.
- [Sections 4.1, 4.3, 4.4] The paper repeatedly claims that DreamSim's score curves 'align with human judgement' (Section 4.1), that its foreground/background sensitivity 'align[s] more closely with human judgement' (Section 4.3), and that small pixel crops are 'imperceptible' and do not affect 'scene understanding' (Sections 3.4 and 4.4). No human opinion scores are collected or cited for this specific benchmark, and DreamSim's training on human similarity judgments (reference [2]) does not by itself establish that its monotonic curves track human utility for NVS renders; this is a partial circularity. Please add a small human-rating study on a subset of corrupted and cropped images, or compare against an established subjective NVS quality dataset such as references [6] or [9], to support the human-alignment claims.
- [Section 3.3, Figures 9, 12, 13] The paper presents median score curves without any error bars, confidence intervals, or per-condition sample sizes. The NVS Scenes dataset has only 24 images, and the foreground/background experiment is limited to approximately 500 images due to manual mask validation. Because the central claims about discriminative power and relative robustness rely on the shapes of these median curves, the absence of variance measures makes it impossible to assess whether observed differences (e.g., DreamSim versus LPIPS plateau behavior, or the foreground/background gap) are statistically meaningful. Please report per-condition quantiles or bootstrap confidence intervals, and state the number of images per condition for each figure.
- [Section 3.2 and Section 4.4] The phrase 'imperceptible to the human eye' is used to describe pixel crops of 5–10 pixels (Section 3.2) and the paper later refers to 'small, imperceptible artifacts' (Section 4.4). Whether a given crop is perceptually invisible depends on image resolution, viewing distance, and content; this is asserted rather than measured. Without a perceptual verification (e.g., two-alternative forced choice with human observers, or a published visibility threshold), the claim that DreamSim is 'less sensitive to corruptions imperceptible by humans' conflates small pixel shifts with imperceptibility. Please either verify imperceptibility empirically or rephrase the claim to refer to 'small pixel crops' rather than 'imperceptible corruptions.'
minor comments (6)
- [Section 4.1] The text states that fog 'has little perceptual change across severity levels, depicted in Figure 10,' but Figure 10 shows pixelation and Figure 11 shows fog. Please correct the cross-reference.
- [Figure 9 caption] The caption says 'our NVS rendered dataset' while the text and Table 1 refer to the 'NVS Scenes' dataset. Use consistent terminology throughout.
- [Table 1] The row for ImageNet-C lists '10k' images but does not explain that this is 200 classes × 50 images per class. Please include this derivation in the text or table note.
- [Figure 13(b) caption] The caption states that 'DreamSim and SSIM are normalized to the same 0 to 1 scale' but does not explain how PSNR and LPIPS are normalized in the same plot. Please clarify the normalization applied to each metric.
- [References] Reference [12] has garbled author names ('Benand Mildenhall,' 'Matthewand Tancik,' 'Raviand Ramamoorthi,' 'Andreaand Vedaldi,' etc.). Please correct these entries.
- [Abstract] The sentence 'traditional metrics are unable to effectively differentiate between images with minor pixel-level changes and those with substantial corruption' is too strong given that the benchmark uses specific synthetic corruptions; it would be more accurate to say 'in the tested corruption scenarios.'
Circularity Check
No significant circularity: the benchmark evaluates metric behavior on synthetic corruptions without fitting any parameter to the conclusions; the human-alignment premise is imported from external prior work, not re-derived from the present paper's outputs.
full rationale
The paper makes no fitted-parameter predictions and contains no equations whose outputs are definitionally equal to their inputs. Its experimental contribution is a set of sensitivity and discriminability measurements (median metric scores vs. corruption severity) on its own NVS Corruptions suite and on ImageNet-C. Those measurements are self-contained: the claim that DreamSim decreases gradually and differentiates severity levels is directly read from Figure 9 and Figure 13 and does not presuppose the conclusion. The recommendation to use DreamSim for NVS evaluation does rely on the premise that DreamSim aligns with human judgment, but that premise is cited to the external DreamSim training work [2] (which collected human similarity ratings) and to human-vision literature [8]. No author of the present paper is an author of those sources, so this is external evidentiary support rather than a self-citation loop. The chief weakness is that the paper never validates that its synthetic corruptions match real NVS artifact distributions or collects human ratings on this benchmark; however, those are unverified assumptions about external validity, not circular derivations. Accordingly, under the specified criteria no step reduces to its own inputs, and the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Artificial corruptions emulate NVS render artifacts
- domain assumption DreamSim's pretrained similarity scores reflect human judgement
- domain assumption Foreground clarity bias is a human-like trait
- domain assumption Normalizing each metric to 0-1 makes scores comparable across metrics
Cite this review
Pith. "Pith review of Benchmarking Image Similarity Metrics for Novel View Synthesis Applications." pith.science (2026). https://pith.science/paper/562MMD2I
@misc{pith2026250612563,
author = {Pith},
title = {Pith review of: Benchmarking Image Similarity Metrics for Novel View Synthesis Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/562MMD2I}},
note = {Machine review of arXiv:2506.12563}
}
read the original abstract
Traditional image similarity metrics are ineffective at evaluating the similarity between a real image of a scene and an artificially generated version of that viewpoint [6, 9, 13, 14]. Our research evaluates the effectiveness of a new, perceptual-based similarity metric, DreamSim [2], and three popular image similarity metrics: Structural Similarity (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Learned Perceptual Image Patch Similarity (LPIPS) [18, 19] in novel view synthesis (NVS) applications. We create a corpus of artificially corrupted images to quantify the sensitivity and discriminative power of each of the image similarity metrics. These tests reveal that traditional metrics are unable to effectively differentiate between images with minor pixel-level changes and those with substantial corruption, whereas DreamSim is more robust to minor defects and can effectively evaluate the high-level similarity of the image. Additionally, our results demonstrate that DreamSim provides a more effective and useful evaluation of render quality, especially for evaluating NVS renders in real-world use cases where slight rendering corruptions are common, but do not affect image utility for human tasks.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 2 Pith papers
-
Walk through Paintings: Egocentric World Models from Internet Priors
A lightweight action-conditioning layer inserted into pre-trained video diffusion models turns them into egocentric world models that follow 3-DoF and 25-DoF action commands and generalize to unseen environments such ...
-
Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning
A layout-aware, alignment-first framework decomposes two PDF versions into typed elements, aligns them, and reports detection, localization, and structure-aware changes, outperforming element-specific baselines on a p...
Reference graph
Works this paper leans on
-
[2]
Dream- sim: Learning new dimensions of human visual similarity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similarity using synthetic data. InAdvances in Neural Information Pro- cessing Systems, volume 36. Curran Associates, Inc., 2023. 1, 3
work page 2023
-
[6]
Hanxue Liang, Tianhao Wu, Param Hanji, Francesco Ban- terle, Hongyun Gao, Rafal Mantiuk, and Cengiz Oztireli. Perceptual quality assessment of nerf and neural view syn- thesis methods for front-facing views, 2023. 1, 2
work page 2023
-
[9]
Nerf view synthesis: Subjective quality assessment and objective metrics evaluation, 2024
Pedro Martin, Antonio Rodrigues, Joao Ascenso, and Maria Paula Queluz. Nerf view synthesis: Subjective quality assessment and objective metrics evaluation, 2024. 1, 2
work page 2024
-
[1]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009. 3
work page 2009
-
[3]
Benchmarking neu- ral network robustness to common corruptions and perturba- tions
Dan Hendrycks and Thomas Dietterich. Benchmarking neu- ral network robustness to common corruptions and perturba- tions. arXiv preprint arXiv:1903.12261, 2019. 3
arXiv 1903
-
[4]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4), 2023. 2
work page 2023
-
[5]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. In 2023 IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , 2023. 4
work page 2023
-
[7]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 4
arXiv 2023
Show all 19 references
-
[8]
Across the planes: Dif- fering impacts of foreground and background information on visual search in scenes
Louisa Man and Monica Castelhano. Across the planes: Dif- fering impacts of foreground and background information on visual search in scenes. Journal of Vision, 18, 2018. 6
2018
-
[10]
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duck- worth. Nerf in the wild: Neural radiance fields for uncon- strained photo collections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2021
-
[11]
lang-segment-anything
Luca Mederios. lang-segment-anything. https : / / github . com / luca - medeiros / lang - segment - anything, 2024. 4
2024
-
[12]
Srinivasan, Matthewand Tan- cik, Jonathan T
Benand Mildenhall, Pratul P. Srinivasan, Matthewand Tan- cik, Jonathan T. Barron, Raviand Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In Andreaand Vedaldi, Horstand Bischof, Thomasand Brox, and Jan-Michael Frahm, editors, Co...
2020
-
[13]
An evalu- ation of quality metrics for neural radiance field
Chibuike Onuoha, Jean Atsumi Flaherty, Shihao Luo, Truong Thu Huong, and Truong Cong Thang. An evalu- ation of quality metrics for neural radiance field. In 2023 IEEE 15th International Conference on Computational In- telligence and Communication Networks (CICN) , 2023. 1, 2
2023
-
[14]
Nerf-nqa: No-reference quality assess- ment for scenes generated by nerf and neural view synthesis methods
Qiang Qu, Hanxue Liang, Xiaoming Chen, Yuk Ying Chung, and Yiran Shen. Nerf-nqa: No-reference quality assess- ment for scenes generated by nerf and neural view synthesis methods. IEEE Transactions on Visualization and Computer Graphics, 30(5):2129–2139, 2024. 1, 2
2024
-
[15]
Schonberger and Jan-Michael Frahm
Johannes L. Schonberger and Jan-Michael Frahm. Structure- from-motion revisited. In Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[16]
Doi- baa-wriva-fy22-02, 2022
Acquisition Services Directorate (AQD) The Department of the Interior (DOI), Interior Business Center (IBC). Doi- baa-wriva-fy22-02, 2022. 2
2022
-
[17]
Benchmarking robustness in neural radiance fields
Chen Wang, Angtian Wang, Junbo Li, Alan Yuille, and Cihang Xie. Benchmarking robustness in neural radiance fields. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) Workshops ,
-
[18]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 1, 2
2004
-
[19]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 1, 2
2018
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.