{"id":"7d1e1ce7-3c37-455a-b8e5-b970d5a736cc","arxiv_id":"2501.06431","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Aug3D augments large outdoor datasets with synthetic views rendered from SfM reconstructions, improving PixelNeRF's PSNR from 20.03 to 21.80 on UrbanScene3D Campus, though reducing cluster size alone reached 22.94.","lead":"This paper proposes Aug3D, a method that reconstructs large outdoor drone scenes and renders extra synthetic views to help train generalizable novel-view-synthesis models like PixelNeRF. The authors report modest PSNR gains, but their main comparison omits their own stronger real-data baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own reported cluster-size-10 baseline (PSNR 22.94) is omitted from Table II, and it beats the proposed Aug3D Semantic result (21.80); the central claim that Aug3D improves real-data GNVS performance is therefore unsupported by the paper's numbers.","rationale":"I read the paper as aiming to show that synthetic novel views generated by Aug3D improve a feed-forward NVS model on real outdoor data. For that central claim to hold, the evaluation must compare against the strongest real-data curation baseline. The paper itself establishes that SfM shared grouping with cluster size 10 gives PSNR 22.94, yet Table II only includes the cluster-size-20 baseline of 20.03 when evaluating Aug3D. This is an internal inconsistency, not merely a disagreement with external consensus: the authors' own numbers show the simpler baseline beating their augmentation by 1.14 dB. The synthetic-only results are also not a valid generalization test because the test renders come from the same reconstruction used to generate the synthetic training views. The absence of error bars and the varying GPU setups further weaken confidence in the small reported differences. The underlying idea is plausible and worth testing, but as presented the central claim is not established; the most direct fix is a fair comparison against the cluster-size-10 baseline and, ideally, a combination of Aug3D with cluster size 10.","tokens_in":10568,"tokens_out":3080,"duration_ms":25670,"concrete_test":"Re-run the real-data PixelNeRF experiments with the SfM shared grouping and cluster size 10 (the paper's own best configuration) on the same Campus test split used for Table II, and report average and per-cluster PSNR. If the cluster-size-10 result remains at or above 22.94, then Aug3D Semantic at 21.80 does not improve on the real-data baseline. Additionally, train PixelNeRF on Aug3D plus real data with cluster size 10; if this does not exceed both 22.94 and 21.80, the claimed additive benefit of augmentation is not supported.","verdict_should_be":"REJECT","load_bearing_attack":"Section V reports that reducing cluster size from 20 to 10 with SfM shared grouping improves best PSNR to 22.94, yet Table II compares Aug3D only against the cluster-size-20 baseline (20.03). Under the paper's own numbers, the simple cluster-size reduction outperforms Aug3D Semantic (21.80) by 1.14 dB, so the statement 'These results validate the effectiveness of the Aug3D dataset in enhancing GNVS performance' does not follow. The central contribution—that synthetic augmented views improve real-data generalization—requires comparison against the best real-data configuration, ideally with cluster size 10, and a demonstration that augmentation adds to it. A secondary issue: the synthetic-only scores (29.12, 28.79) are measured on renders from the same SfM/MVS reconstruction used to create the synthetic training views, so they reflect reconstruction consistency rather than generalization to real test views. No error bars or standard deviations are reported, and varying compute setups across experiments make it hard to rule out training noise. These gaps, not the plausibility of the idea, are what block the conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses generalizable novel view synthesis (GNVS) on large-scale outdoor scenes by curating training clusters from the UrbanScene3D dataset and proposing Aug3D, a reconstruction-based augmentation method that generates synthetic views via multiscale grid or semantic plane-fitting sampling. The authors train PixelNeRF and compare four clustering strategies, finding SfM shared-point grouping to be the best. They report that reducing the cluster size from 20 to 10 images improves PSNR from 20.03 to 22.94. They then claim that augmenting the real dataset with synthetic views achieves a best PSNR of 21.80, surpassing the real-data baseline, and conclude that this validates Aug3D's effectiveness in enhancing GNVS performance.","tokens_in":10717,"tokens_out":3865,"duration_ms":35870,"significance":"If properly validated, a data curation and augmentation pipeline for outdoor GNVS would be valuable, as feed-forward NVS models are typically limited to small object-centric scenes. The systematic comparison of clustering methods and the idea of using reconstructed scenes to generate well-conditioned novel views are interesting and potentially useful. However, the current evidence does not support the central claim; the paper's own numbers contradict it, and the synthetic evaluation is circular.","major_comments":[{"comment":"The claimed validation of Aug3D is unsupported because Table II compares Aug3D (best PSNR 21.80) only against the cluster-size-20 real baseline (20.03), while Section V itself reports that reducing the cluster size from 20 to 10 improves PSNR to 22.94. Under the paper's own numbers, the simple cluster-size reduction outperforms Aug3D Semantic by 1.14 dB. The statement 'These results validate the effectiveness of the Aug3D dataset in enhancing GNVS performance' therefore does not follow. The authors must compare against the best real-data configuration (cluster size 10) and, ideally, include an ablation in which Aug3D is added to that configuration.","section":"Section V, Table II"},{"comment":"The synthetic-only PSNR values (29.12 for Grid Sampling, 28.79 for Semantic Plane Fitting) are evaluated on renders from the same SfM/MVS reconstruction that was used to generate the synthetic training views. This evaluation is circular: the model is trained and tested on views derived from the same mesh, so the high PSNR reflects reconstruction consistency rather than generalization to real novel views. The approximately 8 dB gap between synthetic-only and real-data PSNR indicates a substantial domain gap. The authors should evaluate models trained on synthetic data against held-out real views or on a different scene to demonstrate generalization.","section":"Section V, Table II, Synthetic Dataset rows"},{"comment":"No error bars, standard deviations, or multiple-seed results are reported. The compute setup varies across experiments: the real-dataset experiments use two 32GB Tesla V100 GPUs, the Grid-based augmentation uses a single 24GB RTX 3090 Ti, and all other experiments use a 10GB RTX 3080. The difference between Aug3D Grid (21.67) and Aug3D Semantic (21.80) is only 0.13 dB, and without variance estimates or fixed hardware, training noise cannot be ruled out as an explanation. Report the mean and standard deviation over at least three independent runs on identical hardware.","section":"Section IV, Compute Setup and Table II"}],"minor_comments":[{"comment":"The phrase 'scene sentric dome sampling' appears to contain a typo; it should read 'scene-centric dome sampling.'","section":"Section III-B"},{"comment":"The paper states that PixelNeRF is run with '256 hidden layers'; this is likely a typo for '256 hidden units' or 'hidden features,' since PixelNeRF's architecture uses fully connected layers with 256 hidden units.","section":"Section IV, Dataset and Metric"},{"comment":"The column labeled 'Low PSNR' in Table I is not defined; please clarify whether it refers to the minimum PSNR across test views, the worst cluster, or some other quantity.","section":"Section V, Table I"},{"comment":"Section IV says 'we focus exclusively on the Campus scene from the UrbanScene3D dataset,' but Figure 7 in the Appendix reports qualitative results on a 'Residence scene.' Please clarify whether quantitative results also exist for that scene or remove the inconsistency.","section":"Section IV and Appendix B"},{"comment":"The abstract reports that reducing the cluster size from 20 to 10 'improves PSNR by 10%,' but the text gives values 20.03 and 22.94, which correspond to a relative improvement of about 14.5%. Please make the percentage calculation consistent.","section":"Abstract and Section V"}],"recommendation":"reject","confidential_remarks":"The central quantitative claim is internally contradicted by the paper's own reported cluster-size-10 baseline (22.94 dB), which beats the proposed Aug3D result (21.80 dB). This is not a presentation issue but a load-bearing flaw in the evaluation. The synthetic-only evaluation is also circular, and the lack of error bars makes the small reported differences uninterpretable. I recommend rejection, though a substantially revised version with proper baselines, non-circular evaluation, and variance reporting might be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2501.06431. First, the underlying idea is sensible and, as far as the cited literature goes, new: take a large outdoor drone dataset, cluster images by SfM shared points, reconstruct the scene, sample new object-centric views via multiscale grids or semantic plane fitting, and train PixelNeRF on real plus synthetic views. That is a reasonable recipe for making city-scale data usable by feed-forward NVS models, and the qualitative clustering comparison in Fig. 4 is convincing evidence that SfM shared grouping gives better overlap than the other three heuristics. Second, the central quantitative claim is not supported by the paper's own numbers: Section V says reducing cluster size from 20 to 10 improves PSNR to 22.94, but Table II compares Aug3D Semantic (21.80) only against the cluster-20 baseline (20.03). The simple clustering change beats the proposed augmentation by 1.14 dB. That omission is hard to read as an oversight, and the sentence 'These results validate the effectiveness of the Aug3D dataset' does not follow.\n\nThe synthetic-only numbers (29.12, 28.79) are also self-referential: they evaluate on renders from the same SfM/MVS mesh used to create the synthetic training set, so they measure reconstruction consistency, not generalization. No error bars or standard deviations are reported, and the compute setup differs across experiments (V100 vs 3090 Ti vs 3080), which makes it hard to rule out training noise. The paper releases no code or data, so independent reproduction of even the clustering baseline would require re-implementing the pipeline.\n\nWhat the paper does well: the clustering study is a useful empirical comparison, the multiscale grid and semantic sampling are concrete and described clearly, and the discussion of why raw UrbanScene3D is badly conditioned for feed-forward NVS is honest. The citation pattern is fine. The flaw is not in the idea; it is in the evaluation design. The fix is straightforward: compare against cluster-10, report mean and variance over runs, and evaluate synthetic-trained models on held-out real test views. If the augmentation adds to cluster-10, the claim stands; if not, the paper becomes a clustering study, which is still publishable as a smaller contribution.\n\nWho is this for? People working on generalizable NVS for large scenes, especially anyone who has tried to train PixelNeRF on drone data. It deserves a serious referee: the idea is plausible, the problem is real, and the missing baseline is correctable. I would not cite it yet, but I would engage with a revised version.","headline":"Plausible augmentation idea undermined by an omitted baseline: the paper's own cluster-10 result beats the proposed method.","tokens_in":11377,"tokens_out":1449,"would_cite":false,"duration_ms":94396,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adding synthetic novel views rendered from a reconstruction of a large outdoor scene improves how well a feed-forward neural network predicts new views, with semantic sampling (21.80 PSNR) edging out grid sampling…","keywords":["novel view synthesis","generalizable NeRF","data augmentation","structure-from-motion","outdoor scenes","semantic sampling","UrbanScene3D","PixelNeRF"],"falsifier":"Train PixelNeRF on the real Campus dataset with cluster size 10 (which the paper reports gives 22.94 PSNR) and compare it against the same model trained on real data plus Aug3D synthetic views under identical evaluation; if the augmented model does not exceed 22.94 PSNR on real held-out views, the augmentation's claimed benefit is falsified. Additionally, measuring real-image PSNR of a model trained only on synthetic views would directly quantify the domain gap implied by the 29.12 synthetic-only result.","tokens_in":10253,"feed_emoji":"🏙️","tokens_out":4708,"duration_ms":43373,"temperature":0.7,"pith_summary":"The paper addresses the problem of training generalizable novel view synthesis models on large-scale outdoor scenes captured by drones, where consecutive images share little overlap. It proposes clustering images by shared structure-from-motion points, which it shows is far more effective than sequence-, grid-, or ray-based grouping, and reports that reducing cluster size from 20 to 10 images improves PSNR from 20.03 to 22.94. To further help training, it introduces Aug3D, an augmentation pipeline that reconstructs the scene and renders synthetic novel views from virtual cameras placed by multiscale grid sampling or semantic building sampling. When these synthetic views are combined with real data, PixelNeRF reaches 21.80 PSNR with semantic sampling and 21.67 with grid sampling, which the paper presents as validating the augmentation's effectiveness.","feed_headline":"Adding synthetic views lifts outdoor NVS to 21.80 PSNR","feed_subtitle":"Semantic-sampled synthetic views beat grid in mixed training (21.80 vs 21.67), yet a 10-view real baseline already reaches 22.94.","key_machinery":"The central objects are the clustering metric and the two sampling strategies. The clustering metric is an SfM shared-point similarity matrix: cameras observing the same structures share many matched points, so top-K neighbors by this similarity form coherent clusters. The augmentation strategies are Multiscale Grid Sampling, which places virtual domes over dynamically sized grid cells, and Semantic Building Sampling, which fits a plane to the top percentile of points by height, renders a top-down mask, extracts bounding boxes, and merges nearby boxes to place domes preferentially over urban regions. These domes sample synthetic camera poses on the reconstructed mesh, and the rendered views are added to the real training set plus PixelNeRF, a feed-forward NeRF conditioned on pixel-aligned features.","core_discovery":"On the paper's own terms, the central discovery is that a data curation and augmentation pipeline can make feed-forward NeRF models viable on large outdoor scenes. The paper reports four clustering strategies tested on the UrbanScene3D Campus scene, finding that grouping images by shared SfM points yields the best PixelNeRF performance (best PSNR 20.03, average 14.6), far above sequence grouping (9.7), grid grouping (12.2), and ray-intersection grouping (13.6). It further reports that shrinking the cluster size from 20 to 10 images raises best PSNR to 22.94. The paper's proposed Aug3D augmentation renders synthetic views from a reconstructed mesh using either multiscale grid sampling or semantic plane-fitting sampling; synthetic-only training reaches 29.12 and 28.79 PSNR respectively, and mixing these synthetic views with the real dataset yields best PSNR of 21.80 (semantic) and 21.67 (grid), slightly surpassing the cluster-size-20 real baseline of 20.03.","pith_inferences":["The paper's claim that Aug3D 'enhances' GNVS performance would be stronger if benchmarked against the best real-data baseline: the cluster-size-10 baseline (22.94 PSNR) already exceeds the best augmented result (21.80 PSNR), so the marginal benefit of synthetic views on real data is not established by the presented comparison.","The large gap between synthetic-only PSNR (29.12) and real-data PSNR (around 20 to 21) suggests a substantial domain gap; a natural test is whether synthetic pretraining followed by fine-tuning on real data narrows that gap.","The semantic plane-fitting sampler could be extended to other semantic classes (roads, vegetation) or combined with more robust building detectors; the paper notes that SAM-based detection was shadow-sensitive, so better segmentation would likely improve view diversity.","The key clustering insight—that shared SfM points define coherence—could be combined with the augmentation strategy for other feed-forward models and should be tested on held-out real captures from unseen drone trajectories."],"forward_implications":["Reducing cluster size from 20 to 10 images improves PSNR by roughly 10 percent, indicating that high view overlap within input clusters is a key factor for feed-forward NVS on outdoor scenes.","Semantic sampling around urban regions outperforms uniform grid sampling when synthetic views are mixed with real data, suggesting that directing augmentation toward underrepresented scene content helps.","The SfM shared-point grouping method can be applied to any feed-forward NVS model that expects DTU-like clustered inputs, not only PixelNeRF.","The full pipeline trained successfully on the UrbanScene3D Campus scene, implying it could extend to other large outdoor datasets captured in similar drone grid patterns, such as Mill-19."],"supporting_citations":[{"why":"Supplies the PixelNeRF feed-forward NVS model that is trained and evaluated throughout the paper.","marker":"[38]"},{"why":"Provides the UrbanScene3D dataset, specifically the Campus scene used for all experiments.","marker":"[16]"},{"why":"Motivates large-scale outdoor scene reconstruction and the drone grid-scan capture pattern that creates low-overlap challenges.","marker":"[30]"},{"why":"Defines the DTU-style object-centric input format that the clustering methods aim to replicate for compatibility with NVS models.","marker":"[11]"},{"why":"Documents urban drone capture characteristics and the need for high overlap, informing the sampling strategies.","marker":"[21]"},{"why":"Represents recent large-scale 3D reconstruction advances that support the feasibility of using reconstructed meshes for synthetic view generation.","marker":"[15]"}],"fun_headline_variants":["Aug3D: synthetic views boost outdoor novel view synthesis","Outdoor NVS improved via Aug3D semantic sampling","Aug3D lifts feed-forward NVS on large outdoor scenes","Synthetic views from SfM scenes aid outdoor NVS training","Aug3D: grid and semantic sampling for outdoor NVS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that synthetic views rendered from an SfM/MVS reconstruction of the same scene are a valid proxy for real novel-view generalization, such that adding them to the real training set improves real-data performance; the presented evidence only compares against a weaker cluster-size-20 baseline, not the stronger cluster-size-10 real baseline.","fun_headline_variants_meta":{"raw":{"variants":["Aug3D: synthetic views boost outdoor novel view synthesis","Outdoor NVS improved via Aug3D semantic sampling","Aug3D lifts feed-forward NVS on large outdoor scenes","Synthetic views from SfM scenes aid outdoor NVS training","Aug3D: grid and semantic sampling for outdoor NVS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1450,"prompt_tokens":979,"completion_tokens":471,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":386}},"tokens_in":595,"tokens_out":471,"duration_ms":5458,"temperature":1.0,"reasoning_tokens":386,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:00:00.618962+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train PixelNeRF on the real Campus dataset with cluster size 10 (which the paper reports gives 22.94 PSNR) and compare it against the same model trained on real data plus Aug3D synthetic views under identical evaluation; if the augmented model does not exceed 22.94 PSNR on real held-out views, the augmentation's claimed benefit is falsified. Additionally, measuring real-image PSNR of a model trained only on synthetic views would directly quantify the domain gap implied by the 29.12 synthetic-only result.","supporting_citations":[{"cited_title":"In: CVPR (2021)","cited_arxiv_id":null,"evidence_quote":"Supplies the PixelNeRF feed-forward NVS model that is trained and evaluated throughout the paper."},{"cited_title":"In: 2014 IEEE Conference on Computer Vision and Pattern Recognition","cited_arxiv_id":null,"evidence_quote":"Defines the DTU-style object-centric input format that the clustering methods aim to replicate for compatibility with NVS models."},{"cited_title":"In: CVPR (2024)","cited_arxiv_id":null,"evidence_quote":"Represents recent large-scale 3D reconstruction advances that support the feasibility of using reconstructed meshes for synthetic view generation."}],"review_version":1}