{"id":"0a3e71aa-829f-48c0-a0ad-c96a0fea920e","arxiv_id":"2606.18826","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"EDoF-NeRF inserts a coded aperture into the NeRF camera model to accept defocused images and synthesize novel views with larger depth of field than standard aperture cameras.","lead":"The paper proposes EDoF-NeRF, a technique that adds a coded aperture to the camera model inside neural radiance fields so that defocused images can still produce sharp novel views with extended depth of field. A smart generalist might read it because the approach targets a basic hardware limit that affects any 3D reconstruction method built from ordinary photographs.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Preservation of spatial frequencies by coded aperture under defocus is asserted but not derived from first principles in the camera model","rationale":"The reader's weakest assumption directly identifies the same un-derived optical property that the NeRF inversion step depends on. Because the review was performed on the abstract, the full-text derivation (if present) could close the gap; the concrete test above would settle whether the assumption is merely stated or actually satisfied by the model.","tokens_in":1653,"tokens_out":349,"duration_ms":12946,"concrete_test":"Extract the exact pupil function and rendering equation from §3 (camera model); recompute the OTF analytically or numerically for at least three defocus distances (in-focus, ±2× Rayleigh) using the reported aperture code; verify whether spatial frequencies above 0.3 cycles/pixel remain above 10% modulation for all three distances. If any distance drops below threshold, the frequency-preservation assumption does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the coded aperture (placed at the pupil) transmits sufficient high-frequency content even for out-of-focus scene points so that the NeRF can invert the measurements to recover a sharp radiance field. The abstract states this preservation occurs, yet the integration into the NeRF volume-rendering integral (presumably replacing the standard pinhole or circular-aperture integral with a coded one) implicitly assumes the optical transfer function remains invertible across depths without explicit derivation of the coded pupil function, its Fourier transform, or the resulting point-spread-function support. If the code only preserves frequencies for a narrow defocus range or specific scene content, the direct-input claim fails.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes EDoF-NeRF, a NeRF variant that incorporates a coded aperture placed at the camera pupil into the forward camera model. This allows direct use of defocused coded images as input while generating novel views with extended depth-of-field, by asserting that the coded aperture preserves spatial frequency content under defocus. Validation is claimed via simulations and real experiments showing superior performance over conventional circular apertures.","tokens_in":1791,"tokens_out":530,"duration_ms":17411,"significance":"If the optical model is shown to be invertible across depths, the approach would address a practical bottleneck in NeRF data capture—the inherent DoF versus light-throughput trade-off—potentially enabling larger apertures without focus stacking. The work introduces a concrete hardware modification (coded aperture) directly into the differentiable rendering pipeline, which is a clear engineering contribution even if the frequency-preservation claim requires further substantiation.","major_comments":[{"comment":"Abstract and camera-model description: the central claim that the coded aperture 'preserving spatial frequency components under defocused conditions' is asserted without derivation of the pupil function, its Fourier transform, or the resulting depth-dependent OTF/PSF. This assumption is load-bearing for the statement that coded images can be fed directly into NeRF to recover a sharp radiance field; without an explicit expression replacing the standard pinhole integral, it is impossible to verify that high-frequency content remains recoverable for out-of-focus points.","section":"Abstract / camera model"},{"comment":"Validation section: the abstract states that superiority is demonstrated 'through simulations and experiments,' yet no error metrics (PSNR, SSIM, depth-range coverage), scene parameters, or comparison baselines (e.g., focus-stacking NeRF or small-aperture capture) are supplied. Without these, the empirical support for the extended-DoF claim cannot be assessed and the result remains unverifiable from the given text.","section":"Validation / experiments"}],"minor_comments":[{"comment":"Notation for the coded pupil function and its integration into the volume-rendering equation should be introduced with an explicit equation number rather than left at the level of prose description.","section":"Method"},{"comment":"The abstract would benefit from a single quantitative statement (e.g., 'X dB PSNR gain over Y mm DoF range') to allow readers to gauge the magnitude of the improvement before reading further.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the optical model and validation. We address each major comment below and will revise the manuscript to improve clarity and verifiability.","responses":[{"response":"We agree that the current presentation would benefit from an explicit derivation. In the revised manuscript we will add a new subsection deriving the pupil function of the coded aperture, its Fourier transform, and the resulting depth-dependent OTF. This will include the modified forward imaging integral that replaces the standard pinhole model and will explicitly show the frequency-preservation property relative to a circular aperture.","revision_made":"yes","referee_comment":"[Abstract / camera model] Abstract and camera-model description: the central claim that the coded aperture 'preserving spatial frequency components under defocused conditions' is asserted without derivation of the pupil function, its Fourier transform, or the resulting depth-dependent OTF/PSF. This assumption is load-bearing for the statement that coded images can be fed directly into NeRF to recover a sharp radiance field; without an explicit expression replacing the standard pinhole integral, it is impossible to verify that high-frequency content remains recoverable for out-of-focus points."},{"response":"We acknowledge that the abstract alone does not contain quantitative details. The full manuscript reports simulation and real-world results; however, to address the concern we will expand the experiments section with a summary table of PSNR, SSIM, and depth-range metrics, explicit scene parameters, and direct numerical comparisons against focus-stacking NeRF and small-aperture baselines.","revision_made":"yes","referee_comment":"[Validation / experiments] Validation section: the abstract states that superiority is demonstrated 'through simulations and experiments,' yet no error metrics (PSNR, SSIM, depth-range coverage), scene parameters, or comparison baselines (e.g., focus-stacking NeRF or small-aperture capture) are supplied. Without these, the empirical support for the extended-DoF claim cannot be assessed and the result remains unverifiable from the given text."}],"tokens_in":1353,"tokens_out":441,"duration_ms":15582,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core claim is that inserting a coded aperture at the pupil lets NeRF train directly on defocused images while still producing sharp novel views across a larger depth range.\n\nThe new piece is the explicit camera model that replaces the usual pinhole or circular-aperture integral with one that includes the coded pupil function. Earlier NeRF work treated the aperture as fixed and simple; this version makes the aperture part of the differentiable forward model. That hardware-aware step is the actual addition.\n\nThe paper reports simulations and real experiments that supposedly outperform standard apertures. If those include side-by-side PSNR or SSIM numbers on the same scenes and a physical coded-aperture prototype, the result is practically useful for anyone capturing NeRF datasets with ordinary lenses.\n\nThe weak point is the optical justification. The abstract asserts that the code preserves high spatial frequencies even for out-of-focus points, yet gives no derivation of the pupil function, its Fourier transform, or the resulting depth-dependent OTF. Without that step, it is unclear whether the model remains invertible across the claimed depth range or only for particular codes and scene content. The stress-test note correctly flags this gap; if the full manuscript does not supply the explicit PSF calculation or show that the chosen code actually transmits the needed frequencies, the central claim rests on an unverified assumption.\n\nThe work targets people building NeRF pipelines for real-world capture where depth of field is a constraint. Readers who care about computational imaging plus neural rendering will find the direction relevant. The problem is concrete and the proposed fix is specific, so the paper merits a serious referee even if the optical modeling needs tightening.","headline":"The paper folds a coded aperture into the NeRF rendering model to handle wider depth ranges, but the optical derivation for frequency preservation under defocus is not shown in the abstract.","tokens_in":2295,"tokens_out":411,"would_cite":false,"duration_ms":17281,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Coded aperture integration into NeRF allows direct use of defocused images for high-fidelity novel views with extended depth of field.","keywords":["neural radiance fields","coded aperture","depth of field","novel view synthesis","camera model","defocus","photorealistic rendering","extended depth of field"],"falsifier":"Capture the same scene with both a conventional aperture and the coded aperture at identical focus settings, then compare the peak signal-to-noise ratio of novel views synthesized by each NeRF; a drop below the conventional case would falsify the claim.","tokens_in":2544,"feed_emoji":"📷","tokens_out":595,"duration_ms":19165,"temperature":0.7,"pith_summary":"NeRF training datasets inherit the depth-of-field versus light-quantity trade-off from conventional cameras. The paper introduces a coded aperture at the pupil that preserves spatial frequency content even when images are defocused. A new camera model is built inside the NeRF framework so that these coded images can be fed directly into the training process. The resulting EDoF-NeRF then renders novel views that remain sharp over a larger depth range than standard aperture models permit. Simulations and real-camera experiments show improved fidelity relative to conventional capture.","feed_headline":"Coded aperture extends NeRF depth of field","feed_subtitle":"New camera model accepts defocused coded images and renders sharp novel views across wider focus ranges than standard apertures allow.","key_machinery":"Coded aperture placed at the camera pupil, integrated into the NeRF forward imaging model to handle defocus without frequency loss.","core_discovery":"We develop a camera model that incorporates coded apertures into NeRF, allowing direct input of coded images and enabling the generation of novel views with an extended DoF while maintaining high fidelity.","pith_inferences":["The approach may generalize to other implicit scene representations that rely on differentiable camera models.","Optimized aperture codes could be scene-specific, learned jointly with the radiance field rather than fixed in advance.","Extending the method to dynamic scenes would require the coded pattern to preserve frequencies across both space and time."],"forward_implications":["NeRF training sets can be acquired with wider apertures, shortening exposure times while still covering larger scene depths.","Novel-view synthesis quality remains high even when input images contain intentional defocus from the coded pattern.","The same coded-aperture model can be swapped into existing NeRF pipelines without changing the underlying radiance-field representation.","Real-world experiments confirm the simulated gains, indicating the method transfers from rendered to physical cameras."],"fun_headline_variants":["Coded aperture extends NeRF DoF","Extended DoF NeRF via coded aperture camera","NeRF uses coded images for wider focus range","Camera model incorporates coded apertures into NeRF","Coded aperture input enables extended DoF neural radiance fields"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The chosen coded aperture pattern must continue to transmit the scene's spatial frequencies even when the sensor plane lies outside the focal range.","fun_headline_variants_meta":{"raw":{"variants":["Coded aperture extends NeRF DoF","Extended DoF NeRF via coded aperture camera","NeRF uses coded images for wider focus range","Camera model incorporates coded apertures into NeRF","Coded aperture input enables extended DoF neural radiance fields"]},"model":"grok-4.3","cost_usd":0.004301,"raw_usage":{"total_tokens":2110,"prompt_tokens":564,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":43012000,"prompt_tokens_details":{"text_tokens":564,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1479,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":564,"tokens_out":67,"duration_ms":13234,"temperature":1.0,"reasoning_tokens":1479,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T20:19:56.200244+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Capture the same scene with both a conventional aperture and the coded aperture at identical focus settings, then compare the peak signal-to-noise ratio of novel views synthesized by each NeRF; a drop below the conventional case would falsify the claim.","supporting_citations":[],"review_version":1}