{"id":"d23dd0e8-c6ba-4d7e-98d4-253aa2a1472f","arxiv_id":"2412.02052","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Depth-prior-guided foveation of SPAD LiDAR histograms cuts memory and improves ambient-light robustness.","lead":"A new imaging pipeline uses an external depth guess, such as a monocular depth map or optical flow, to tell a single-photon LiDAR sensor where in time to look, so it stores fewer histogram bins and rejects more ambient light. The approach is tested in simulation and on real single-photon datasets, reporting up to a 1548-fold memory reduction while preserving depth accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The efficiency claim is conditioned on the foveation window containing the true histogram peak, but the paper never quantifies failure as a function of prior error, so the 'maintained depth accuracy' claim is not established.","rationale":"The reader's weakest-assumption diagnosis is correct: the foveation gains all hinge on the depth prior placing the true histogram peak inside the M-bin window. I do not see a reason to change the reader's CONDITIONAL verdict. The paper is honest about this dependence and provides qualitative discussion, but it does not quantify the failure rate or show that the efficiency gains survive realistic prior errors. Since the abstract and algorithm claims are stated as categorical improvements ('reduces raw data... while also improving resilience'), the missing prior-error robustness analysis is a genuine condition on acceptance, not a manufactured objection. The secondary theoretical errors in Eqs. (4) and (7) are independent correctness issues that the reader already flagged; they weaken the theory section but are not the primary reason for the conditional verdict. My proposed test is a direct empirical measurement of the conditioning event, so it would settle whether the concern actually lands. I agree with the reader's assessment, and no verdict change is needed.","tokens_in":19695,"tokens_out":3697,"duration_ms":39717,"concrete_test":"Run a controlled prior-error sweep on the NYUv2 simulation: take the locally calibrated ZoeDepth prior, apply known per-pixel depth offsets or scale perturbations (e.g., adding delta in {0, 0.1, 0.5, 1.0 m} or multiplying by {0.8, 0.9, 1.1, 1.2}), then execute Algorithm 1 with M=1/16 and N'=16. Report RMSE and the fraction of pixels whose true SPAD peak falls outside the foveation window, and compare this measured failure rate with the prediction of Eqs. (9)-(10). If small offsets sharply degrade RMSE or produce a large out-of-window fraction, the 'maintained accuracy' claim is conditional on prior quality; if accuracy degrades gracefully, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims (e.g., the 1548x memory reduction, the SBR gains, and the Table II/III accuracy numbers) are derived under the explicit assumption in Sec. III-C that 'the desired histogram peak is captured by these bins.' Eqs. (2)-(8) and the simulation protocol all condition on this event, and the paper admits this dependence in Sec. III-A and Sec. VIII. The proposed 'worst-case' analysis in Sec. VIII, Eqs. (9)-(10), does not close the gap: it is a stochastic expression for total depth-detection failure, not a validation against actual prior-error statistics, and it is not tested against any experiment. The monocular experiments scale and calibrate ZoeDepth locally using SPAD ground truth, yet no prior-error statistics or per-pixel window-inclusion rates are reported; the optical-flow experiments already show a failure of the conditioning event in the first CARLA scene (RMSE 101.9m at M=1/10N). Thus the paper's headline efficiency results conflate conditional gains, valid only when the prior is accurate, with the unconditional claim that foveation 'maintains depth accuracy.' This is not a disagreement about the value of foveation; it is a missing error-propagation analysis that is load-bearing for the central claim. The additional algebra errors in Eq. (4) (the cycle-count condition should scale as sqrt(N/M), not N^2/M^2) and Eq. (7) (the background denominator is omitted) reinforce that the theoretical support for the claim is weaker than presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces FoveaSPAD, a family of adaptive capture policies for SPAD-based LiDAR in which an external depth prior (monocular depth, optical flow, or low-resolution sampling) is used to restrict or reallocate histogram bins around an expected echo time. The authors distinguish memory foveation (fewer bins at full width) and depth foveation (fixed bin budget concentrated in a window), and they claim theoretical gains in SNR/SBR, memory, and depth resolution. The results are demonstrated on simulated datasets (NYUv2, CARLA) and hardware emulation of real SPAD data (Lindell et al., Gutierrez-Barragan et al.), including a reported 1548x memory savings in a spatio-temporal variant. The paper explicitly acknowledges that the strategies are dependent on the accuracy of the depth prior, and it provides a worst-case stochastic analysis in Sec. VIII.","tokens_in":20041,"tokens_out":5192,"duration_ms":46269,"significance":"Foveated capture is an important and timely idea for SPAD LiDAR because raw histogram bandwidth is a known bottleneck, and the paper makes a credible case that a prior can shift the sampling budget. The simulations and emulations are performed on standard public datasets, and the hardware-emulation results against real SPAD data add value. The reported memory savings are substantial. However, the stated theoretical support for the efficiency gains contains concrete errors in Eqs. (4) and (7), and the empirical claims are conditional on the prior correctly localizing the histogram peak, a condition that is acknowledged but never quantified. If these issues are corrected and an explicit error-propagation analysis is added, the paper could be a useful contribution to adaptive single-photon imaging.","major_comments":[{"comment":"The central efficiency claim is conditional on the foveation window containing the true histogram peak. Eqs. (2)-(8) and the simulation protocol all assume this event, and Sec. III-A states that the strategies are fundamentally dependent on prior accuracy. The worst-case analysis in Sec. VIII, Eqs. (9)-(10), is a stochastic expression for total depth-detection failure under multipath and noise, but it is not tied to measured prior-error statistics and is not validated experimentally. The optical-flow results in Sec. VI already exhibit a failure of the conditioning event: in the first CARLA scene at M=1/10N the reported RMSE is 101.9 m. Without a quantitative characterization of window-inclusion probability as a function of prior error, the claim that foveation 'maintains depth accuracy' is not established for realistic biased priors. This is a load-bearing gap because all reported memory and SNR gains depend on the peak being captured.","section":"Sec. III-C and Sec. VIII"},{"comment":"Eq. (4) states that depth foveation requires C_new/C >= N^2/M^2 to match conventional SNR. From Eq. (3), SNR is proportional to C sqrt(M T / N^2), which equals C sqrt(M/N) sqrt(T/N), while the conventional SNR in Eq. (1) is C sqrt(T/N). Equating these gives C_new/C = sqrt(N/M), not (N/M)^2. The current expression overstates the required exposure increase by a factor of (N/M)^{3/2} for typical M << N. This error directly affects the theoretical claim in Sec. III-C that depth foveation can be compensated by more laser cycles.","section":"Eq. (4)"},{"comment":"Eq. (7) omits the denominator in the SBR expression. Starting from Eq. (6) with j=i, the numerator becomes (1 - e^{-(Phi_sig+Phi_bkg)}), but the denominator is not 1: it contains p^i_bkg = (1 - e^{-Phi_bkg}) e^{-Sum_{1}^{i-1} Phi_bkg}. The perfect-foveation limit therefore still depends on the ambient level through the probability that a background photon is detected at the correct bin. The text's assertion that foveation 'removes the dependence on prior photon arrival' is only valid in the limit Phi_bkg approaching 0, which is the opposite of the strong-ambient-light regime this subsection addresses. This error weakens the claimed SBR advantage of memory foveation.","section":"Eq. (7)"}],"minor_comments":[{"comment":"The mathematical expression in Eq. (2) is garbled and should be rewritten for readability; the intended formula appears to be SNR proportional to sqrt(T/N).","section":"Eq. (2)"},{"comment":"The notation is overloaded: in the paragraph after Eq. (4), 'the foveated bins N are given to us' uses N to mean the earlier M, which is confusing.","section":"Sec. III-C"},{"comment":"The Table II header contains the typo 'ERRROR'; please correct to 'ERROR'.","section":"Table II"},{"comment":"The author affiliation line lists 'Gainsville, FL'; the correct spelling of the city is 'Gainesville'.","section":"Affiliation"},{"comment":"Eq. (5) contains an extra closing parenthesis after (1 - e^{-(Phi_sig+Phi_bkg)}), resulting in '(1 - e^{-(Phi_sig+Phi_bkg)}))'.","section":"Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a computational imaging journal. The main concern is not the value of the foveation idea but the gap between the conditional analysis and the unconditional claims; with the requested revisions the contribution would be publishable. No citation concerns beyond noting that the authors co-author several cited prior works, which is standard in this area."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Sarah — quick take on FoveaSPAD (2412.02052). The core idea is genuinely new: use an external depth prior to place the SPAD histogram bins during capture, rather than capturing full histograms or doing post-capture upsampling. The memory-foveation vs depth-foveation distinction is clear, and the spatio-temporal quantization variant (1/16 bins x ~1% pixels, giving a 1548x memory reduction) is a concrete, sensible mechanism. The hardware emulation on the Lindell and Gutierrez-Barragan datasets gives the paper real evidentiary weight; it's not pure simulation. The authors also state their scope honestly — no commercial foveation-capable SPAD array exists, and they propose a speculative pixel design.\n\nThe soft spots are real but mostly fixable. Eq. (4) has an algebra error: matching conventional SNR for depth foveation requires C_new/C ~ sqrt(N/M), not N^2/M^2. Eq. (7) drops the background denominator in the perfect-foveation SBR; the full expression still has a (1 - exp(-Phi_bkg)) in the denominator. These are mistakes a careful referee should catch, and they undermine the theory section as written, but they don't invalidate the qualitative claim.\n\nThe bigger conceptual gap is the one the authors themselves flag in Sec. III-A and VIII: all the efficiency numbers are conditional on the M foveated bins containing the true peak, and the paper never quantifies the failure rate as a function of prior error. The optical-flow experiment shows the problem concretely — at M=1/10N the first CARLA scene gives RMSE 101.9 m. The 'worst-case' stochastic analysis (Eqs. 9-10) is about degenerate mirror-like scenes, not about realistic monocular or flow prior errors. Without window-inclusion statistics, the claim that foveation 'maintains depth accuracy' is not yet established unconditionally.\n\nAlso minor: no error bars in the tables, no code release, and the local polynomial scaling of ZoeDepth is fit using full-resolution SPAD histograms on the same sensor, which is close to calibrating the prior on the test instrument. That's not fatal — the scaling uses a small set of pixels — but it should be disclosed more prominently.\n\nBottom line: this is a solid, worth-reading paper that should go to peer review. The idea is novel enough to merit referee time, and the emulations give it empirical grounding. The authors need to fix the algebra, add prior-error statistics, and either release code or provide more detailed experimental protocols. I'd accept it for review with major revisions.","headline":"A genuinely useful idea about adaptive SPAD histogram capture that deserves review, but the theory has correctable algebra errors and the headline efficiency numbers are conditional on an unquantified prior-accuracy assumption.","tokens_in":20585,"tokens_out":3939,"would_cite":false,"duration_ms":34537,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FoveaSPAD claims that guiding SPAD histogram capture with depth priors cuts raw data by 1548-fold while keeping depth accuracy and improving ambient-light resilience.","keywords":["single-photon avalanche diode","time-of-flight LiDAR","foveation","depth prior","photon timing histogram","memory efficiency","ambient light rejection","computational imaging"],"falsifier":"Use a static scene with known ground-truth depth, introduce a controlled bias into the depth prior that moves the window more than half its width away from the true peak, and compare foveated depth error to full-histogram depth error. If the foveated capture still recovers the correct peak, or if the error grows only as fast as bin width rather than jumping to the noise floor, the paper's dependence on prior accuracy would be contradicted; the paper's own worst-case analysis predicts the jump.","tokens_in":19486,"feed_emoji":"📡","tokens_out":7091,"duration_ms":64250,"temperature":0.7,"pith_summary":"FoveaSPAD argues that a SPAD time-of-flight camera does not need to store the full photon-timing histogram. If an external depth prior, from a monocular network, optical flow, or a coarse scan, places the laser echo inside a small window of M histogram bins, the sensor can capture only those bins. The paper derives why this works: memory foveation keeps the same per-bin SNR while cutting stored data by M/N, and it raises the signal-to-background ratio because ambient photons arriving before the signal bin no longer contribute to detector dead time. Depth foveation instead packs a fixed memory budget into the window to improve depth resolution, at an SNR cost that more laser cycles can repay. If the claim holds, SPAD LiDAR can approach the depth quality of full histograms with a fraction of the bandwidth, which matters for autonomous vehicles and power-constrained sensors.","feed_headline":"Depth priors cut SPAD LiDAR memory 1548-fold","feed_subtitle":"Foveating photon histograms to expected depths keeps depth accuracy while shrinking data and rejecting ambient light.","key_machinery":"The central object is the foveation window: a per-pixel subset of M histogram bins (M much smaller than N) placed around an estimated depth from a prior. The argument runs through two identities. First, SNR scales with bin width, so keeping the original bin width inside the M-bin window preserves SNR while storing only M/N of the histogram. Second, SBR depends on the probability that an ambient photon in an earlier bin resets the detector before the laser echo; Eq. (6) shows foveation removes those early-bin terms from the denominator, and perfect foveation reduces SBR to the direct signal-vs-background ratio. Depth foveation reuses the memory savings to place more, narrower bins inside the window, which is how the paper converts bandwidth savings into resolution.","core_discovery":"The paper's central claim is that foveated capture, gating each SPAD pixel to a window of M bins centered on a depth prior, converts the SPAD histogram bottleneck into a tunable trade. In memory foveation the bin width stays T/N and only the M bins around the predicted peak are stored; SNR is unchanged and memory falls by M/N. In depth foveation the window is re-divided into the same number of bins one would otherwise spread over the full range, so depth resolution improves while SNR drops by sqrt(M/N), recoverable by increasing the number of laser cycles. Under ambient light, memory foveation raises SBR because photons arriving before the foveation window no longer reset the detector; with a perfect prior the denominator's prior-bin dependence vanishes, leaving SBR proportional to the direct signal-vs-background odds. The paper supports this with simulations on an indoor RGB-D benchmark using a monocular prior, spatio-temporal quantized sampling that yields a 1548-fold memory reduction, optical-flow-driven foveation on driving scenes, and hardware emulation on real SPAD datasets where even simple peak detection improves after foveation.","pith_inferences":["A natural extension the paper does not explore is making the window size adaptive to prior confidence: pixels with high prior uncertainty could keep larger M, trading memory for robustness exactly where the prior is weakest.","The same M/N storage saving implies that once per-pixel gating hardware exists, a SPAD array could raise spatial resolution or frame rate by roughly N/M without increasing off-chip bandwidth, for scenes with accurate priors.","Because the related-work section notes compressive histogramming and sketching are complementary, a plausible next test is combining foveated capture with a compressive projection to compound the bandwidth reduction; the paper does not run that experiment.","The paper's worst-case probability analysis suggests a quantitative failure boundary: scenes with multipath effects degrade catastrophically only when the single-bounce probability and multipath probability satisfy specific relations, so a calibration experiment sweeping prior bias against M could map acceptable operating conditions."],"forward_implications":["On an indoor RGB-D benchmark with a monocular depth prior, memory foveation at 1/16 of the histogram bins keeps depth errors close to full-resolution SPAD simulation, and depth foveation with the same memory budget improves resolution over uniformly spread limited bins.","Spatio-temporal foveation that samples a few pixels per quantized depth bucket and foveates each in time achieves a 1548-fold memory reduction while still recovering scene depth.","Memory foveation extends the operable ambient-light range: hardware emulation shows the foveated photon cube has fewer background detections and a simple maxima estimator recovers structure that full-histogram maxima misses.","Optical-flow-driven foveation transfers the depth prior between frames for moving scenes, with a noise-floor comparison that resets pixels whose foveated window has drifted off the signal.","Superpixel-based foveation on real SPAD scans without a co-located camera reduces per-pixel memory by about 64 times for over 99 percent of pixels, using one full histogram per segment to anchor the window."],"supporting_citations":[{"why":"Supplies the SPAD simulation framework and the single-pixel scanned datasets used for hardware emulation.","marker":"[5]"},{"why":"Supplies real SPAD data with a co-aligned camera and the denoising network used as a baseline in hardware emulation.","marker":"[4]"},{"why":"Defines the signal-to-background ratio and the binomial SPAD capture model that the foveation SBR analysis builds on.","marker":"[46]"},{"why":"Provides the Poisson/binomial photon-detection probability model used to write the conventional and foveated SBR expressions.","marker":"[47]"},{"why":"Provides the monocular depth estimator used to generate the depth prior for the indoor RGB-D simulations.","marker":"[51]"},{"why":"Cited as a proof-of-concept reconfigurable SPAD array with in-pixel histogramming that motivates per-pixel foveation hardware.","marker":"[6]"},{"why":"Provides the driving simulator whose optical flow drives dynamic-scene foveation experiments.","marker":"[54]"}],"fun_headline_variants":["FoveaSPAD cuts SPAD LiDAR memory 1548-fold","Depth priors foveate SPAD captures, cutting memory 1548-fold","Foveated photon windows trim SPAD LiDAR memory 1548-fold","Adaptive foveation gating cuts SPAD LiDAR memory 1548-fold","Depth-prior foveation yields 1548-fold SPAD memory cut"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the depth prior being accurate enough that the true laser echo lands inside the M-bin foveation window: if the prior is biased by more than half the window width, the saved memory and the SNR/SBR gains collapse because the sensor never records the signal.","fun_headline_variants_meta":{"raw":{"variants":["FoveaSPAD cuts SPAD LiDAR memory 1548-fold","Depth priors foveate SPAD captures, cutting memory 1548-fold","Foveated photon windows trim SPAD LiDAR memory 1548-fold","Adaptive foveation gating cuts SPAD LiDAR memory 1548-fold","Depth-prior foveation yields 1548-fold SPAD memory cut"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":4065,"prompt_tokens":1048,"completion_tokens":3017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2913}},"tokens_in":664,"tokens_out":3017,"duration_ms":21819,"temperature":1.0,"reasoning_tokens":2913,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:53:28.994330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use a static scene with known ground-truth depth, introduce a controlled bias into the depth prior that moves the window more than half its width away from the true peak, and compare foveated depth error to full-histogram depth error. If the foveated capture still recovers the correct peak, or if the error grows only as fast as bin width rather than jumping to the noise floor, the paper's dependence on prior accuracy would be contradicted; the paper's own worst-case analysis predicts the jump.","supporting_citations":[{"cited_title":"Compressive single-photon 3d cameras,","cited_arxiv_id":null,"evidence_quote":"Supplies the SPAD simulation framework and the single-pixel scanned datasets used for hardware emulation."},{"cited_title":"Single-photon 3d imaging with deep sensor fusion,","cited_arxiv_id":null,"evidence_quote":"Supplies real SPAD data with a co-aligned camera and the denoising network used as a baseline in hardware emulation."},{"cited_title":"Asynchronous single-photon 3d imaging,","cited_arxiv_id":null,"evidence_quote":"Defines the signal-to-background ratio and the binomial SPAD capture model that the foveation SBR analysis builds on."},{"cited_title":"Photon-flooded single- photon 3d cameras,","cited_arxiv_id":null,"evidence_quote":"Provides the Poisson/binomial photon-detection probability model used to write the conventional and foveated SBR expressions."},{"cited_title":"A reconfigurable 3-d-stacked spad imager with in-pixel histogramming for flash lidar or high-speed time-of-flight imaging,","cited_arxiv_id":null,"evidence_quote":"Cited as a proof-of-concept reconfigurable SPAD array with in-pixel histogramming that motivates per-pixel foveation hardware."}],"review_version":1}