{"id":"de19288e-3b19-4454-9add-4970f9cb61eb","arxiv_id":"2509.04402","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A self-supervised neural framework simultaneously reconstructs the object and the unknown probe in X-ray ptychography, improving low-dose and low-overlap reconstructions.","lead":"PtyINR uses two neural networks, one for the sample and one for the illuminating X-ray beam, to reconstruct both at the same time from measured diffraction patterns without the beam being characterized in advance. The authors report sharper reconstructions than conventional ptychography algorithms, especially at low X-ray dose or low scan overlap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported 'consistent' superiority lacks repeated-seed statistics despite the paper's own admission of non-negligible failure probability, and key per-scenario hyperparameters (λ, k, ω0, β) are unreported; the SOTA claim is not yet supported.","rationale":"I read the paper in good faith: PtyINR is a plausible self-supervised INR approach, the forward model is standard, and the real-data comparisons are extensive. The reader's weakest assumption (single-mode probe and exact positions) is a valid limitation and is explicitly acknowledged in the Conclusion, but it does not by itself falsify the tested claims because the data may be approximately single-mode and the paper already flags those caveats. The more load-bearing weakness is internal to the evidence: the authors claim consistent superiority while their own supplementary text discloses a non-negligible failure probability for exactly the joint object-probe recovery setting, yet no statistics on success rate, seed variance, or hyperparameter sensitivity are provided. Since the only quantitative support for SOTA is single-run PSNR and FRC numbers, and since key hyperparameters λ, k, ω0, and β are either unreported or tuned per scenario, the central empirical claim is not yet reproducible in a statistically meaningful sense. This does not require rejecting the paper; it does require the authors to supply repeated-seed statistics and exact hyperparameters, which is consistent with the reader's CONDITIONAL verdict. My concern points at the same conclusion but through a different, more directly testable route, hence 'partial' agreement with the reader's stated weakest assumption.","tokens_in":22045,"tokens_out":4452,"duration_ms":46508,"concrete_test":"Using the released code, rerun the simulation overlap experiments (40%, -540%, -800%) and the 0.003 s low-dose experiment with at least 10 random seeds per configuration, using the exact λ, k, β, and ω0 values that produced the reported figures. Report mean±std PSNR for object/probe amplitude and phase, plus the fraction of seeds that converge within 3 dB of the reported best PSNR. If any configuration has a failure rate above ~20% or the mean PSNR is within the baseline variance, the \"consistently outperform\" claim fails. Also rerun with λ=0 after step k and with k varied by ±50% to quantify sensitivity to the unreported regularization schedule.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that PtyINR consistently outperforms existing methods, especially at low overlap or low SNR. The paper's own Supplementary Note (p. 24, discussion around Fig. S15) admits that when object and probe are both unknown, \"the reconstruction may fail entirely with a non-negligible probability,\" and that without the probe regularization/normalization, an FZP probe \"diverges.\" Yet the Main text reports no success rates, no repeated-seed error bars, and no PSNR distributions. All comparisons are single-run point estimates. Moreover, key controls of the method are tuned per experiment and not reported: β is varied over 1e-4–1e-1, ω0 is chosen per scan step size (e.g., 30 vs. 90 in Fig. S17), and the regularization coefficient λ and cutoff k in Eq. (2) are never given numerical values. Consequently, the reported PSNR improvements could reflect best-of-N or hand-tuned runs, and the \"consistent\" robustness claim is not established. This is load-bearing because the headline contribution is an empirical superiority claim; if the failure rate at -800% overlap or 0.003 s exposure is appreciable, PtyINR is not a reliable SOTA in exactly the regimes emphasized.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PtyINR, a self-supervised implicit neural representation framework for joint recovery of the object and the unknown illumination probe in X-ray ptychography. The object is parameterized by a SIREN and the probe by a multi-resolution hash-encoded ReLU MLP; training minimizes a SmoothL1 diffraction-consistency loss with a probe-amplitude regularization term. The authors report simulations over scanning overlap ratios, noise conditions, and real datasets (LiCoO2 and Siemens star, FZP and MLL optics) and claim consistent state-of-the-art performance, especially in sparse and low-signal regimes.","tokens_in":22423,"tokens_out":5072,"duration_ms":49976,"significance":"If the empirical claims are substantiated, this is a useful advance: it removes probe pre-characterization and paired training data, is applicable to standard ptychographic forward models, and could benefit low-dose experiments. The paper ships released code, includes extensive ablation studies in the supplementary material, and evaluates on real data from two focusing systems. However, the headline 'consistently outperforms' claim is currently supported only by single-run point estimates, and key controls are unreported, so the significance is conditional.","major_comments":[{"comment":"The central claim is that PtyINR 'consistently outperform[s] existing techniques' and shows 'remarkable robustness' in low-overlap/low-dose regimes. The evidence consists of single-run PSNR values. The supplement itself states that when object and probe are jointly unknown, 'the reconstruction may fail entirely with a non-negligible probability' (p. 24). No success rates, repeated-seed error bars, medians, or failure statistics are reported for the final architecture or for the baselines. Because random initialization and Adam make the method stochastic, a point estimate cannot support a consistency/robustness claim. Please report seed/initialization statistics (e.g., at least 10 runs) and success probabilities, particularly at -800% overlap and 0.003 s exposure.","section":"Results, Figs. 2–5; Supplementary Note, Fig. S15 (p. 24)"},{"comment":"Several per-scenario controls are tuned but not quantified. Eq. (2) defines the regularization coefficient λ and cutoff k, but no numerical values are given anywhere in the paper. The Methods section only reports ranges (β between 1e-4 and 1e-1, learning rate 1e-5 to 1e-4), and Fig. S17 shows ω0 varied among 30, 90, and 300 without stating which value is used for each experiment in Figs. 2–5. Without a table of per-experiment hyperparameters, the reported gains may reflect per-scenario tuning rather than the method itself. Code release helps but does not replace reporting in the manuscript.","section":"Methods, Eq. (2); Supplementary Note, Fig. S17"},{"comment":"The FRC-based resolution claim requires two independent reconstructions F1 and F2, but the paper never states how these were obtained (e.g., splitting scan positions, different random seeds, or separate measurement sets). If the two inputs are not truly independent, the half-bit threshold overestimates resolution. The protocol must be specified or replaced with a controlled metric.","section":"Methods, Eq. (11); Results, Fig. 5e"},{"comment":"The paper acknowledges single-mode probe and exact position assumptions in the Conclusion, but the Abstract describes the method as 'generalizable' and applicable to 'a wide range of computational microscopy problems.' In practice, partial coherence or scan-position errors can be absorbed into the reconstructed object and probe fields, which would undermine the claim of recovering the true probe under low-signal conditions. Please either scope the claims to the tested assumptions or include a mismatched-model experiment (e.g., a two-mode probe or position error).","section":"Conclusion; Abstract"}],"minor_comments":[{"comment":"The negative overlap labels (-540%, -800%) are unintuitive; the text explains the actual overlap (95%, 50%, 30%) but the figure axes should carry a clarifying note.","section":"Results, Fig. 2"},{"comment":"The LiCoO2 amplitude images are excluded as 'non-informative.' If this is due to weak absorption, that is plausible, but the criterion should be stated and the images shown in the supplement so readers can assess the choice.","section":"Results, Fig. 4"},{"comment":"The panel label '(Appropriate loss should be 0)' is unclear: it appears to refer to the regularization loss, but the text is ambiguous. Please rephrase and specify which loss should be zero.","section":"Supplementary Note, Fig. S15"},{"comment":"A few grammatical issues: 'we evaluates,' 'these methods fails,' and 'it fail to preserve' should be corrected. Also, the claim 'state-of-the-art' should be qualified because recent methods such as the stochastic ADMM of Ref. [22] are cited but not benchmarked.","section":"Introduction and Results"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible application of INRs to ptychographic reconstruction and the code release is a genuine strength. The main risk is that the 'consistent robustness' claim is undercut by the absence of repeated-seed statistics and by the supplement's own admission of non-negligible failure probability in joint recovery. These concerns are addressable with additional experiments and reporting, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this paper is a credible INR-based blind ptychography method with the strongest claims in the low-overlap/noisy regimes, but the empirical evidence for 'consistently outperforms' isn't as solid as the prose makes it sound.\n\nWhat is new: prior INR work in Fourier ptychography, adaptive optics, and tomography either assumed a known probe or didn't treat ptychographic blind recovery. PtyINR parameterizes the object with a SIREN and the probe with a hash-encoding ReLU network, and uses a SmoothL1 physics loss plus a probe-amplitude regularization during early training. When the probe is unknown, that's a real step forward. The simulation suite is extensive — overlap ratios down to -800%, three noise models, real data from FZP and MLL setups, exposure times down to 0.003s. The ablations in the supplement (loss function, omega0, regularization) suggest the authors looked at their failure modes and engineered around them. Code is reportedly released. That's a solid methods package.\n\nThe soft spots are about evidence for the headline claim. All PSNR numbers are single-run point estimates; the supplement admits that with both object and probe unknown, 'the reconstruction may fail entirely with a non-negligible probability.' That admission is honest, but it cuts directly against the word 'consistently.' No success rates, no error bars, no seed variations. The hyperparameters are tuned per scenario: omega0 varies with step size, beta ranges over three orders of magnitude, and the regularization strength lambda and cutoff k in Eq. 2 are never given values. The LiCoO2 amplitude channel is excluded post-hoc as 'non-informative' — a red flag unless explained. And the single-mode probe/known-position assumptions limit the real-data conclusions; the authors acknowledge that in the conclusion. The FRC analysis covers only object phase. These are not fatal, but they are exactly the places where a 'state-of-the-art' claim needs support.\n\nWho it's for: researchers in X-ray ptychography/computational microscopy who want to see whether blind INR-based reconstruction is viable, especially for low-dose experiments. A referee should engage with it — the architecture design and the experimental breadth deserve serious evaluation. The authors should be pushed to provide repeated-seed statistics, actual hyperparameter values, and a clear account of the amplitude-channel exclusion before the claim should be accepted.\n\nRecommendation: send it to peer review rather than desk-rejecting.","headline":"PtyINR: a credible blind probe+object INR for ptychography, but the 'consistent SOTA' claim needs repeated-seed stats and hyperparameter values.","tokens_in":22882,"tokens_out":3527,"would_cite":true,"duration_ms":28106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PtyINR claims that X-ray ptychography can jointly recover both the sample and the unknown illuminating probe from raw diffraction patterns by parameterizing each as a continuous neural field, and that this beats established solvers especial","keywords":["X-ray ptychography","implicit neural representation","blind reconstruction","unknown probe recovery","self-supervised learning","low-dose imaging","phase retrieval","coherent diffractive imaging"],"falsifier":"Simulate ptychographic data from a known two-mode probe, or impose known scan-position jitter, then run PtyINR and check whether the object error grows with the mode-mixture or position-error magnitude. A clean control is to give a conventional solver the true probe modes and exact positions under the same noise and see whether it recovers the object with higher fidelity than PtyINR; if so, the single-mode/exact-position assumption is the crucial limit of the central claim.","tokens_in":21941,"feed_emoji":"🔬","tokens_out":5253,"duration_ms":50589,"temperature":0.7,"pith_summary":"X-ray ptychography records only the intensity of diffracted X-rays while the illuminating beam (the probe) is usually unknown, making image recovery a blind phase-retrieval problem. This paper proposes PtyINR, a self-supervised method that represents both the sample and the probe as continuous functions learned by neural networks, and trains them end-to-end by simulating diffraction patterns and comparing them with measurements. The central claim is that one unified network can recover both fields without probe pre-characterization or labeled training data, and that it does so more accurately than established iterative solvers (ePIE, DM, WASP, APG) and learning-based baselines, especially when overlap is low or the signal is noisy. If correct, this removes a major practical bottleneck: probes no longer need to be measured separately, and low-dose or radiation-sensitive samples become more accessible to high-resolution ptychography.","feed_headline":"Two neural fields recover X-ray probe and object in one pass","feed_subtitle":"A self-supervised network reconstructs both sample and illumination from noisy, sparse diffraction data without probe pre-characterization.","key_machinery":"The load-bearing object is the pair of implicit neural representations: a SIREN (sine-activated MLP) for the object and a multi-resolution hash-encoding network (Instant-NGP style) for the probe. The asymmetry matters because the probe contributes to every diffraction pattern, so its representation must train stably, while the object contributes only locally and needs expressive high-frequency modeling. The other key piece is the hybrid loss: a SmoothL1 interpolation between l1 and l2 in intensity space, plus a probe mean-amplitude regularization applied only in early training and a normalization constraint. Together these carry the blind joint recovery; the paper's ablations attribute the s","core_discovery":"PtyINR's discovery is that ptychographic blind reconstruction can be formulated as pure self-supervised neural-field fitting. The object is parameterized as two coordinate-to-value MLPs with periodic activations, giving amplitude and phase; the probe is parameterized by a multi-resolution hash-encoded ReLU network. The forward model computes the predicted far-field diffraction intensity by Fourier transform of the exit wave, and the networks are trained by minimizing a SmoothL1 loss between predicted and measured intensities. To make joint recovery stable, the probe amplitude is normalized and a mean-amplitude regularization is applied in early training. On simulated data with 40%, -540%, an","pith_inferences":["A testable generalization is that the stable-probe/expressive-object split of representations may be a general recipe for blind inverse problems: parameterize the globally shared unknown with a low-variance hash-encoded network and the target with a high-frequency sine network.","The early-training probe regularization suggests a curriculum strategy where the probe is constrained until the object forms a rough structure and then released; the same trick could be exported to other self-supervised joint estimation tasks.","Because the probe neural field is a continuous function of coordinates, a probe network trained on one beamline configuration might be transferred to new samples measured with the same optics, making reconstructions faster and more stable than starting from random probe initialization.","The single-mode assumption is the natural place to test limits: if the method is extended to multi-mode probes by adding a mode dimension to the probe network, one can check whether the regularization still prevents degenerate solutions."],"forward_implications":["Separate probe characterization becomes unnecessary for high-quality ptychography; the probe is recovered from the same data as the object.","Sparse overlap and very short exposures become usable, extending ptychography to dose-sensitive and dynamic samples.","The two-network physics-loss recipe can be carried to other probe-dependent imaging geometries such as near-field and Bragg ptychography by replacing the forward model.","Because the object is a continuous function rather than a pixel grid, irregular scan geometries can be handled without grid-interpolation artifacts.","With abundant clean data the method is comparable to established solvers; its reported advantage concentrates in degraded, low-overlap, and low-dose regimes."],"supporting_citations":[{"why":"Supplies the SIREN sine-activated MLP architecture and initialization used for the object network.","marker":"[13]"},{"why":"Supplies the multi-resolution hash encoding and ReLU-MLP used for the probe network.","marker":"[49]"},{"why":"Fast R-CNN SmoothL1 loss, the basis of the hybrid loss used for intensity matching.","marker":"[48]"},{"why":"ePIE, the main iterative baseline that PtyINR is compared against.","marker":"[14]"},{"why":"Difference Map baseline and a foundational reference for probe retrieval in ptychography.","marker":"[16]"},{"why":"APG, a proximal baseline designed for noisy data, used in the noise and exposure-time experiments.","marker":"[19]"},{"why":"PtychoNN baseline and the source of the simulation object used in the synthetic benchmarks.","marker":"[27]"},{"why":"PINN, the self-supervised known-probe neural baseline contrasted with PtyINR's unknown-probe setting.","marker":"[36]"},{"why":"AD/Adorym, the auto-differentiation matrix-based baseline that serves as an ablation for the neural representation.","marker":"[34]"},{"why":"Fly-scan ptychography, providing the forward-model integration framework and the experimental acquisition mode.","marker":"[44]"}],"fun_headline_variants":["Blind ptychography solved by self-supervised neural fields","One network pair recovers X-ray probe and sample together","Neural fields reconstruct X-ray images without known probe","Self-supervised AI tackles unknown-probe ptychography","PtyINR: joint probe and object recovery from diffractions"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The reconstruction assumes a single coherent probe mode and exactly known scan positions; if the real probe is partially coherent (multi-mode) or the scan positions drift, the recovered object and probe may absorb those mismatches and the claimed fidelity is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Blind ptychography solved by self-supervised neural fields","One network pair recovers X-ray probe and sample together","Neural fields reconstruct X-ray images without known probe","Self-supervised AI tackles unknown-probe ptychography","PtyINR: joint probe and object recovery from diffractions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000504,"raw_usage":{"total_tokens":2281,"prompt_tokens":708,"completion_tokens":1573,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":1504}},"tokens_in":452,"tokens_out":1573,"duration_ms":11010,"temperature":1.0,"reasoning_tokens":1504,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:12:52.787969+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate ptychographic data from a known two-mode probe, or impose known scan-position jitter, then run PtyINR and check whether the object error grows with the mode-mixture or position-error magnitude. A clean control is to give a conventional solver the true probe modes and exact positions under the same noise and see whether it recovers the object with higher fidelity than PtyINR; if so, the single-mode/exact-position assumption is the crucial limit of the central claim.","supporting_citations":[{"cited_title":", author Martel, J","cited_arxiv_id":null,"evidence_quote":"Supplies the SIREN sine-activated MLP architecture and initialization used for the object network."},{"cited_title":", author Evans, A","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-resolution hash encoding and ReLU-MLP used for the probe network."},{"cited_title":"title Fast R - CNN","cited_arxiv_id":null,"evidence_quote":"Fast R-CNN SmoothL1 loss, the basis of the hybrid loss used for intensity matching."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ePIE, the main iterative baseline that PtyINR is compared against."},{"cited_title":", author Dierolf, M","cited_arxiv_id":null,"evidence_quote":"Difference Map baseline and a foundational reference for probe retrieval in ptychography."},{"cited_title":"title Ptychographic phase retrieval by proximal algorithms","cited_arxiv_id":null,"evidence_quote":"APG, a proximal baseline designed for noisy data, used in the noise and exposure-time experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PtychoNN baseline and the source of the simulation object used in the synthetic benchmarks."},{"cited_title":", author Mishra, A","cited_arxiv_id":null,"evidence_quote":"PINN, the self-supervised known-probe neural baseline contrasted with PtyINR's unknown-probe setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"AD/Adorym, the auto-differentiation matrix-based baseline that serves as an ablation for the neural representation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Fly-scan ptychography, providing the forward-model integration framework and the experimental acquisition mode."}],"review_version":1}