{"id":"0e2f33d6-4aec-4cfd-8daa-3df69d6479c5","arxiv_id":"2411.16896","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MFliNet, a differential-transformer network that takes pixelwise IRF as input, estimates fluorescence lifetime parameters robustly under IRF shifts from surface height variations.","lead":"A deep learning model called MFliNet estimates fluorescence lifetimes from time-resolved images while also taking the instrument response function (IRF) as input, which helps it cope with varying sample surface heights. It matches conventional fitting on tissue phantoms and tumor-bearing mice, but runs much faster.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training IRFs (flat diffuser, imaging table, Sec. 2.2) are not shown to contain the 40–160 ps height/anatomy offsets of Fig. 2, so the claimed offset robustness may be untested generalization rather than learned behavior.","rationale":"The reader's weakest_assumption and my stress-test point to the same condition: the training IRF set must be representative of test-time IRF offsets. I find this to be the single most load-bearing assumption because the architecture's novelty and the abstract's headline claim ('robustness and suitability for complex macroscopic FLI applications') rest on it. The shared-forward-model issue raised in the reader's rationale is real but secondary: NLSF is used as a comparator, and both NLSF and MFliNet assume the bi-exponential model of Eq. 1; this limits absolute-accuracy claims, but it would not change the relative comparison at different heights. The offset generalization issue, however, directly targets whether the input-IRF design works as claimed. I do not see an internally inconsistent derivation or a machine-checkable flaw; the architecture is plausible and the runtime improvement is concrete. The missing code/data and the lack of quantitative phantom numbers already justify the CONDITIONAL verdict. My proposed retraining experiment would settle whether the apparent offset invariance is learned or accidental; until then the verdict should remain CONDITIONAL, not stronger. No change to the reader's verdict is needed.","tokens_in":9875,"tokens_out":5281,"duration_ms":53975,"concrete_test":"Augment the training set with the measured height-dependent IRFs (or shifted copies of the existing IRF set) at offsets 0, 40, 80, 120, and 160 ps, generating corresponding shifted TPSFs via Eq. 1; retrain MFliNet from scratch with identical hyperparameters; and rerun the Fig. 3 phantom comparison and the Fig. 4 in-vivo analysis. If the augmented model matches NLSF but the original model does not, the published robustness was an untested generalization. If both models match NLSF equally, the original training IRFs evidently already contained sufficient offset variability, and the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that feeding the pixelwise IRF alongside the TPSF makes MFliNet robust to IRF offsets caused by sample topography (Figs. 2–4). That claim depends on the training distribution covering those offsets. Section 2.2 states only that experimental IRFs were captured from a white diffuser paper placed on the imaging table and that each simulated TPSF was generated by convolving a randomly selected IRF from this dataset. It does not state that height-induced time-of-flight shifts (the 40–160 ps offsets reported in Sec. 3 for the ladder phantom, or the organ/tumor IRFs of Fig. 2(b,c)) were included or simulated. A flat diffuser on the table does not, by itself, produce such offsets. If the training IRFs span only a narrow offset range, then at test time MFliNet must generalize common temporal shifts of both inputs—a nontrivial property for a transformer with absolute positional structure. The phantom and in-vivo agreement with NLSF is the only evidence for this generalization, and it cannot distinguish 'learned handling of offset' from 'accidental invariance on these particular cases.' Since the IRF input is the paper's main novelty, this missing distributional overlap is the most load-bearing weakness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MFliNet, a deep learning architecture that estimates fluorescence lifetime parameters (short lifetime, long lifetime, fractional amplitude) from a temporal point spread function (TPSF) and a pixelwise instrument response function (IRF) input. The model uses a differential transformer encoder–decoder with three output branches, trained on synthetic data generated from the bi-exponential convolution model of Eq. (1), where experimental IRFs were captured from a flat white diffuser paper placed on the imaging table. The authors validate MFliNet on a step-ladder phantom that introduces height-dependent IRF offsets of 40–160 ps and on HER2+ tumor xenografts in mice, comparing against NLSF (AlliGator), FLI-Net, and a standard transformer of the same architecture. The paper claims that MFliNet matches NLSF accuracy while being much faster, and that it avoids the height-dependent bias seen in FLI-Net and offset-free NLSF.","tokens_in":10155,"tokens_out":3011,"duration_ms":28838,"significance":"If the results hold, MFliNet would be a timely contribution: it is, to my knowledge, the first macroscopic FLI network that explicitly ingests the pixelwise IRF, and the differential-transformer design is a plausible way to focus on informative temporal features. The inclusion of a stepped-phantom experiment and an in-vivo xenograft comparison is commendable, and the reported speed advantage over NLSF is substantial. However, the central accuracy claim is not yet established: the reference method (NLSF) shares the same bi-exponential convolution model used to generate the training data, so agreement between MFliNet and NLSF demonstrates consistency rather than absolute accuracy; and phantom MFliNet values are not reported numerically. The paper would be significantly stronger with an independent lifetime standard, quantitative phantom values with statistical comparisons, and a demonstration that the training IRF distribution actually contains the height- and anatomy-dependent offsets that the model is claimed to handle.","major_comments":[{"comment":"The training IRFs are described as captured from a flat white diffuser paper placed on the imaging table and convolved with simulated decays; the paper does not state that height-induced time-of-flight shifts (the 40–160 ps offsets reported in Sec. 3) were included in the training set. If these offsets are absent from the training distribution, then the phantom and in-vivo results in Sec. 3 demonstrate an untested generalization rather than a learned offset-invariance. Because the pixelwise IRF input is the paper's main novelty, this gap is load-bearing. Please clarify whether the training data included IRFs spanning the observed offset range, or, if not, provide a controlled experiment (e.g., test-time IRFs with known shifts) showing that the model generalizes to out-of-distribution offsets.","section":"Sec. 2.2 (training data generation)"},{"comment":"The text states that MFliNet results were 'within the same range as NLSF' but reports no numerical mean±SD values for MFliNet at any of the five heights. Without these numbers, the claim that MFliNet avoids the height-dependent bias visible in FLI-Net and offset-free NLSF cannot be verified quantitatively from the manuscript. Please report the per-height means and standard deviations for all methods and include a statistical test (e.g., ANOVA or paired comparisons) for height-dependence.","section":"Sec. 3 (phantom results)"},{"comment":"The evaluation reference (NLSF with AlliGator) and the training data are both based on the same bi-exponential convolution model of Eq. (1). Consequently, agreement between MFliNet and NLSF establishes that the network reproduces the NLSF fit under the same model assumptions, but not that the estimated lifetimes are accurate in an absolute sense. No independent lifetime standard (e.g., a fluorophore with a known lifetime or a physical calibration target) is used. Please add such a validation, or explicitly discuss the limitation that the reported accuracy is relative to the NLSF model.","section":"Sec. 3 and Eq. (1)"}],"minor_comments":[{"comment":"The sentence 'the data was generated using the MNIST dataset' is vague; please clarify how MNIST images determine the spatial distribution of the three lifetime parameters (τ1, τ2, AR) across the 28×28 pixel frame.","section":"Sec. 2.2"},{"comment":"The differential attention formula, as written, computes a difference of two softmax-weighted value sums: softmax(Q1K1^T/√dk)V1 − λ·softmax(Q2K2^T/√dk)V2. Please verify that this matches the definition in Ref. [33], or state explicitly any adaptation made for the FLI setting.","section":"Eq. (2)"},{"comment":"The caption says 'Violin plots of NLSF analysis, FLI-Net, transformer model and MFliNet for all outputs,' but the figure panels appear to show lifetime maps and violin plots for the amplitude-weighted mean lifetime. Please make the caption more specific about which output (τM, τ1, τ2, or AR) is displayed.","section":"Fig. 3 caption"},{"comment":"The sentence about processing speed reports that NLSF took 6 hours for 598 pixels 'covering only the tumor area' while MFliNet processed 90,480 pixels; since this appears in the phantom section, please clarify whether the NLSF analysis was restricted to the phantom fluorescence embeddings or to the in-vivo tumor region.","section":"Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially suitable for the journal after revision, but the central robustness claim hinges on whether the training IRFs actually cover the height- and anatomy-dependent offsets observed at test time. The authors should be asked to supply the missing numerical results and either retrain with shifted IRFs or explicitly test out-of-distribution generalization. The absence of an independent lifetime standard is also a significant gap for a journal that publishes methodological validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a legitimate step forward for macroscopic FLI, but the central robustness claim is softer than it looks. The model is new—first to feed pixelwise IRF alongside the TPSF into a differential-transformer encoder-decoder for lifetime estimation—and the phantom and in-vivo results are encouraging: MFliNet matches NLSF within about 0.04 ns, and the speed gain is real (63 s for 90k pixels vs. 6 h for 598 pixels). Benchmarking against FLI-Net and a same-architecture standard transformer is a good design choice.\n\nThe soft spot I care most about is the training distribution. The IRFs used for training were captured from a flat diffuser on the imaging table; the paper never says height-shifted IRFs were included. The phantom experiment shows 40–160 ps shifts with height, and in-vivo IRFs from different organs show similar offsets. If those offsets were not in the training set, then the model's apparent invariance to them is an untested generalization. It could still work—if the network truly learns to use the IRF as an input, the deconvolution is shift-invariant—but the paper needs to show that explicitly, either by training with shifted IRFs or by testing on a controlled shift set. As written, that's a load-bearing gap.\n\nThe evaluation also has the usual weaknesses: NLSF is the reference, but NLSF uses the same bi-exponential convolution model as the training data, so agreement is consistency, not absolute accuracy. No independent lifetime standard, no statistical tests, and the phantom MFliNet values are described only qualitatively in the text. No code or data released, so the numbers can't be checked. One minor note: the speed comparison is apples-to-oranges, but the qualitative conclusion is probably fine.\n\nBottom line: this is a paper for the FLI community, especially people working on fit-free real-time processing for surgical guidance or drug-target imaging. It deserves a serious referee. I'd send it to review with a request for code/pretrained model, an explicit statement (or experiment) covering IRF shift in training, quantitative phantom numbers with error bars, and ideally a non-NLSF ground truth. With those additions, the robustness claim would actually be supported.","headline":"A real engineering advance for real-time FLI, but the headline claim about robustness to surface-height IRF shifts is not actually demonstrated by the training data or the validation.","tokens_in":10714,"tokens_out":2713,"would_cite":true,"duration_ms":27719,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Feeding each pixel's instrument-response function alongside the photon histogram into a differential-transformer encoder-decoder yields fluorescence lifetime estimates that stay accurate when the sample surface moves through the imaging…","keywords":["fluorescence lifetime imaging","macroscopic FLI","instrument response function","differential transformer","deep learning","time-resolved imaging","IRF offset","in-vivo imaging"],"falsifier":"Check the training dataset for IRFs with the 40–160 ps offsets observed in the phantom; then measure MFliNet's lifetime error on a phantom placed at heights whose IRF shift lies outside that range or whose IRF shape changes (not just shifts). If the error grows with unseen offsets or with IRF shape changes, the central robustness claim is limited to interpolation over the training IRF distribution.","tokens_in":9697,"feed_emoji":"⏱️","tokens_out":5582,"duration_ms":44979,"temperature":0.7,"pith_summary":"This paper introduces MFliNet, a deep-learning model for macroscopic fluorescence lifetime imaging that takes the pixelwise instrument response function (IRF) as an explicit second input beside the measured photon time-of-arrival histogram. The claim is that this extra input, combined with a differential-attention transformer, lets the network absorb the temporal offsets caused by surface topography, so lifetime estimates do not drift when the sample is not flat. The paper supports this with a five-step ladder phantom whose heights impose IRF shifts up to 160 ps, where MFliNet stays within about 0.04 ns of nonlinear least-squares fitting while a standard DL baseline and offset-free fitting drift with height. In two HER2+ tumor xenografts, MFliNet reproduces the fitted short and long lifetimes to within a few hundredths of a nanosecond while processing 90,480 pixels in 63 seconds versus roughly 6 hours for the fitter. If correct, this removes a known bias that has limited deep-learning FLI to planar samples and opens the method to in-vivo and surgical imaging.","feed_headline":"One extra input fixes height bias in fluorescence lifetime imaging","feed_subtitle":"Network using each pixel's instrument response matches slow least-squares fits on phantoms and tumors.","key_machinery":"The load-bearing mechanism is the differential attention layer, defined as DiffAttn(X) = softmax(Q1K1^T/$\\sqrt$(dk))V1 − λ·softmax(Q2K2^T/$\\sqrt$(dk))V2, where λ is a learnable scalar. Subtracting two attention maps cancels the noise floor and sharpens attention onto the informative features of the input pair, which in this application are the early-arrival variations that encode the IRF offset. The encoder-decoder processes the TPSF and pixelwise IRF as an input pair, and three parallel output heads regress the short lifetime, long lifetime, and fractional amplitude.","core_discovery":"MFliNet's central discovery is that making the IRF a per-pixel model input, rather than a fixed global correction, de-biases deep-learning lifetime estimation under realistic depth-of-field variations. The model uses a differential attention mechanism that computes the difference between two softmax attention maps, scaled by a learned parameter, to concentrate on the parts of the TPSF–IRF pair that carry lifetime information and suppress noise. On the step-ladder phantom, MFliNet matches NLSF across all five heights, whereas FLI-Net errors grow with height and NLSF without offset correction systematically underestimates. In vivo, its short-lifetime estimates (0.52 ± 0.05 ns and 0.58 ± 0.04 ns for two HCC1954 xenografts) and long-lifetime estimates (1.19 ± 0.02 ns and 1.18 ± 0.03 ns) align with NLSF values to within measurement spread.","pith_inferences":["The paper does not state whether the training set included IRFs shifted by 40–160 ps like those seen in the phantom; if it did not, the demonstrated height invariance is an untested generalization rather than a learned behavior, and the robustness claim would need a dedicated generalization test.","The same architecture could be applied to other deconvolution problems where a per-pixel or per-channel instrument response is known, such as time-correlated single-photon counting arrays or pulse oximetry, since the differential attention is not specific to fluorescence.","A direct test of the mechanism would be to ablate the IRF input while keeping the differential transformer: the paper compares against a standard transformer with IRF input, but an ablation without IRF input would separate the benefit of the extra input from the benefit of differential attention."],"forward_implications":["Macroscopic FLI can be run in real time: 90,480 pixels analyzed in 63 seconds on a single dataset, versus roughly 6 hours for NLSF.","The model removes the need for manual offset correction in NLSF-style analysis, since the IRF input accounts for time-of-arrival shifts per pixel.","FLI becomes usable on non-planar samples such as whole animals and surgical fields, where surface height varies by centimeters.","The differential attention mechanism is the component that gives the gain: a same-architecture standard transformer trained on the same data performs worse across heights."],"supporting_citations":[{"why":"Supplies the FLI-Net baseline whose height-dependent bias motivates the work and against which MFliNet is benchmarked.","marker":"[28]"},{"why":"Supplies the differential attention mechanism that the MFliNet architecture is built on.","marker":"[33]"},{"why":"Describes the MFLI system and the system-derived noise model used to generate realistic training data.","marker":"[34]"},{"why":"Provides the AlliGator NLSF reference estimates used as ground truth in phantom and in-vivo comparisons.","marker":"[37]"},{"why":"Documents the IRF offset effects on fluorescence lifetime uncertainty that MFliNet is designed to handle.","marker":"[22]"}],"fun_headline_variants":["Pixelwise IRF input de-biases deep FLI on curved tissue","MFliNet: differential transformer with pixelwise IRF for unbiased lifetime imaging","Deep lifetime fits match NLSF on curved surfaces with per-pixel IRF","Height bias in fluorescence lifetime imaging fixed by per-pixel IRF input"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training IRFs were captured from a flat white diffuser on the imaging table, and the paper does not state whether height-shifted IRFs (40–160 ps) were included in training, so the model's robustness to those shifts may rest on an untested generalization rather than on having learned to handle them.","fun_headline_variants_meta":{"raw":{"variants":["Pixelwise IRF input de-biases deep FLI on curved tissue","MFliNet: differential transformer with pixelwise IRF for unbiased lifetime imaging","Deep lifetime fits match NLSF on curved surfaces with per-pixel IRF","Height bias in fluorescence lifetime imaging fixed by per-pixel IRF input"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000893,"raw_usage":{"total_tokens":3877,"prompt_tokens":1002,"completion_tokens":2875,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":2793}},"tokens_in":618,"tokens_out":2875,"duration_ms":20681,"temperature":1.0,"reasoning_tokens":2793,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:46:31.516285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the training dataset for IRFs with the 40–160 ps offsets observed in the phantom; then measure MFliNet's lifetime error on a phantom placed at heights whose IRF shift lies outside that range or whose IRF shape changes (not just shifts). If the error grows with unseen offsets or with IRF shape changes, the central robustness claim is limited to interpolation over the training IRF distribution.","supporting_citations":[{"cited_title":"Fast fit-free analysis of fluorescence lifetime imaging via deep learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the FLI-Net baseline whose height-dependent bias motivates the work and against which MFliNet is benchmarked."},{"cited_title":"Differential transformer,","cited_arxiv_id":null,"evidence_quote":"Supplies the differential attention mechanism that the MFliNet architecture is built on."},{"cited_title":"Venugopal,A small animal time-resolved optical tomography platform using wide-field excitation(Rensselaer Polytechnic Institute, 2011)","cited_arxiv_id":null,"evidence_quote":"Describes the MFLI system and the system-derived noise model used to generate realistic training data."},{"cited_title":"Alligator: A phasor computational platform for fast in vivo lifetime analysis,","cited_arxiv_id":null,"evidence_quote":"Provides the AlliGator NLSF reference estimates used as ground truth in phantom and in-vivo comparisons."},{"cited_title":"Experimental study of fluorescence lifetime uncertainty in time- gated iccd-based macroscopic fluorescence lifetime imaging,","cited_arxiv_id":null,"evidence_quote":"Documents the IRF offset effects on fluorescence lifetime uncertainty that MFliNet is designed to handle."}],"review_version":1}