{"id":"5128fa19-54f2-40f0-87ff-291c2aa19c27","arxiv_id":"2501.11520","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review of fundus image quality assessment and enhancement, covering handcrafted and learning-based paradigms and their clinical deployment.","lead":"A survey of algorithms that assess and improve the quality of retinal photographs, organized into assessment and enhancement paradigms. The paper maps the field's methods, datasets, and deployment challenges for clinical screening.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No reproducible search protocol is described, so the claim of a 'systematic' review is not auditable and coverage may be biased toward the authors' own prior work.","rationale":"The reader's weakest assumption exactly matches the key load-bearing concern: the absence of a reproducible systematic search protocol. My stress-test of the full text confirms that no search strategy, inclusion/exclusion criteria, or PRISMA-style flow diagram is presented despite the title and abstract claiming a systematic review. The reference list contains a high density of self-citations (e.g., refs 14, 23, 24, 34, 68, 91, 101, 102), which is a red flag for selection bias in a survey. The paper has useful content: the imaging-system background, the FR/NR IQA split, the handcrafted/learning IQE split, and the discussion of clinical deployment challenges are all clearly organized and informative. The integrated IQA+IQE perspective is a genuinely useful framing. However, the 'systematic' label is not justified by the manuscript as written. The appropriate remedy is to either add a detailed Methods section describing the search, screening, and data-extraction procedure, or to retitle the paper as a narrative review. Since this is a fixable methodological gap, a CONDITIONAL verdict is appropriate, matching the reader's decision. I also noticed minor errors (e.g., PSNR claimed to have 'greater stability' than MSE when PSNR is a monotone transform of MSE; BRISQUE incorrectly expanded; duplicate label in Fig. 1) but these do not affect the core claim as much as the missing methodology. Therefore I agree with the reader's assessment and recommend no change to the verdict.","tokens_in":28178,"tokens_out":5163,"duration_ms":52344,"concrete_test":"Perform an independent structured search in PubMed and IEEE Xplore for 2015-2024 using queries (fundus OR retinal) AND (image quality assessment OR image enhancement OR restoration). Compare the retrieved papers against the review's reference list and compute recall (the fraction of retrieved papers that are cited). Additionally, compute the proportion of self-citations among all references. If recall is below 70% or if the missing set includes seminal or high-citation works that would alter the taxonomy or the qualitative judgments, the systematicity claim is not supported. Alternatively, ask the authors to provide the search protocol and selection log; if it cannot be provided or the same queries do not reproduce the reference list, the review should be reframed as a narrative review.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the paper provides a thorough, systematic review integrating fundus IQA and IQE. For this claim to hold, the set of surveyed papers must be representative of the literature. However, no systematic search methodology is documented: the manuscript does not state which databases were queried, which keywords were used, what inclusion or exclusion criteria were applied, or how the references were selected. This makes the corpus non-reproducible and the coverage potentially biased. Inspection of the reference list shows a substantial fraction of citations to the authors' own prior work (e.g., refs [14], [23], [24], [34], [68], [91], [101], [102]), which raises the risk that the taxonomy and the qualitative strengths/weaknesses assigned to methods are shaped by convenience rather than by an unbiased survey. Because the paper's value proposition is its integrated, comprehensive perspective, an unrepresentative sample would directly weaken the central claim. This is a problem of external validity: the review may still be a legitimate narrative summary, but it is not a systematic review as claimed in the title and abstract.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a review of fundus image quality assessment (IQA) and enhancement (IQE). It first describes the fundus camera imaging system and common sources of image degradation, then categorizes IQA methods into full-reference and no-reference approaches, subdividing NR-IQA into numerical metrics, end-to-end learning-based methods, and comprehensive methods that introduce intermediate sub-metrics. For IQE, the paper distinguishes handcrafted methods (pixel-based processing, image filtering, statistical priors) from learning-based methods (paired and unpaired data paradigms) and further separates real-world from synthesized paired data. The latter half of the paper discusses practical deployment challenges such as data availability, generalizability, and interpretability, reviews applications of IQA and IQE in diagnosis and segmentation, and closes with six future research directions. The paper is written as a tutorial-style survey with equations for standard metrics and enhancement operations, tables of representative models and datasets, and example figures illustrating degradations.","tokens_in":28346,"tokens_out":6409,"duration_ms":66034,"significance":"If the manuscript's scope claims are taken at face value, it fills a gap by co-reviewing fundus IQA and IQE, which earlier surveys treat separately. The tutorial material is solid: standard formulas for MSE, PSNR, SSIM, histogram equalization, CLAHE, gamma correction, the dark-channel model, and CycleGAN losses are presented in a self-contained way, and the tables of representative IQA/IQE models and datasets provide a useful quick reference. The discussion of synthesized versus real-world paired data, domain adaptation/generalization, and the use of IQA to guide enhancement is timely and practically relevant. However, the contribution is currently weakened by two structural issues: the 'systematic review' claim is not backed by a reproducible search protocol, and the promised integrated perspective is not fully realized because the IQA and IQE halves are largely separate surveys with only limited cross-referencing. These issues are fixable within the manuscript's scope and are the main reasons for requesting major revision.","major_comments":[{"comment":"The manuscript is presented as a 'Systematic Review' in the title and abstract, but no systematic methodology is described anywhere. There is no statement of the databases queried, the search strings used, the inclusion/exclusion criteria, the screening procedure, or the time window of coverage, and no PRISMA-style flow diagram or study-selection protocol. The reference list also contains a substantial number of the authors' own prior publications (e.g., [14], [23], [24], [34], [68], [91], [101], [102]), which increases the risk that the corpus is a convenience sample rather than an auditable comprehensive survey. Because the paper's central value proposition is that it provides a thorough, comprehensive, integrated review, the absence of a reproducible corpus is a load-bearing omission. The authors should either add a complete search and selection protocol, or explicitly reframe the manuscript as a narrative review in the title and abstract.","section":"Title, Abstract, Section I"},{"comment":"The Introduction and Abstract promise an integrated perspective grounded in the 'intimate relationship' between IQA and IQE, but the body of the paper does not deliver this synthesis. Sections III and IV are two largely independent surveys, and the connections between the two fields appear only in isolated remarks, such as Hou et al. [90] using IQA to guide enhancement and the quality-feature fusion examples in Section V.B. The paper would substantially strengthen its central claim if it made the interplay explicit, for example by mapping the sub-metrics used in comprehensive NR-IQA (Section III.B.3) to the degradation targets addressed by IQE methods, or by discussing how the same datasets and quality labels support both tasks. Without such a conceptual bridge, the integrative contribution asserted in the Introduction is not fully demonstrated.","section":"Sections III-V"}],"minor_comments":[{"comment":"The caption repeats '(b)' for the second and third images; the third image should be labeled '(c)'.","section":"Fig. 1 caption"},{"comment":"'Biaswas et al. [41]' is a misspelling of Biswas; later in the same paragraph 'NF-IQA' should be 'NR-IQA'.","section":"Section III.B.3"},{"comment":"The prose definition 'CFD(l) = sum_{l=0}^{L-1} H(l)/N' is not the cumulative distribution function; it should be 'CDF(l) = sum_{i=0}^{l} H(i)/N'. The actual computation in Eq. (7) uses the correct lower limit, so only the defining sentence needs correction.","section":"Section IV.A.1, Eq. (7) text"},{"comment":"'Scematics' should be 'Schematics'.","section":"Fig. 6 caption"},{"comment":"There are two typos: 'funds image data' should be 'fundus image data', and 'paried' should be 'paired'.","section":"Section IV.B"},{"comment":"The column headings 'Instance #' and 'Quality #' are ambiguous; the table would be clearer if the headers were 'Images' and 'Quality levels'. The superscript footnote markers for ODIR and iSee are also placed awkwardly in the dataset-name column.","section":"Table IV"},{"comment":"References [83] and [103] appear to be the same paper (Guo et al., 'Bridging synthetic and real images: a transferable and multiple consistency aided fundus image enhancement framework'); one should be removed and the in-text citations renumbered and checked.","section":"References"},{"comment":"The 'Aspect' entry for SSIM is listed as 'Pixel-level and structure fidelity', but Eq. (3) defines SSIM in terms of luminance, contrast, and structure; consider revising to 'Luminance, contrast, and structure fidelity' for consistency.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The duplicate references [83]/[103] and the high density of self-citations in the absence of a documented selection protocol are editorial concerns worth monitoring during revision. If the authors add a transparent search and screening methodology, or reframe the paper as a narrative review, the manuscript could become suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the integrated IQA+IQE framing is genuinely useful, and the deployment discussion is the best part. But \"Systematic Review\" overstates what's here; no search protocol is described, so the coverage is not auditable.\n\nWhat's new: the paper isn't a new method, it's a synthesis. The taxonomy (FR vs NR-IQA; handcrafted vs learning-based IQE) is standard but applied consistently. The tables summarizing datasets and representative models are handy. The sections on data availability, domain adaptation/generalization, and interpretability are practical and readable, and they reflect real challenges in clinical deployment. The future outlook items are sensible, if speculative.\n\nSoft spots: The title and abstract promise a systematic review, but the methods for selecting papers are never stated. No databases, keywords, inclusion/exclusion criteria, or PRISMA-style flow. That's a legitimate gap. For a survey, the corpus is not reproducible, and the reader can't judge coverage bias. That is the main problem. The editorial glitches (duplicate '(b)' in Fig. 1 caption, 'Biaswas' for Biswas, 'Scematics' for Schematics) are minor but should be fixed. Self-citations are noticeable but the authors are among the main contributors to fundus enhancement, so citing their own work is not obviously distorting. I'm more troubled by the lack of critical comparison: the review lists methods but rarely discusses comparative performance numbers. That's typical for a narrative review, though, and doesn't undermine the structural value.\n\nWho it's for: newcomers wanting a map of the field, and practitioners looking for a compact reference on data and deployment pitfalls. It won't change a research direction, but it's a fair overview.\n\nRecommendation: send it to peer review, with a request that the authors either document the search methodology or revise the title to 'A Review' rather than 'A Systematic Review.' Fix the typos. That would make it a solid contribution to the literature.","headline":"A useful integrated narrative review of fundus IQA and IQE, but the 'Systematic Review' label overpromises given no documented search protocol.","tokens_in":28897,"tokens_out":2201,"would_cite":false,"duration_ms":23147,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that assessing and enhancing fundus image quality should be treated as one problem, not two, and maps the algorithms, datasets, and deployment hurdles that follow.","keywords":["fundus photography","image quality assessment","image quality enhancement","no-reference image quality","full-reference image quality","domain adaptation","clinical deployment","systematic review"],"falsifier":"A concrete check would be to run a documented, reproducible literature search (fixed databases, query strings, inclusion criteria) for fundus IQA and IQE papers and compare the resulting corpus with the tables in this review; if substantial clusters of published methods fall outside the taxonomy, or if the integrated-perspective premise fails because few papers actually combine IQA with IQE, the review’s central claims would need revision.","tokens_in":27966,"feed_emoji":"👁️","tokens_out":5481,"duration_ms":55743,"temperature":0.7,"pith_summary":"This paper is a systematic review whose central claim is that fundus image quality assessment (IQA) and enhancement (IQE) are two halves of one clinical problem and should be reviewed together, not in separate surveys. The authors note that fundus photography captures clear critical structures in fewer than half of patients in some settings, and roughly a fifth of clinical images are unsuitable for automated diagnosis, which makes quality control a practical bottleneck for eye screening. They organise IQA into full-reference methods (which compare against an ideal image) and no-reference methods (which score quality without one), and IQE into handcrafted techniques versus learning-based models trained on paired or unpaired data. Across both tasks they identify shared datasets, shared metrics, and a common set of deployment obstacles—data scarcity, annotation inconsistency, domain shift, and interpretability—and argue that future progress lies in embedding medical knowledge, interpreting image interferences, decoupling visual from semantic content, and adapting continuously to each clinic.","feed_headline":"Assessing and enhancing eye-scan quality is one pipeline","feed_subtitle":"The review unifies the two tasks under one taxonomy and maps the data, domain-shift, and interpretability hurdles for clinics.","key_machinery":"The organising device is a two-axis taxonomy: IQA methods are classified by whether they require a reference image (full-reference versus no-reference) and IQE methods by whether they are handcrafted or learning-based, with learning-based IQE further split into paired-data and unpaired-data training. Within this taxonomy the load-bearing technical objects are the numerical metrics (MSE, PSNR, SSIM and its variants; BRISQUE, NIQE, PIQE), the CycleGAN unpaired translation architecture with adversarial and cycle-consistency losses, and the imaging-prior models such as the dark channel prior adapted from dehazing to cataractous fundus images. These objects carry the argument because they show how assessment and enhancement share a common vocabulary: the same metrics used to score quality are used to evaluate enhancers, and the same degradation models used to synthesise training pairs for IQE are the distortions that IQA must detect. The taxonomy is what lets the review claim that IQA and IQE are one field rather than two.","core_discovery":"The central claim is that the field should be understood as a single pipeline with two cooperating functions. IQA determines whether a captured fundus image is usable; IQE restores images that are degraded but recoverable; together they decide the flow of accept, enhance, or recapture before diagnosis. The review systematises IQA by reference dependence (numerical and learning-based full-reference methods; numerical, end-to-end, and comprehensive no-reference methods) and IQE by technique origin (pixel processing, filtering, and statistical priors versus paired- and unpaired-data deep models), and it argues that the two literatures are converging: handcrafted priors such as the dark channel or green channel are being embedded into deep enhancers, quality sub-metrics are being predicted alongside overall grades for interpretability, and multi-task systems couple quality assessment with disease grading. On this view the main obstacles to clinical use are data availability, inconsistent quality labels across datasets, domain shift, and black-box behaviour, and the next advances will come from medical knowledge embedding, interference interpretation, information decoupling, personalisation, continual optimisation, and cross-modality generalisation.","pith_inferences":["A testable extension the paper leaves implicit is an explicit ‘enhanceability’ label: instead of binary good or bad quality grades, datasets could carry a third label for images that are degraded but restorable, which would make the IQA–IQE loop directly trainable.","The integrated perspective suggests that evaluation protocols for IQE should be standardised around downstream clinical tasks (segmentation, grading) rather than only around perceptual metrics, since the review repeatedly ties quality to diagnostic utility.","One could check the review’s central claim empirically by measuring how often IQA and IQE methods are actually co-cited or jointly trained in the literature; if the two literatures barely overlap in practice, the ‘intimate relationship’ may be a design goal rather than a current fact.","The emphasis on source-free domain adaptation and continual learning implies that post-deployment data, normally discarded, could become the primary resource for keeping quality models aligned with each clinic’s imaging conditions."],"forward_implications":["Treating IQA and IQE as one pipeline suggests that future systems should be built as closed loops, where an assessment module decides between accept, enhance, or recapture before diagnosis.","Because both tasks share degradation models and datasets, progress in synthesizing realistic low-quality fundus images should improve both IQA training and IQE training at once.","The review implies that handcrafted priors (dark channel, green channel, structure) will increasingly be embedded inside learning-based models to gain interpretability without giving up performance.","Domain adaptation and domain generalization, not just benchmark accuracy, become the decisive criteria for whether an IQA and IQE method can be deployed in a new clinic.","Multi-task models that couple quality assessment with disease grading or vessel segmentation are a natural end point of the integrated view, since they let quality features act as inductive bias for diagnosis."],"supporting_citations":[{"why":"Provides the earlier fundus IQA survey that this review updates and explicitly integrates with enhancement rather than treating separately.","marker":"[12]"},{"why":"Supplies the recent review of single fundus image restoration that this paper pairs with IQA to form its integrated scope.","marker":"[13]"},{"why":"Is the broad fundus deep-learning review whose limited IQA and IQE coverage motivates the gap this paper addresses.","marker":"[15]"},{"why":"Offers the recent performance-focused comparison of fundus IQA algorithms that the review positions as narrower than its combined treatment.","marker":"[18]"},{"why":"Supplies the cataract degradation model, domain adaptation approach, and paired RCF dataset used as a recurring IQE example.","marker":"[23]"},{"why":"Introduces the EyeQ dataset whose three-level quality labels and color-space features anchor the no-reference IQA and quality-label discussion.","marker":"[39]"},{"why":"Provides the CycleGAN unpaired image-to-image translation framework on which most unpaired IQE models in the review are built.","marker":"[84]"}],"fun_headline_variants":["Fundus IQA and IQE: one pipeline for eye-scan quality","Unifying eye-scan quality assessment and enhancement","Assess, enhance, or recapture: fundus image pipeline","Fundus image quality: one review, two tasks, one pipeline","Tackling fundus image degradation with combined IQA and IQE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the papers it surveys, grouped under its full-reference versus no-reference and handcrafted versus learning-based taxonomy, are representative of the whole field, and that its qualitative claims about their strengths and weaknesses are accurate; because no reproducible search protocol is described, that coverage cannot be audited or independently reconstructed.","fun_headline_variants_meta":{"raw":{"variants":["Fundus IQA and IQE: one pipeline for eye-scan quality","Unifying eye-scan quality assessment and enhancement","Assess, enhance, or recapture: fundus image pipeline","Fundus image quality: one review, two tasks, one pipeline","Tackling fundus image degradation with combined IQA and IQE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2862,"prompt_tokens":937,"completion_tokens":1925,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":1845}},"tokens_in":553,"tokens_out":1925,"duration_ms":13937,"temperature":1.0,"reasoning_tokens":1845,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:08:23.347841+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to run a documented, reproducible literature search (fixed databases, query strings, inclusion criteria) for fundus IQA and IQE papers and compare the resulting corpus with the tables in this review; if substantial clusters of published methods fall outside the taxonomy, or if the integrated-perspective premise fails because few papers actually combine IQA with IQE, the review’s central claims would need revision.","supporting_citations":[],"review_version":1}