{"id":"599d292c-ae5d-46b3-9f12-f95ab2236fe8","arxiv_id":"2502.08540","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of image quality assessment methods that argues for scenario-specific, interpretable, and practical metrics.","lead":"This paper surveys the many ways computers judge image quality, from simple formulas like PSNR to deep learning and Transformer models. It argues that different applications, such as medical imaging and portrait photography, need their own dedicated quality measures.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Universal-IQA conclusion rests on unrepresentative scenarios and ignores cross-domain metrics; missing comparative benchmark evidence makes the Sec. 3 claim under-supported.","rationale":"The survey provides a clear taxonomy and mostly accurate method descriptions, and its advice to design scenario-aware metrics is sensible. However, the headline conclusion is a strong universal negative that needs more evidence than five examples. My concern is not that the claim is false; it may be true. The issue is internal to the argument: the paper does not define 'universal' precisely, does not compare general and specialized metrics on common benchmarks, and omits widely used universal metrics and datasets. Its own observation that PSNR and SSIM dominate practice further suggests universal methods remain practical despite suboptimal performance. These gaps do not invalidate the survey, but they support the reader's CONDITIONAL verdict: the central claim should either be softened to a recommendation or hypothesis, or backed by cross-scenario validation. I therefore see no reason to change the reader's verdict; the condition is exactly that the authors add such evidence or narrow the claim. This is consistent with the reader's weakest assumption about representativeness, though I would extend it from 'scenarios' to 'scenarios and metrics/benchmarks.'","tokens_in":12897,"tokens_out":5113,"duration_ms":57547,"concrete_test":"Evaluate one general-purpose metric not covered by the survey (e.g., LPIPS or MUSIQ) on public datasets from all five surveyed scenario families: medical (e.g., structural MRI-IQA with expert ratings), dehazing (e.g., synthetic-hazy image sets with MOS), portrait (NTIRE 2024 Deep Portrait Quality Assessment), blur (e.g., LIVE or BID), and JPEG compression (LIVE). Compute per-scenario Spearman rank correlation between the metric and subjective scores. If one metric achieves acceptable correlation in all five (say rho > 0.8), the claim 'universal IQA is impractical' is contradicted; if it fails in one or more, the survey's conclusion is supported. This directly probes the representativeness assumption underlying Section 3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in Section 3 as 'proposing a universal IQA method applicable across all scenarios is impractical,' is a universal negative supported only by a handful of specialized methods (medical, dehazing, portrait, blur, JPEG) and one anecdote about motion blur. For the claim to hold, it must be that no single practical metric can serve all these and other scenarios. The survey never tests this: it omits standard general-purpose metrics (NIQE, LPIPS), omits benchmark datasets (LIVE, TID2013, KADID, PIPAL), and provides no quantitative comparison between general and specialized methods. It also implicitly equates 'universal' with a single fixed scalar criterion, ignoring multi-task or conditional designs that could combine a shared backbone with scenario-specific heads. In addition, the paper itself says PSNR and SSIM remain the predominant methods because of simplicity and interpretability (Section 3), which shows universal metrics are used in practice, weakening the impracticality claim. The load-bearing premise—that the reviewed scenarios are representative and that no cross-scenario method works—is therefore not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys image quality assessment (IQA) methods, organizing them into general-scene methods (statistical and machine-learning based, further split into model- and framework-based approaches) and specific-scene methods (medical imaging, dehazing, portrait, blur, and JPEG artifacts). It provides bibliographic tables with citation counts, chronological figures, and a taxonomy in Figure 1. The survey's central thesis is stated in Section 3: proposing a universal IQA method applicable to all scenarios is impractical, so the field should move toward scenario-specific, user-oriented, and interpretable metrics. The paper reports no new algorithms, experiments, or benchmark comparisons, and its conclusions are derived from the reviewed literature and the authors' own commentary.","tokens_in":13040,"tokens_out":5594,"duration_ms":55707,"significance":"The survey has a useful organizing principle—application scenario rather than distortion type or architecture—and it collects recent work, including NTIRE challenge entries, that is not always covered in earlier surveys. The chronological figures and citation tables provide a convenient starting point for newcomers to the field. However, the significance of the central claim is limited by its evidentiary basis: the claim is a universal negative about the entire IQA field, but the paper reviews only a small set of scenario families and does not include standard benchmark evidence or several widely used general-purpose metrics. Because the conclusion is framed as a field-level prescription, these omissions are not cosmetic. If the gaps are addressed, the survey could serve as a valuable roadmap; in its current form, the practical conclusion is not established.","major_comments":[{"comment":"The central claim that a universal IQA method is impractical is a universal negative, but the support offered is limited to the specific scene methods reviewed in Section 2.2 (medical, dehazing, portrait, blur, JPEG) and a single motion-blur/medical counterexample. No benchmark evidence is provided: the paper never compares general-purpose metrics against specialized methods on standard datasets such as LIVE, TID2013, KADID, or PIPAL, and it does not discuss cross-scenario generalization results of methods such as UNIQUE (cited in Section 2.1) that were specifically designed to generalize across distortions. As stated, the claim is not supported by the selected examples. Either the claim should be weakened to \"different scenarios impose different requirements that need to be considered in metric design,\" or the paper should add a comparative analysis showing failure of general-purpose metrics on representative scenarios.","section":"Section 3 (and Sections 2.2–2.3)"},{"comment":"The paper states that \"the predominant IQA methods in use continue to be the traditional PSNR and SSIM, primarily due to simplicity and interpretability.\" PSNR and SSIM are universal, content-agnostic metrics; if they remain the de facto standard, the assertion that a universal method is impractical requires qualification. The term \"universal\" is also never defined: it could mean a single fixed scalar criterion, a family of metrics sharing a common backbone, or a benchmark protocol. Multi-task or conditional architectures—e.g., a shared representation with scenario-specific regression heads—are not discussed, even though they would represent a middle ground between fully universal and fully bespoke methods. This ambiguity is load-bearing because the paper's final recommendation depends on ruling out such designs.","section":"Section 3, first paragraph"},{"comment":"The description of FSIM states that it \"selects phase congruency and gradient deviation (GD) as predictive features of quality.\" FSIM actually uses phase congruency and gradient magnitude (GM). This is a factual error in a core method description and should be corrected; the same paragraph's discussion of GMSD and PSIM already refers to gradient magnitude, so the inconsistency is internal.","section":"Section 2.1, HVS-Based Methods"},{"comment":"Several widely used general-purpose methods are absent from the taxonomy. NIQE (Mittal et al., 2013) does not appear in Table 5 or in the discussion of NR methods, even though it is one of the most common no-reference baselines; LPIPS appears only as a citation to [Zhang et al., 2018] for the value of deep features and is not named or classified as a perceptual metric. Given that the survey's thesis rests on a contrast between general and specific methods, omitting standard general-purpose baselines makes the contrast incomplete. A survey that explicitly aims for comprehensiveness should cover these methods and, ideally, report their benchmark performance.","section":"Section 2.1, Tables 5–8"}],"minor_comments":[{"comment":"The VSNR citation [Chandler and Hemami, ] is missing its publication year in both the text and the reference list; the reference also lacks volume and page information.","section":"References"},{"comment":"Several citations use the placeholder form \"and et al.\" (e.g., [Bosse and et al., 2017], [Egiazarian and et al., 2006]) rather than naming the first author and co-authors; this will cause formatting and indexing problems.","section":"References"},{"comment":"If the table is intended as a representative selection rather than an exhaustive list, the caption should say so; the absence of NIQE is otherwise conspicuous for a survey of NR-IQA methods.","section":"Table 5"},{"comment":"The acronym IFC is used without expansion; the reader must infer that it stands for Information Fidelity Criterion.","section":"Section 2.1, NSS-Based Methods"},{"comment":"Some entries in the figures use different naming conventions than the corresponding table entries, and the bracket notation explained in the captions is applied inconsistently; aligning the figure labels with the tables would improve readability.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an early preprint with incomplete proofreading, most visibly in the reference formatting and the missing year for VSNR. For a survey with a field-level recommendation, the absence of benchmark datasets and standard baselines such as NIQE and LPIPS is a more serious concern than novelty. I recommend major revision rather than rejection because the organizational framework is sound and the identified weaknesses are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, readable survey of IQA methods organized by application scenario, with some useful tables and a sensible discussion of why generic metrics like PSNR/SSIM persist. But it is not the comprehensive survey it claims to be, and its central argument—that universal IQA is impractical—is asserted rather than shown. The paper is fine as an orientation for a newcomer, but it needs work before I'd trust it as a reference.\n\nWhat's genuinely useful: the taxonomy (general vs. specific scenes, then statistical/ML, then CNN/Transformer/framework) is clear, and the chronological figures give a quick sense of how the field moved from statistics to deep learning. The specific-scene section covers medical, dehazing, portrait, blur, and JPEG, which are real applications with distinct quality criteria. The discussion of why PSNR/SSIM remain popular—simplicity and interpretability—is honest and matches practice. The tables with citation counts are a nice touch, though they age quickly.\n\nSoft spots, in order of severity. First, coverage gaps: NIQE and LPIPS are both absent as named metrics. NIQE is one of the most used no-reference metrics, and LPIPS is the de facto perceptual metric for generated images in the last five years. Their omission from a survey claiming to cover 'both general and specific IQA methodologies' is hard to justify. Second, the central claim rests on a few examples. The paper never defines what 'universal' means, never considers multi-task or conditional designs that could combine a shared backbone with scenario-specific heads, and provides no quantitative comparison between general and specialized metrics on the same benchmarks. The one concrete example—motion blur desirable in portraits but bad in medical images—is fine as an intuition but doesn't establish a universal negative. Third, citation errors: the VSNR reference is missing its year, and FSIM is described as using 'gradient deviation' when the paper's own description elsewhere says 'gradient magnitude'—that's a real slip. Fourth, no benchmark datasets (LIVE, TID2013, KADID, PIPAL) are mentioned, which any serious IQA survey should at least list.\n\nThe paper's own statement that PSNR and SSIM dominate practice doesn't contradict the scenario-specificity thesis—it just means simplicity wins—but the thesis would need a broader set of scenarios and some evidence that generic metrics actually fail in those scenarios.\n\nBottom line: this is a candidate survey but not a definitive one. With a proper revision that fills the gaps and softens the universal claim, it could be a useful reference. For now, I'd point a new student to it as an overview but not cite it as authoritative.\n\nRecommendation: send it to review with major revisions expected. The topic is important, the structure is sound, and the flaws are fixable. A serious referee would catch the errors and push for better coverage.","headline":"A readable but not comprehensive IQA survey whose scenario-specificity thesis is argued more than demonstrated; usable as an entry point, not as a definitive reference.","tokens_in":13562,"tokens_out":2011,"would_cite":false,"duration_ms":20442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a universal image quality assessment method is impractical because different application scenarios demand conflicting quality criteria, so the field should move toward scenario-specific, user-oriented metrics.","keywords":["image quality assessment","IQA survey","scenario-specific IQA","blind image quality assessment","full-reference IQA","deep learning IQA","perceptual metrics","PSNR and SSIM"],"falsifier":"Take a general-purpose IQA model and specialized IQA models for medical imaging, dehazing, and portraiture, and evaluate all of them on human opinion scores from each domain; if one general model ranks first in every domain or matches the specialized models everywhere, the claim that universal IQA is impractical would be refuted. A weaker falsifier would be finding a single quality criterion (for example, sharpness or structural fidelity) that human raters consistently prioritize across all these domains regardless of scenario.","tokens_in":12687,"feed_emoji":"🖼️","tokens_out":5443,"duration_ms":49903,"temperature":0.7,"pith_summary":"This survey argues that no universal image quality assessment (IQA) method can serve all application scenarios, because different scenes impose conflicting quality criteria. It reaches this conclusion by organizing the field into general-purpose methods and scenario-specific methods, and by showing how the same distortion can be acceptable or even desirable in one context and unacceptable in another. A sympathetic reader should care because the claim redirects the field's research agenda: instead of chasing a single metric that works everywhere, IQA should be designed per application, from the user's perspective, with attention to practicality and interpretability. The survey also observes that despite progress in deep learning, traditional metrics like PSNR and SSIM remain the most used, precisely because they are simple and interpretable.","feed_headline":"Survey: one image-quality metric cannot fit all scenes","feed_subtitle":"Medical imaging prizes lesion visibility, portraits prize aesthetics; IQA should become scenario-specific and user-centric.","key_machinery":"The mechanism carrying the argument is the survey's scenario taxonomy, which splits IQA methods into general-scene methods (statistical, HVS-based, transform-domain, NSS-based, and machine-learning approaches) and specific-scene methods (medical, dehazing, portrait, and specific distortions). The taxonomy does conceptual work: by placing methods side by side, it exposes that their evaluation criteria are not commensurable. The load-bearing contrast is the same distortion receiving opposite valuations in different scenes, such as motion blur being aesthetically acceptable in portraits but diagnostically harmful in medical images, which drives the conclusion that quality criteria must be derived from the image user's perspective.","core_discovery":"The paper's central claim is that \"proposing a universal IQA method applicable across all scenarios is impractical, as different scenes demand contrasting quality criteria.\" The evidence comes from reviewing IQA methods in medical imaging, dehazing, portrait photography, blur, and JPEG compression: medical imaging prioritizes lesion visibility and local topology, portraits prioritize the facial region and aesthetics, dehazing evaluation must account for contrast and color adjustment, and deblurring evaluation must detect artifacts like ringing. In each case, a generic metric misses what actually matters for the user. The paper therefore concludes that future IQA should be scenario-specific and user-centric, and that deep learning metrics will only see adoption if they become more interpretable and easier to use, not merely more accurate.","pith_inferences":["Beyond the paper: the argument implies that leaderboard-style comparisons of general IQA metrics across mixed datasets may be measuring the wrong target, since the criteria themselves shift by scenario.","Beyond the paper: a testable extension would be constructing paired metrics with identical architectures, one specialized for medical images and one for portraits, and measuring how much cross-scenario agreement drops; the paper's claim predicts a large drop.","Beyond the paper: the survey's examples leave out video, point clouds, and AI-generated images; testing whether those domains also demand dedicated quality criteria would either extend or bound the impracticality claim."],"forward_implications":["IQA research should shift from designing one universal metric to designing per-scenario metrics, with evaluation criteria derived from the user's task.","Adoption of deep-learning IQA methods will depend on improving interpretability and ease of deployment, since the field still leans on PSNR and SSIM for those reasons.","Specialized metrics should be built around the distortions and artifacts that matter in each application, such as ringing in deblurring, contrast and color changes in dehazing, and facial-region quality in portraits.","Quality criteria should be defined from the perspective of image users, for example aesthetic criteria like composition and lighting for aesthetic evaluation."],"supporting_citations":[{"why":"Supplies the aesthetic criteria (composition, lighting, focus, color) used to argue quality criteria are user- and scene-dependent.","marker":"[Luo and Tang, 2008]"},{"why":"Shows medical image quality is best captured by a local topological measure (Euler number) rather than global metrics, grounding the medical scenario.","marker":"[Rosen et al., 2018]"},{"why":"Demonstrates dehazing needs a dedicated metric covering structure restoration, color rendition, and over-enhancement.","marker":"[Min et al., 2019]"},{"why":"Provides portrait-quality challenge methods that split facial and background analysis, grounding the portrait scenario.","marker":"[Chahine et al., 2024]"},{"why":"Introduces artifact-specific features for deblurring, grounding the claim that deblurring quality needs specialized evaluation.","marker":"[Liu et al., 2013]"},{"why":"Shows blur assessment requires HVS-based just-noticeable-blur modeling rather than generic metrics.","marker":"[Ferzli and Karam, 2009]"},{"why":"Shows JPEG-specific metrics capture blocking and blurring effects better than conventional IQA methods.","marker":"[Gore and Gupta, 2015]"}],"fun_headline_variants":["No universal image quality metric: scenes demand different criteria","IQA must become scene-specific and user-centric","Universal image quality metric? Survey says impractical","One metric can't judge medical, portrait, and dehazed images alike","Scene-specific, user-centric IQA: the future of image quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that the scenarios it reviews—medical imaging, dehazing, portraits, blur, and JPEG compression—are representative enough that their conflicting criteria generalize to all possible image applications; if other domains such as video, point clouds, or AI-generated images turn out to share one compatible quality standard, the case against universal IQA weakens.","fun_headline_variants_meta":{"raw":{"variants":["No universal image quality metric: scenes demand different criteria","IQA must become scene-specific and user-centric","Universal image quality metric? Survey says impractical","One metric can't judge medical, portrait, and dehazed images alike","Scene-specific, user-centric IQA: the future of image quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000877,"raw_usage":{"total_tokens":3739,"prompt_tokens":837,"completion_tokens":2902,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":2822}},"tokens_in":453,"tokens_out":2902,"duration_ms":21342,"temperature":1.0,"reasoning_tokens":2822,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T04:40:46.989231+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a general-purpose IQA model and specialized IQA models for medical imaging, dehazing, and portraiture, and evaluate all of them on human opinion scores from each domain; if one general model ranks first in every domain or matches the specialized models everywhere, the claim that universal IQA is impractical would be refuted. A weaker falsifier would be finding a single quality criterion (for example, sharpness or structural fidelity) that human raters consistently prioritize across all these domains regardless of scenario.","supporting_citations":[{"cited_title":"Photo and video quality evaluation: Focusing on the subject","cited_arxiv_id":null,"evidence_quote":"Supplies the aesthetic criteria (composition, lighting, focus, color) used to argue quality criteria are user- and scene-dependent."},{"cited_title":"Quan- titative assessment of structural image quality","cited_arxiv_id":null,"evidence_quote":"Shows medical image quality is best captured by a local topological measure (Euler number) rather than global metrics, grounding the medical scenario."},{"cited_title":"Quality evaluation of image dehazing methods using syn- thetic hazy images","cited_arxiv_id":null,"evidence_quote":"Demonstrates dehazing needs a dedicated metric covering structure restoration, color rendition, and over-enhancement."},{"cited_title":"Deep portrait quality assessment","cited_arxiv_id":null,"evidence_quote":"Provides portrait-quality challenge methods that split facial and background analysis, grounding the portrait scenario."},{"cited_title":"A no-reference metric for evaluating the quality of motion deblurring","cited_arxiv_id":null,"evidence_quote":"Introduces artifact-specific features for deblurring, grounding the claim that deblurring quality needs specialized evaluation."},{"cited_title":"A no-reference objective image sharpness metric based on the notion of just noticeable blur (jnb)","cited_arxiv_id":null,"evidence_quote":"Shows blur assessment requires HVS-based just-noticeable-blur modeling rather than generic metrics."},{"cited_title":"Full reference image quality metrics for jpeg compressed im- ages","cited_arxiv_id":null,"evidence_quote":"Shows JPEG-specific metrics capture blocking and blurring effects better than conventional IQA methods."}],"review_version":1}