{"id":"a236e445-fd3b-4ddc-9565-44301bae382d","arxiv_id":"2506.18323","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"LucentVisionNet combines multi-scale spatial attention, depthwise separable convolutions, and a MUSIQ-based perceptual loss for zero-reference low-light enhancement, but the claimed state-of-the-art results rest on invalid averaging of raw metrics.","lead":"This paper proposes LucentVisionNet, a zero-reference deep network that brightens low-light images without paired training data. The authors claim it outperforms existing supervised, unsupervised, and zero-shot enhancers, but the headline evidence relies on statistically invalid metric averaging.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'consistently outperforms' claim rests on Average rows that average raw metric scores with incompatible scales (NIMA roughly 3-5, PaQ2PiQ roughly 60-77, MANIQA roughly 0.6), so the averages are scale-dominated and per-metric results are actually mixed.","rationale":"The reader's verdict is REJECT with high confidence, and my independent read agrees. The most load-bearing problem is the statistical construction of the Average rows. The claim of consistent superiority is operationalized only through those rows for no-reference metrics, and the rows are invalid because they average raw scores on different scales without normalization. The failure is concrete and visible in the tables: per-metric rankings are mixed, with LucentVisionNet winning some metrics while losing others, yet it always tops the Average row. I do not see a need to move the verdict: the current evidence does not support acceptance, and the paper would need a corrected evaluation rather than a minor revision. I also note the reader's secondary concern about the MUSIQ-AVA training loss transferring to MUSIQ-Koniq evaluation; that is plausible but not the decisive issue, because even with that set aside the averaging problem remains. The proposed test is inexpensive and uses only published table data, so it can settle the question directly. No code or significance testing is needed to expose the scale-averaging artifact, although public code and error bars would strengthen any resubmission.","tokens_in":22026,"tokens_out":7141,"duration_ms":62012,"concrete_test":"Recompute every Average row in Tables 2-8 after per-metric normalization: for each metric within one table, convert the raw scores of all methods to [0,1] via min-max scaling or z-scores, then average the normalized scores. Also report per-metric win counts for LucentVisionNet. If the normalized average does not place LucentVisionNet first in every dataset, or if it wins fewer than half the individual metrics, the consistently outperforms claim is not supported. This check needs only the numbers already published in Tables 2-8 and takes minutes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Tables 2-8 are the primary evidence for the central claim, and their headline entries are the Average rows (for example, Table 2 ends with 18.06 and Table 4 with 24.76). These rows are computed as the arithmetic mean of raw scores from metrics with very different ranges: NIMA is roughly 3-5, PaQ2PiQ roughly 60-77, DBCNN roughly 30-60, MUSIQ-Koniq roughly 39-66, MANIQA roughly 0.5-0.7, CLIPIQA roughly 0.1-0.6, and HyperIQA, GPR-BIQA, QualityNet, and PIQI roughly 0.3-0.7. Averaging these raw values without normalization gives dominant weight to the largest-scale metric, so the resulting Average is not a meaningful perceptual summary. This is not a pedantic point: the per-metric scores are mixed. On DarkBDD (Table 2), LucentVisionNet is not best on NIMA (3.99 versus 4.13 for Multiscale Retinex), PaQ2PiQ (66.73 versus 68.55 for LIME), MUSIQ-Koniq (42.30 versus 45.92 for BIMEF), or MANIQA (0.55 versus 0.59 for BIMEF), yet it tops the Average row. On DICM (Table 4), it loses NIMA, MANIQA, TReS-Koniq, and HyperIQA but again tops the Average. The conclusion's claim of consistency across PSNR, SSIM, LPIPS, DISTS, and subjective scores is therefore not supported by the analysis as presented. The additional untested premise, using MUSIQ-AVA aesthetic scores as a training loss while evaluating with the sibling MUSIQ-Koniq model, could also bias results, but the invalid averaging alone is sufficient to undercut the headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LucentVisionNet, a low-light image enhancement framework that combines multi-scale depthwise separable convolutions, spatial attention, and a recurrent curve estimation strategy. Training uses a composite loss with six terms, including a no-reference aesthetic loss based on MUSIQ-AVA. The authors evaluate on paired datasets (LOL, LOL-v2) with full-reference metrics and on unpaired datasets (DarkBDD, DarkCityScape, DICM, LIME, MEF, NPE, VV) with no-reference IQA metrics, and they claim that LucentVisionNet consistently outperforms supervised, unsupervised, and zero-shot baselines across all metrics.","tokens_in":22441,"tokens_out":5782,"duration_ms":49077,"significance":"If the empirical claims were sound, the paper would offer a practical contribution to low-light enhancement: the architecture is lightweight (depthwise separable convolutions, multi-scale fusion, residual learning) and the integration of a differentiable aesthetic quality loss is a plausible way to improve perceptual quality without paired data. The full-reference results on LOL and LOL-v2 show some encouraging improvements, notably in LPIPS and DISTS. However, the central claim of consistent superiority rests on statistically invalid aggregation of no-reference metrics and on differences that are not tested for significance. The per-metric results are mixed, so the headline claim is not supported by the evidence as presented.","major_comments":[{"comment":"The 'Average' rows in Tables 2-8 are computed as arithmetic means of raw scores from metrics with incompatible ranges: NIMA is roughly 3-5, PaQ2PiQ roughly 60-77, DBCNN roughly 30-60, MUSIQ-Koniq roughly 39-66, MANIQA roughly 0.5-0.7, CLIPIQA roughly 0.1-0.6, and HyperIQA/GPR-BIQA/QualityNet/PIQI roughly 0.3-0.7. This averaging gives dominant weight to the largest-scale metrics, so the resulting 'Average' is not a meaningful summary of perceptual quality. For example, on DarkBDD (Table 2) the proposed method tops the Average (18.06) despite being below the best method on NIMA (3.99 vs 4.13), PaQ2PiQ (66.73 vs 68.55), MUSIQ-Koniq (42.30 vs 45.92), MANIQA (0.55 vs 0.59), CLIPIQA (0.13 vs 0.14), and HyperIQA (0.30 vs 0.30). Similar patterns appear in Tables 3-8. The Abstract and Conclusion (Section 8) claim consistent outperformance 'across multiple full-reference and no-reference image quality metrics,' but the per-metric tables show mixed results. This is a load-bearing flaw in the central performance claim.","section":"Tables 2-8, Section 6.1.1"},{"comment":"No error bars, confidence intervals, or statistical significance tests are reported for any of the quantitative comparisons. The margins in the average scores are small in several cases (e.g., Table 2: 18.06 vs 18.02; Table 4: 24.76 vs 24.43; Table 5: 24.18 vs 24.03), and the full-reference improvements are also modest (e.g., Table 9: PSNR 18.39 vs 18.33; SSIM tied at 0.85). Without an estimate of variability across the test images, the reader cannot determine whether the reported differences are meaningful. The 'consistently outperforms' claim requires at least per-image distributions or a suitable significance test.","section":"Tables 2-10, Section 6"},{"comment":"The training loss explicitly maximizes the MUSIQ-AVA aesthetic score (Eq. 23), and the evaluation includes MUSIQ-Koniq, a separate fine-tuned variant of the same MUSIQ architecture. The paper does not discuss this potential circularity: the model is optimized for a MUSIQ-family score and then evaluated with a sibling model. This weakens the independence of the MUSIQ-Koniq results as evidence of perceptual superiority. The authors should either exclude MUSIQ-family metrics from the evaluation, report results with and without the MUSIQ-based loss, or provide an argument that the loss transfers without bias to other metrics.","section":"Section 4.6 and Tables 2-8"}],"minor_comments":[{"comment":"The method is trained on 2,422 images from the SICE dataset, so calling it 'zero-shot learning' is misleading in the strict sense. The term 'zero-reference' (as in Zero-DCE) or 'unsupervised' would be more accurate and would avoid confusion with the established zero-shot learning literature.","section":"Section 5.1"},{"comment":"The caption of Figure 2 states that scores are 'scaled to 100,' but the numbers in Tables 2-8 are raw scores, and the 'Average' rows are clearly raw averages (e.g., the DarkBDD average 18.06 matches the sum of raw values divided by 11). This inconsistency should be resolved.","section":"Figure 2 caption and Table headers"},{"comment":"There are typographical inconsistencies: Figure 5 and 6 captions list '(h) Semantic Guide ZERO DCE51 and (h) Ours' with the same label (h) for two entries, and the table headers in Tables 2-8 use 'A verage' with a space in some rows (e.g., Table 2, 3, 7).","section":"Table 8 and Figure 5/6 captions"},{"comment":"The claim of 'real-time capability' is questionable: processing a 1200x900 image in 1-1.5 seconds is not real-time for video. If the intended meaning is 'interactive,' this should be stated more precisely.","section":"Introduction, paragraph on real-time"}],"recommendation":"reject","confidential_remarks":"The central flaw is the invalid aggregation of no-reference metrics in Tables 2-8, which directly undermines the paper's headline claim. The mixed per-metric results would likely not support the 'consistently outperforms' conclusion even after proper normalization, and the absence of significance testing compounds the problem. I would not recommend major revision; the evaluation protocol and the corresponding claims would need to be substantially reworked, and it is unclear whether the method would still come out ahead. If the authors resubmit, they should provide code and per-image results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the architecture is a sensible recombination, but the 'consistently outperforms' claim rests on an invalid averaging procedure. The Average rows in Tables 2-8 are arithmetic means of raw metric values with wildly different scales (NIMA ~4, PaQ2PiQ ~70, MANIQA ~0.6). That gives the largest-scale metric all the weight, so the average is not a meaningful perceptual summary. The per-metric results are actually mixed: on DarkBDD you lose NIMA, PaQ2PiQ, MUSIQ-Koniq, and MANIQA to simpler methods; on DICM the same pattern holds. The abstract and conclusion claim consistent state-of-the-art performance that the tables do not support.\n\nWhat is genuinely new: the multi-scale spatial attention over Zero-DCE-style curve estimation, depthwise separable convolutions, and the MUSIQ-AVA aesthetic loss. This is a legitimate extension, not a copy. The evaluation is broad: nine no-reference and seven full-reference metrics across ten datasets. The full-reference numbers on LOL and LOL-v2 are plausible and occasionally best (e.g., PSNR 21.25 on LOL-v2), though gains are modest and there are no error bars or significance tests.\n\nThe other soft spots are minor by comparison. No code is released; for an empirical ZSL paper that hurts reproducibility. Using your own GPR-BIQA, PIQI, and QualityNet metrics as evaluation, and optimizing a MUSIQ-AVA loss while evaluating with sibling MUSIQ-Koniq, does not make the evaluation invalid but it weakens independence. A proper analysis with normalized scores, confidence intervals, and code would make the paper publishable.\n\nI would not accept as is, but I would send it to peer review: the method deserves a careful look, and the evaluation is fixable. The right referee will catch the averaging and ask for proper statistics.","headline":"A competent architectural mashup whose headline result is an artifact of averaging raw metric scores on incompatible scales; the per-metric evidence is mixed.","tokens_in":22977,"tokens_out":3732,"would_cite":false,"duration_ms":33156,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Zero-shot low-light enhancement can beat paired-supervision methods on every benchmark tested.","keywords":["low-light image enhancement","zero-shot learning","multi-scale spatial attention","deep curve estimation","no-reference image quality assessment","depthwise separable convolution","perceptual loss","image enhancement benchmarks"],"falsifier":"Recompute the Average rows in Tables 2 through 8 after standardizing each no-reference metric, for example by converting each method's score to a rank per metric before averaging; if LucentVisionNet does not rank first on a majority of the unpaired datasets, the headline claim is an artifact of raw-scale averaging. As a second check, evaluate all methods with a BIQA model outside the MUSIQ family that was not used in training; if the ordering changes substantially, the claimed perceptual advantage does not transfer.","tokens_in":21798,"feed_emoji":"🌙","tokens_out":7763,"duration_ms":71434,"temperature":0.7,"pith_summary":"This paper claims that low-light image enhancement can be done without any paired training data and still outperform fully supervised and unsupervised rivals. The proposed LucentVisionNet combines depthwise separable convolutions at three image scales, a spatial attention block, and a recurrent quadratic curve-estimation step, trained with six losses, including a new no-reference aesthetic-quality loss from MUSIQ-AVA. Across seven unpaired real-world datasets and two paired benchmarks, LOL and LOL-v2, the authors report the best or tied-best scores on most full-reference and no-reference quality metrics, plus a runtime of about 1 to 1.5 seconds for a 1200 by 900 image on one GPU. If true, this would make high-quality enhancement practical for mobile photography, surveillance, and autonomous driving without collecting reference images.","feed_headline":"Zero-shot enhancer outranks paired baselines on every test set","feed_subtitle":"Multi-scale attention and a perception-guided loss beat Zero-DCE, EnlightenGAN, and supervised baselines without paired data.","key_machinery":"The load-bearing object is the multi-scale spatial curve estimation network: the input image is processed at full, half, and quarter resolution by parallel stacks of depthwise separable convolutions, fused hierarchically, passed through a spatial attention block, and mapped by a final depthwise separable convolution layer with tanh activation into an enhancement curve. Enhancement is applied recurrently through the residual quadratic update $X_t = X_{t-1} + D(X_{t-1}^2 - X_{t-1})$, where $D$ is a diagonal matrix of per-pixel curve parameters predicted by the network. This structure lets the model refine exposure over iterations while staying cheap enough for deployment. The other load-bearing piece is the six-term composite loss, including total variation, spatial consistency, color constancy, exposure control, segmentation guidance, and the MUSIQ-AVA no-reference aesthetic loss, which supplies the perceptual and semantic pressure that replaces ground-truth supervision.","core_discovery":"The paper's central claim is that LucentVisionNet consistently outperforms state-of-the-art supervised, unsupervised, and zero-shot low-light enhancement methods on both paired and unpaired benchmarks. On the paired LOL and LOL-v2 datasets it reports the highest PSNR among all compared methods, tied-highest SSIM and VSI, the lowest or tied-lowest LPIPS and DISTS on LOL, and the lowest MAD; on seven unpaired datasets it reports the highest average no-reference score in every table. The mechanism credited for this is the integration of multi-scale spatial attention into a deep curve estimation network, a recurrent residual enhancement step, and a composite loss whose sixth term is a no-reference aesthetic score from MUSIQ-AVA that rewards perceptually pleasing outputs during training. The authors present this as evidence that zero-shot enhancement can exceed paired methods while staying computationally light enough for near-real-time deployment.","pith_inferences":["The reported average no-reference scores in Tables 2 through 8 are arithmetic means of raw values on different scales, such as NIMA around 4, PaQ2PiQ around 60, DBCNN around 30, and CLIPIQA around 0.1; redoing the comparison with per-metric ranks or z-score normalization could change the winner, so the headline highest-average claim should be read with that caveat.","Because the same MUSIQ architecture family appears both as the training loss, MUSIQ-AVA, and as an evaluation metric, MUSIQ-Koniq, the evaluation may be biased in the model's favor; a held-out BIQA model outside that family would test how much of the gain is real perceptual improvement.","A natural next experiment is to keep the architecture fixed and ablate the MUSIQ-AVA loss term against the other five losses, since the contribution of the paper's one novel loss has not been isolated in the experiments."],"forward_implications":["Zero-shot enhancement can match or beat paired-supervision methods on public paired benchmarks, so collecting aligned low-light and normal-light pairs may not be necessary for strong PSNR or perceptual results.","Recurrent application of a learned quadratic curve, rather than a single pass, is a viable way to trade a little latency for better exposure and structure preservation.","Using a differentiable no-reference aesthetic model as a training loss can steer enhancement toward human-preferred outputs without any reference image, and the same principle could transfer to other image restoration tasks.","At roughly 1 to 1.5 seconds for 1200 by 900 images on one GPU, the architecture is near real-time enough for mobile and edge deployment, with the depthwise separable backbone being the main reason for the low cost."],"supporting_citations":[{"why":"Supplies the base quadratic curve-estimation model and the first four loss terms that LucentVisionNet builds on.","marker":"[49]"},{"why":"Zero-DCE++, the recurrent and fast variant whose improvement path the paper extends.","marker":"[50]"},{"why":"Semantic-guided zero-shot learning; contributes the segmentation guidance loss and the DarkBDD and DarkCityScape test sets, and is a stated baseline to beat.","marker":"[51]"},{"why":"MUSIQ; provides both the no-reference aesthetic model used as the novel training loss, MUSIQ-AVA, and one of the evaluation metrics, MUSIQ-Koniq.","marker":"[52]"},{"why":"EnlightenGAN; supplies the unsupervised baseline and the train and validation split protocol from SICE used in the experiments.","marker":"[53]"},{"why":"AVA dataset; the aesthetic annotations that MUSIQ-AVA is trained on, making the sixth loss term possible.","marker":"[64]"},{"why":"SICE dataset; source of the 360 multi-exposure sequences and 3,022 training images used to fit the model.","marker":"[65]"},{"why":"LOL dataset; one of the two paired benchmarks supporting the full-reference PSNR, SSIM, and LPIPS claims.","marker":"[70]"},{"why":"LOL-v2 dataset; the second paired benchmark supporting the full-reference claims.","marker":"[71]"}],"fun_headline_variants":["Zero-shot low-light enhancer surpasses supervised rivals on all tests","LucentVisionNet wins zero-shot low-light enhancement against paired models","No-pair training powers low-light enhancement that beats supervised baselines","Multi-scale attention and human-perception loss drive zero-shot low-light win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The superiority claims rest on averaging raw scores from no-reference metrics that have different ranges, roughly NIMA around 4, PaQ2PiQ around 60, DBCNN around 30, and CLIPIQA around 0.1, and treating the resulting average as a meaningful summary, plus the unstated assumption that optimizing the MUSIQ-AVA aesthetic score does not bias the MUSIQ-family evaluation metrics.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot low-light enhancer surpasses supervised rivals on all tests","LucentVisionNet wins zero-shot low-light enhancement against paired models","No-pair training powers low-light enhancement that beats supervised baselines","Multi-scale attention and human-perception loss drive zero-shot low-light win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2333,"prompt_tokens":894,"completion_tokens":1439,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":1363}},"tokens_in":510,"tokens_out":1439,"duration_ms":11080,"temperature":1.0,"reasoning_tokens":1363,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:51:29.296648+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Average rows in Tables 2 through 8 after standardizing each no-reference metric, for example by converting each method's score to a rank per metric before averaging; if LucentVisionNet does not rank first on a majority of the unpaired datasets, the headline claim is an artifact of raw-scale averaging. As a second check, evaluate all methods with a BIQA model outside the MUSIQ family that was not used in training; if the ordering changes substantially, the claimed perceptual advantage does not transfer.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the base quadratic curve-estimation model and the first four loss terms that LucentVisionNet builds on."},{"cited_title":"& Gupta, G","cited_arxiv_id":null,"evidence_quote":"Semantic-guided zero-shot learning; contributes the segmentation guidance loss and the DarkBDD and DarkCityScape test sets, and is a stated baseline to beat."},{"cited_title":"& Yang, F","cited_arxiv_id":null,"evidence_quote":"MUSIQ; provides both the no-reference aesthetic model used as the novel training loss, MUSIQ-AVA, and one of the evaluation metrics, MUSIQ-Koniq."},{"cited_title":"Enlightengan: Deep light enhancement without paired supervision","cited_arxiv_id":null,"evidence_quote":"EnlightenGAN; supplies the unsupervised baseline and the train and validation split protocol from SICE used in the experiments."},{"cited_title":"& Perronnin, F","cited_arxiv_id":null,"evidence_quote":"AVA dataset; the aesthetic annotations that MUSIQ-AVA is trained on, making the sixth loss term possible."},{"cited_title":"& Zhang, L","cited_arxiv_id":null,"evidence_quote":"SICE dataset; source of the 360 multi-exposure sequences and 3,022 training images used to fit the model."},{"cited_title":"& Liu, J","cited_arxiv_id":null,"evidence_quote":"LOL dataset; one of the two paired benchmarks supporting the full-reference PSNR, SSIM, and LPIPS claims."},{"cited_title":"& Liu, J","cited_arxiv_id":null,"evidence_quote":"LOL-v2 dataset; the second paired benchmark supporting the full-reference claims."}],"review_version":1}