{"id":"87dd54b8-6384-4dcb-94ff-0ef3005c120d","arxiv_id":"2501.00669","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The work applies existing CNN architectures to leaf disease classification and claims high accuracy, but offers no fundamentally new scientific result.","lead":"This thesis applies deep learning models, including a proposed multi-scale CNN, to classify leaf diseases from images. It reports high accuracy on bean, seed, and tomato datasets, but the methods are incremental variations of known architectures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported DMCNN superiority in §8.4.5.3 depends on the test split never informing model selection; the thesis never states this, and its parameter-tuning protocol (§8.1.2, Table 8.15) creates a real leakage risk.","rationale":"The reader's weakest_assumption correctly identifies the test-set-selection risk. My read of the thesis text reinforces it: the abstract promises a tuning algorithm for 'optimal performance of each model', and the experimental sections document sweeping learning-rate, batch-size, and epoch choices without stating that the final test split was never inspected. Since the strongest claim is exclusively an empirical performance comparison, this is the single most load-bearing point. I do not see a stronger technical flaw: the proposed multi-scale architecture is described at the level of a block diagram (Figs. 8.33–8.36), and a reproducibility concern by itself would not justify rejection if the protocol were clarified. The suggested re-run is feasible for the authors, who still possess the code and data, and it would settle whether 'outperforms' is real or an artifact of selection. Thus no change to the conditional verdict is needed.","tokens_in":42259,"tokens_out":3471,"duration_ms":37566,"concrete_test":"Obtain the code/data or an exact split specification and rerun the §8.4 experiment under a nested protocol: tune DMCNN and every pre-trained baseline on the training portion with inner cross-validation, freeze the selected hyperparameters, and then evaluate once on the untouched test split. Recompute Table 8.18 (and the per-class metrics in Table 8.17) from this single held-out evaluation; if the DMCNN advantage over the best baseline disappears or falls within run-to-run noise, the headline claim does not survive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical one: DMCNN 'outperforms other models regarding accuracy, F1 score, precision, and recall' (Ch. 1; §8.4.5.3, Table 8.18). The load-bearing condition is that the reported test metrics were computed on images never used for model or hyperparameter selection. The thesis says a 'parameter-tuning algorithm was developed to identify the optimal performance of each model' (Abstract) and reports tuning of learning rate, batch size, and epochs (§8.1.2.2, §8.1.2.3, §8.4.5.2); Table 8.15 explicitly documents tuning details for each pre-trained model. Nowhere does the thesis state that the final test split was fixed before tuning or that the chosen configurations were selected using only a validation fold. If the same split was consulted to pick the best configuration, the accuracy gap between DMCNN and its baselines is optimistically biased and the 'outperforms' conclusion is unsupported. The absence of released code, data, and explicit split metadata prevents an external check of this condition; this is a validity risk, not a novelty objection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript, based on a 2023 PhD thesis, applies deep learning to plant leaf disease detection. It contains a broad literature review, experiments with MobileNet on bean leaf datasets, a study of dataset impact, a CNN model for Brassica seeds, and a proposed Deep Multi-Scale CNN (DMCNN) for tomato leaf disease classification. The central empirical claim is that DMCNN outperforms several pre-trained models in accuracy, F1 score, precision, and recall. Evaluation uses public datasets plus a newly collected dataset, and the thesis concludes that the proposed approach yields very satisfactory performance.","tokens_in":42510,"tokens_out":4775,"duration_ms":47713,"significance":"If the evaluation protocol is sound, the claimed DMCNN architecture and its systematic comparison against pre-trained models would be a useful empirical contribution to agricultural computer vision. The paper's strengths include a wide literature review, a detailed account of hyperparameter tuning, the use of GradCAM for qualitative analysis, and the curation of a new leaf disease dataset. However, the significance is currently limited by the absence of statistical uncertainty quantification, the lack of explicit test-set hygiene documentation, and the unavailability of code and data for external verification. The central performance claim is only as strong as the evaluation protocol, and that protocol is not fully described.","major_comments":[{"comment":"The central claim that DMCNN outperforms other models depends on the test set never being used for model or hyperparameter selection. The abstract states that a parameter-tuning algorithm was developed, and Sections 8.1.2.2 and 8.4.5.2 describe tuning learning rate, batch size, and epochs; Table 8.15 documents tuning details for each pre-trained model on the Tomato Leaf dataset. The manuscript never states whether the test split was fixed before tuning or whether model selection used only a validation fold. If the test split informed the final configuration, the reported accuracy advantage is optimistically biased. Please provide the exact split-generation procedure and explicitly confirm that the test set was not used for selection; if nested cross-validation or a validation holdout was used, describe it precisely.","section":"Abstract; §8.1.2; §8.4.5.2; Table 8.15"},{"comment":"The comparative tables report single point estimates of accuracy, F1 score, precision, and recall for each architecture. No error bars, confidence intervals, or significance tests are provided, and no repeated runs with different random seeds are reported. Because the reported differences between DMCNN and the baselines could lie within run-to-run variation, the claim of consistent superiority is not statistically supported. Please report means and standard deviations over multiple runs and apply a paired statistical test, such as McNemar's test on the test predictions, to justify the superiority claim.","section":"§8.4.5.3; Tables 8.18–8.20"},{"comment":"The manuscript does not include a data-availability statement, code, or the exact per-class partition sizes used for the tomato dataset. It also does not report the random seed or the order of augmentation operations. Consequently, the empirical results cannot be reproduced or independently checked, which is especially important for a paper whose contribution is an empirical comparison. Please provide split metadata and per-class counts, and release code or pre-trained weights, or explain clearly why this is not possible.","section":"§8.4.1.1; §8.4.1.2; §8.4.2"}],"minor_comments":[{"comment":"The manuscript contains numerous typographical errors and inconsistent capitalization, such as 'devlop', 'approachs', 'Accuaracy', 'thsis', and 'Mobilenet'. A careful copyedit is needed.","section":"Throughout"},{"comment":"The caption for Figure 8.26 appears duplicated ('Accuracy of loss, validation, and training of the suggested model' is repeated). Please correct the duplicate caption.","section":"List of Figures; Figure 8.26"},{"comment":"The Softmax formula in Table 5.1 is garbled and should be typeset correctly. Reference names are also inconsistent, for example 'Lacun et al., 1998' in Section 5.2.5 versus 'LeCun et al., 1998' elsewhere, and 'Mukti et al., 2013' in the text.","section":"Table 5.1; §5.2.5"},{"comment":"The Brassica seeds experiment in Section 8.3 is presented as part of the thesis but is never explicitly connected to the tomato-leaf DMCNN claim in Section 8.4. The relationship between these studies should be clarified, or the seed classification study should be framed as a separate contribution.","section":"§8.3; §8.4"}],"recommendation":"major_revision","confidential_remarks":"This is a 2023 PhD thesis rather than a typical research article, and the thesis format makes the contribution hard to extract. For an empirical machine learning submission, the absence of any data or code availability statement is a serious concern. The central DMCNN claim is falsifiable in principle, but it cannot currently be verified because the evaluation protocol is underspecified and no reproducibility materials are provided. The journal should require a clear statement of test-set hygiene, statistical uncertainty quantification, and a real data/code availability mechanism before further consideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading this: it is a PhD thesis, not a paper, and its central claim—that the proposed DMCNN outperforms pre-trained models on tomato leaf disease classification—cannot be checked from the text alone. The thesis never states that the test split was fixed before the parameter-tuning algorithm selected each model's configuration, and the tuning protocol shown in Table 8.15 (learning rates, batch sizes, epochs, per model) creates a real leakage risk. That is the load-bearing weakness, not the lack of architectural novelty.\n\nWhat is actually new: the authors collected and curated bean and Brassica seed datasets that are not in the prior literature, and they report a wide range of experiments—optimizer, learning rate, batch size, epoch effects, plus GradCAM visualizations. The DMCNN itself is a multi-branch multi-scale CNN, which is a known design pattern, but the application to this problem and the thorough comparison against pre-trained models is a legitimate applied effort. The literature review is broad and useful.\n\nWhat is soft, in proportion: first, reproducibility. No code, no data, no error bars, no significance tests. Second, the writing is rough—typos, inconsistent citation details, and a few table/text mismatches (e.g., Brahimi accuracy in the text vs. Table 2.1). Third, the leakage risk above. These are not fatal to the entire thesis; the MobileNet and dataset-comparison studies can stand independently. But the DMCNN superiority claim in Section 8.4.5.3 is exactly where the absence of split metadata hurts most.\n\nWho is this for? A reader interested in a broad survey of applied CNN tuning on leaf images, or a PhD student looking for a checklist of experiments. It is not a methods paper.\n\nMy recommendation: do not send this to peer review as-is. The author should be told to release code and data, describe the exact split and selection protocol, and rerun the comparison with a proper validation-based tuning setup. With that, a condensed paper on the datasets or the DMCNN comparison could be refereeable. Without that, the central claim is just an unsupported number.","headline":"A broad, hard-working PhD thesis with a real dataset contribution, but the headline DMCNN result is unverifiable without code/data and a clear statement that test data never informed model selection.","tokens_in":43011,"tokens_out":1821,"would_cite":false,"duration_ms":22027,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The thesis claims that its deep multi-scale convolutional neural network outperforms pre-trained models on ten-class tomato leaf disease classification across accuracy, F1, precision, and recall.","keywords":["deep learning","CNN","tomato leaf disease","multi-scale convolutional neural network","plant disease detection","transfer learning","MobileNet","image classification"],"falsifier":"Retrain the DMCNN and the pre-trained baselines on the same ten-class tomato dataset with the test split sealed until the final evaluation, using identical folds; if the accuracy gap shrinks, reverses, or falls within run-to-run variance, the claim that DMCNN outperforms the baselines is not supported.","tokens_in":42056,"feed_emoji":"🍅","tokens_out":6268,"duration_ms":51365,"temperature":0.7,"pith_summary":"This thesis tries to establish that a purpose-built deep multi-scale convolutional neural network (DMCNN) can classify tomato leaf diseases more accurately than transferred pre-trained models. The central claim is empirical: on a ten-class tomato leaf image dataset, DMCNN is reported to outperform the state-of-the-art architectures it was compared against on accuracy, F1 score, precision, and recall. The argument is carried by a multi-branch design that reads the same leaf image at several scales and merges the branches into one classifier. If true, the result would give farmers an efficient, custom CNN option for early disease detection and would suggest that multi-scale feature fusion is a better fit for leaf-disease images than single-scale transfer learning.","feed_headline":"Multi-scale CNN tops pretrained models on tomato leaf disease","feed_subtitle":"Parallel CNN streams at different scales merge into one classifier that beats pretrained baselines on ten classes.","key_machinery":"The central object is the deep multi-scale CNN (DMCNN): a set of parallel convolutional streams, each operating at a different scale, merged into a single classifier. The multi-branch structure lets the network capture both fine-grained lesion texture and broader leaf context simultaneously, and the fusion layer combines these representations before the final classification. The thesis's claim is that this multi-scale fusion is the mechanism behind the accuracy gain over the pre-trained baselines.","core_discovery":"The paper's proposed DMCNN is a multi-branch convolutional architecture in which parallel streams process the leaf image at different scales and are fused at the end into a single output. On a tomato dataset with ten disease classes, the model is reported to achieve higher accuracy, F1, precision, and recall than several pre-trained state-of-the-art architectures, and the thesis presents this as evidence that multi-scale feature fusion captures the lesion-level and leaf-level cues needed for disease classification better than single-scale pre-trained networks.","pith_inferences":["If the multi-scale design is the cause of the gain, combining it with transfer learning rather than training from scratch could push accuracy further on smaller datasets.","The thesis's MobileNet experiments suggest dataset composition can shift accuracy as much as architecture choice; testing DMCNN on the same public/collected/merged splits would separate architecture effects from data effects.","A sealed test set and repeated cross-validation would show whether the reported margin over pre-trained models survives without any tuning on the test data.","The seed-image classification work indicates the tuning methodology and possibly the multi-scale architecture could transfer to seed quality sorting and other agricultural image tasks."],"forward_implications":["DMCNN classifies ten tomato leaf disease classes with accuracy, F1 score, precision, and recall above the pre-trained models compared in the thesis.","Multi-scale feature fusion is a more effective design choice for leaf-disease classification than single-scale transfer learning.","A custom CNN trained from scratch can compete with, and in this study beat, large pre-trained networks on a ten-class plant disease dataset.","The same architecture and hyperparameter-tuning procedure can be applied to other crops and to real-time disease detection in the field.","Dataset composition affects model effectiveness, as demonstrated by the MobileNet experiments on public, collected, and merged bean leaf datasets, so reported gains should be read relative to the dataset used."],"supporting_citations":[{"why":"Early tomato-disease CNN classification with AlexNet and GoogLeNet that the thesis takes as a state-of-the-art baseline to beat.","marker":"Brahimi et al., 2017"},{"why":"PlantVillage-based CNN classification across 26 diseases, establishing the benchmark dataset tradition for plant disease recognition.","marker":"Mohanty et al., 2016"},{"why":"Large-scale comparison of CNN architectures on 58 plant diseases, providing pre-trained model baselines and accuracy standards referenced throughout.","marker":"Ferentinos et al., 2019"},{"why":"Deep residual network for tomato leaf diseases that defines one of the prior tomato-specific results the DMCNN compares against.","marker":"Karthik et al., 2019"},{"why":"GoogleNet/VGG16 bean disease study that motivates fine-tuning pre-trained networks, the alternative approach DMCNN is contrasted with.","marker":"Sahu et al., 2021"}],"fun_headline_variants":["Multi-scale fusion beats pretrained nets on leaf disease","Parallel streams lift leaf disease accuracy","Custom CNN outperforms pretrained for tomato leaf disease","Multi-branch CNN edges out pretrained models","Scale-aware CNN tops pretrained leaf classifiers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy comparisons assume the test images were never used, directly or indirectly, to choose the model or tune its hyperparameters, so the reported margins are unbiased estimates of generalization.","fun_headline_variants_meta":{"raw":{"variants":["Multi-scale fusion beats pretrained nets on leaf disease","Parallel streams lift leaf disease accuracy","Custom CNN outperforms pretrained for tomato leaf disease","Multi-branch CNN edges out pretrained models","Scale-aware CNN tops pretrained leaf classifiers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":1107,"prompt_tokens":786,"completion_tokens":321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":402,"completion_tokens_details":{"reasoning_tokens":252}},"tokens_in":402,"tokens_out":321,"duration_ms":3387,"temperature":1.0,"reasoning_tokens":252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:44:30.429025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the DMCNN and the pre-trained baselines on the same ten-class tomato dataset with the test split sealed until the final evaluation, using identical folds; if the accuracy gap shrinks, reverses, or falls within run-to-run variance, the claim that DMCNN outperforms the baselines is not supported.","supporting_citations":[{"cited_title":"Using deep learning for image -based plant disease detection","cited_arxiv_id":null,"evidence_quote":"PlantVillage-based CNN classification across 26 diseases, establishing the benchmark dataset tradition for plant disease recognition."}],"review_version":1}