{"id":"10919d46-e718-494d-81f4-36f74400df7e","arxiv_id":"2507.01590","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A classroom monitoring system combining YOLOv8, MTCNN, and LResNet reports high detection accuracies, but the paper lacks reproducible artifacts and contains inconsistent results.","lead":"This paper describes a classroom surveillance system that uses YOLOv8, MTCNN, and LResNet to detect sleeping students, phone use, and faces for attendance. The authors report high accuracy numbers but provide no code, data, or detailed methodology, and the reported metrics contradict each other.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The face-recognition accuracy and phone mAP are reported as two different numbers in the same paper, so the headline performance claims are internally inconsistent and unverifiable without artifacts.","rationale":"The reader's verdict is REJECT, and this stress-test identifies the same core problem: the reported metrics, which are the entire evidence for the system's claimed performance, contradict each other. The face recognition validation accuracy is 86.45% in the abstract and conclusion but 84% in Table I; the mobile phone mAP@50 is 85.89% in the abstract and Table I but 87.65% in the conclusion. These are explicit internal inconsistencies, not subtle methodological choices. The paper supplies no training logs, code, data splits, or checkpoints that could disambiguate the numbers. The proposed concrete test would settle the concern by re-evaluating the model on the stated validation set and comparing against both reported values. If neither value reproduces, the central claim is unsupported. This does not change the reader's verdict; it strengthens the basis for rejection. I did not base the critique on the likely fabricated references or the absence of code, although those further support the same conclusion.","tokens_in":8491,"tokens_out":3941,"duration_ms":45250,"concrete_test":"Obtain from the authors the exact trained CustomLResNet Occ FC checkpoint, the validation/test split indices used for Table II, and the evaluation script, then run inference over the 329-image validation set with the reported preprocessing and top-1 accuracy. If the resulting accuracy is neither 84.00% nor 86.45%, the abstract's face-recognition headline is contradicted. Independently re-run the YOLOv8 mobile phone detector on the same test images to check whether mAP@50 is 85.89% or 87.65%; a mismatch with both numbers would confirm that the reported results are not reproducible from the stated protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on three headline performance numbers, and two of them are not stable within the manuscript. Face recognition is reported as 86.45% validation accuracy in the abstract and conclusion, but Table I lists Validation Accuracy 84% for the same 'CustomLResNet Occ FC' model. Mobile phone detection is reported as 85.89% mAP@50 in the abstract and Table I, but the conclusion states 87.65%. These are not rounding differences; they are different reported values for the same quantity, model, and dataset. The paper provides no code, no model checkpoints, and no evaluation script, so there is no way to determine which number is the actual result or whether either came from a properly held-out test set. Because the paper's contribution is precisely these metrics, an internal contradiction in the headline numbers makes the central claim not merely unverified but ill-defined. This is the load-bearing weakness: if the contradictory metrics are not resolved, the abstract's performance claims cannot be accepted as a coherent statement of what the system achieves.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an integrated classroom surveillance system combining YOLOv8-based sleep and mobile-phone detection, an LResNet Occ FC face-recognition model, and SORT tracking, deployed on ESP32-CAM hardware with a PHP web interface. The authors report three headline performance numbers: 97.42% mAP@50 for sleep detection, 86.45% validation accuracy for face recognition, and 85.89% mAP@50 for mobile-phone detection. The paper includes standard equations for softmax classification, cosine similarity, YOLO loss, and the SORT/Kalman-filter tracking pipeline, along with dataset summaries and a self-declared limitations section. The main claimed contribution is the integration of these components into a real-time monitoring system for educational settings.","tokens_in":8817,"tokens_out":2011,"duration_ms":23772,"significance":"If the reported metrics were reliable and the system were fully described, the integrated monitoring pipeline could be a useful engineering contribution to the smart-classroom literature, particularly for its use of low-cost ESP32-CAM hardware and explicit attention to real-time constraints. The paper also deserves credit for openly listing limitations in Section V-B, including occlusion, lighting, and pose issues. However, the central value of the paper rests entirely on the accuracy of the reported performance numbers and on the reproducibility of the training and evaluation protocol. Those numbers are internally inconsistent, and the protocol is not described in sufficient detail for independent verification. Because the headline claims are the paper's primary contribution, the current manuscript does not establish a sound basis for accepting those claims.","major_comments":[{"comment":"The face-recognition validation accuracy is reported inconsistently: the abstract and conclusion state 86.45%, while Table I lists a Validation Accuracy of 84% for the same 'CustomLResNet Occ FC' model. These are materially different values for the same quantity, and no explanation or corrected number is provided. This internal contradiction is load-bearing because the paper's central claim is precisely these performance figures.","section":"Abstract / Table I / Section V"},{"comment":"The mobile-phone detection mAP@50 is reported as 85.89% in the abstract and Table I, but the conclusion states 87.65%. Again, these are different values for the same metric, model, and dataset. The discrepancy cannot be dismissed as a rounding artifact, and without code, checkpoints, or an evaluation script there is no way to determine which value is correct. The central performance claim for this component is therefore ambiguous.","section":"Abstract / Table I / Section V"},{"comment":"The dataset summary is internally inconsistent and incomplete. For the face-recognition dataset, the stated training/validation/testing counts (3000 + 329 + 186 = 3515) do not sum to the stated total of 3524 images, and the 'Classes' field lists three entries ('Masked face, Normal face, Glasses') while claiming 4 classes. For the sleep-detection dataset, no total image count is given, and the paper does not describe whether the reported metrics were computed on the specified test subsets or on some other held-out split. This undermines confidence in the experimental basis of the results.","section":"Section IV-B, Table II"},{"comment":"Training details essential for evaluating the results are missing. The paper does not specify learning rates, batch sizes, number of training epochs for the YOLO models, optimizer choices, data augmentation, image resolutions, or how the ESP32-CAM frames were preprocessed for the reported inference times. Without these details, the reported precision, recall, and mAP values cannot be reproduced or independently validated, which is particularly problematic given the internal metric inconsistencies noted above.","section":"Sections III and IV"}],"minor_comments":[{"comment":"The paragraph on mobile-phone detection is duplicated verbatim in the Related Work section: the sentences beginning 'For example, (Nguyen et al., 2023) proposed a YOLOv5-based system...' appear twice, and a similar duplication occurs with the 'Similarly, (Li et al., 2022)...' sentence. This should be removed.","section":"Section II"},{"comment":"Several references appear to be fabricated or contain clearly invalid bibliographic data. For example, references [13], [14], [16], [17], [18], [19], [20], [23], [24], [25], [26], [27], [28], [29], [30], [31], [32], [33], [34], and [35] use DOI patterns such as '10.1016/j.jair.2022.12345' and '10.1109/tip.2023.123456' that are placeholders rather than real DOIs. Reference [22] begins with ', 12(4), 345–357' with no author names or title. These entries must be corrected or removed.","section":"References"},{"comment":"The softmax equation 'P (Sleep|x) = ezsleep / P i ezi' is missing a subscript on the denominator sum to indicate that the sum runs over all classes; the notation should be made explicit for clarity.","section":"Section III-A"},{"comment":"The cosine similarity formula 'Similarity(A, B) = A · B / ∥A∥∥B∥' omits the norm notation in the denominator; it should read '\\|A\\| \\|B\\|' with explicit vector norms, and the threshold value (e.g., 0.7) should be stated as a tuned hyperparameter rather than a fixed assumption.","section":"Section III-B"},{"comment":"Table I mixes training metrics (train loss, train accuracy) with validation metrics without labeling the columns clearly, and it omits the evaluation protocol for the 'Fitness' and 'Inference Time' rows. Clarifying which numbers are computed on which split would improve interpretability.","section":"Section IV-A, Table I"}],"recommendation":"reject","confidential_remarks":"The manuscript shows strong signs of incomplete or unreliable scholarship: the two headline metrics are numerically inconsistent across sections, the reference list contains multiple placeholder-like DOIs and at least one malformed entry, and the dataset counts do not add up. The paper also lacks any code or data release, and the experimental setup is described in insufficient detail. These issues go beyond presentation and affect the core claims, so I recommend rejection. The editor may also wish to flag the reference concerns for integrity review, as many entries appear to be fabricated or generated by a language model."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper before deciding what to do with it: the integration is routine, and the reported numbers do not hold together. The authors bolt YOLOv8 for phone/sleep detection, MTCNN + LResNet for face recognition, and SORT tracking onto an ESP32-CAM/PHP web app. That is a sensible engineering sketch, and the architecture and workflow are described clearly enough that someone could replicate the overall setup. The limitations section is also refreshingly candid about occlusions, low-light issues, and ESP32-CAM compute limits. Those are real positives, and I want to give credit for them.\n\nThe problems start with the numbers. Face recognition validation accuracy is 86.45% in the abstract and conclusion, but Table I says 84% for the same model. Phone detection mAP@50 is 85.89% in the abstract and Table I, but the conclusion says 87.65%. These are not rounding differences. The paper's sole contribution is these metrics, and they contradict each other. Without code, checkpoints, or an evaluation script, there is no way to tell which number is real, or whether any of them came from a held-out test set. The test sets are also tiny: 42 images for phone detection, 186 for face recognition. That alone would make the headline percentages fragile.\n\nThe citation pattern is a second, independent red flag. Several DOIs look fake (e.g., 10.1016/j.jair.2022.12345), reference [22] is malformed with a missing author line, references [21] and [35] are duplicates, and the Related Work section contains a duplicated paragraph verbatim. This is not a mere style issue; it makes the literature grounding unreliable and violates basic scholarly norms.\n\nI want to be fair: the equations are standard and correctly stated, and nothing here is circular or fraudulent in the sense of a fabricated experiment. But the central performance claims are ill-defined because the same quantity is reported as two different values in the same manuscript. The paper is a modest integration note, not a research contribution. A reader wanting a rough template for a classroom demo might skim the architecture section, but as a paper it does not meet the bar for serious scrutiny.\n\nMy recommendation: desk reject. Do not spend referee time on it. If the authors return with code/data, consistent metrics, and a cleaned reference list, it could become a weak workshop paper, but not in its current form.","headline":"A straightforward classroom-monitoring integration whose reported metrics are internally inconsistent and whose reference list looks partly fabricated; not ready for peer review.","tokens_in":9206,"tokens_out":1835,"would_cite":false,"duration_ms":22564,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An integrated YOLOv8 and LResNet pipeline on ESP32-CAM hardware can monitor sleep, phone use, and attendance in real time.","keywords":["classroom surveillance","YOLOv8","face recognition","sleep detection","mobile phone detection","ESP32-CAM","SORT tracking","multimodal deep learning"],"falsifier":"Run the same three models on an independent, class-balanced test set of real classroom videos with varied lighting, occlusion, and viewing angles, using a documented train/validation split; if sleep mAP@50 falls far below 97%, phone mAP@50 below 86%, or face recognition accuracy below 84%, the claimed performance does not transfer outside the paper's own datasets.","tokens_in":8316,"feed_emoji":"📷","tokens_out":10141,"duration_ms":94173,"temperature":0.7,"pith_summary":"This paper tries to establish that one low-cost camera system can monitor a classroom automatically: detect students who are asleep, detect students using phones, and recognize faces for attendance. The system combines YOLOv8 models for sleep and phone detection with a custom LResNet Occ FC face-recognition model, all tied together with SORT tracking on ESP32-CAM footage. The authors report sleep detection at 97.42% mAP@50 (mean average precision at the 0.5 overlap threshold), face recognition at 86.45% validation accuracy, and mobile phone detection at 85.89% mAP@50. A sympathetic reader would care because real-time automated monitoring could replace subjective manual observation in classrooms, giving teachers objective, immediate feedback on engagement and attendance.","feed_headline":"Classroom AI watcher detects sleep, phones, and faces in real time","feed_subtitle":"Three deep-learning models on a cheap ESP32 camera give automatic attendance and behavior alerts.","key_machinery":"The load-bearing mechanism is the three-model pipeline: YOLOv8 produces the sleep and mobile-phone detections, a custom LResNet Occ FC network (a residual-network face recognizer for occluded faces) produces identity embeddings matched by cosine similarity, and MTCNN supplies face detection. The SORT algorithm, built on a Kalman filter and Hungarian assignment, keeps detections attached to the same student across video frames. The ESP32-CAM provides the real-time video stream, and a PHP web application records events and attendance. What carries the argument is the combination itself: the authors claim that these components work together on cheap hardware while preserving high per-task accuracy.","core_discovery":"The central claim is that an integrated multimodal surveillance pipeline can assess student attentiveness in real time at a low hardware cost. YOLOv8 detects two disengagement behaviors—sleeping and phone use—while a custom LResNet Occ FC network recognizes faces, with MTCNN finding faces and the SORT algorithm tracking identities across frames. The pipeline runs on ESP32-CAM hardware and is orchestrated by a PHP web application that logs events and marks attendance when a student's face matches the database. On the paper's specialized datasets, the sleep detector reaches 97.42% mAP@50, the phone detector 85.89% mAP@50, and the face recognizer 86.45% validation accuracy (Table I reports 84%). The authors present this as evidence that automatic, affordable classroom monitoring can be built from off-the-shelf components.","pith_inferences":["Editorial inference: the reported metrics were computed on very small datasets (1,048 phone images and 3,524 face images), so real classroom performance in large, varied environments is likely to be lower than the headline numbers.","Editorial inference: the mismatch between 86.45% and 84% for face recognition accuracy, and between 85.89% and 87.65% for phone mAP@50, means the paper's own reporting is not yet internally consistent enough to support a deployment decision.","Editorial inference: continuous camera surveillance of students raises privacy and consent concerns that the paper does not engage with; any field deployment would need a data-protection review before piloting.","Editorial inference: because the pipeline is modular, the same sleeping/phone/face stack could be retrained to monitor attention in video-conference calls by swapping the ESP32-CAM input for screen-capture frames."],"forward_implications":["If the reported numbers hold, a single ESP32-CAM can perform sleep detection, phone detection, and face recognition simultaneously, so classroom monitoring no longer requires expensive GPUs or multiple cameras per room.","Teachers could receive real-time alerts when students fall asleep or use phones, and attendance could be recorded automatically whenever a known face is detected.","Because SORT tracking maintains identities across frames, each student's behavior can be logged per session, enabling trend analysis of engagement over time.","The modular design means adding a new behavior class to YOLOv8 would extend the same pipeline to other distractions, such as eating or talking, without changing the tracking or attendance components."],"supporting_citations":[{"why":"Cited as the source for using YOLOv8 to detect both mobile phone and sleep usage.","marker":"(Ghatge et al., 2024)"},{"why":"Cited for the MTCNN-based face detection/recognition component used in the pipeline.","marker":"(Durai et al., 2024)"},{"why":"Review claiming multimodal machine-learning monitoring gives insight into student engagement, grounding the system's purpose.","marker":"(S et al., 2023)"},{"why":"Cited for the ESP32-CAM data-capture hardware that supplies real-time video frames.","marker":"(Neto et al., 2024)"},{"why":"Cited for automatic classroom attendance via face recognition, the use case the face module serves.","marker":"(Banada, 2025)"},{"why":"Cited for the YOLOv8-based real-time mobile phone detection model.","marker":"(Anderson et al., 2023)"},{"why":"Supplies the Kalman-filter and Hungarian-algorithm tracking machinery used by SORT in the pipeline.","marker":"(Wilson et al., 2022)"},{"why":"Cited for the LResNet-based face-recognition approach used to associate faces with student identities.","marker":"(Robinson et al., 2023)"},{"why":"The YOLOv5 mobile-phone-detection baseline whose performance the YOLOv8 phone detector is positioned against.","marker":"(Nguyen et al., 2023)"}],"fun_headline_variants":["AI classroom camera flags sleep, phone use, and faces","YOLO plus face net: real-time attention tracking on ESP32","Cheap ESP32 runs deep learning to monitor students","Multimodal AI watches for drowsiness and distraction","Sleep, phone, face: one AI pipeline for classes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim is load-bearing on the reported accuracy numbers being computed honestly on properly held-out data, but the paper does not describe the training/validation split for the sleep and phone models and reports face recognition accuracy as both 86.45% and 84%.","fun_headline_variants_meta":{"raw":{"variants":["AI classroom camera flags sleep, phone use, and faces","YOLO plus face net: real-time attention tracking on ESP32","Cheap ESP32 runs deep learning to monitor students","Multimodal AI watches for drowsiness and distraction","Sleep, phone, face: one AI pipeline for classes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1674,"prompt_tokens":960,"completion_tokens":714,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":632}},"tokens_in":576,"tokens_out":714,"duration_ms":8630,"temperature":1.0,"reasoning_tokens":632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:46:46.448817+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three models on an independent, class-balanced test set of real classroom videos with varied lighting, occlusion, and viewing angles, using a documented train/validation split; if sleep mAP@50 falls far below 97%, phone mAP@50 below 86%, or face recognition accuracy below 84%, the claimed performance does not transfer outside the paper's own datasets.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the source for using YOLOv8 to detect both mobile phone and sleep usage."},{"cited_title":"Massively Annotated Datasets for Assessment of Synthetic and Real Data in Face Recognition","cited_arxiv_id":"2404.15234","evidence_quote":"Cited for the ESP32-CAM data-capture hardware that supplies real-time video frames."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited for automatic classroom attendance via face recognition, the use case the face module serves."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Kalman-filter and Hungarian-algorithm tracking machinery used by SORT in the pipeline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited for the LResNet-based face-recognition approach used to associate faces with student identities."}],"review_version":1}