{"id":"7ebe993b-e322-411f-a7cb-1609164fc54b","arxiv_id":"2507.18288","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new expert-annotated dataset of 6,719 tongue images with 20 TCM diagnostic feature categories, benchmarked with nine object detection models.","lead":"The authors release a dataset of 6,719 tongue images labeled by traditional Chinese medicine practitioners for 20 diagnostic features. It is meant to give AI researchers a standardized benchmark for automated tongue diagnosis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No inter-annotator agreement or label-noise analysis is reported; the central 'clinically validated' claim, on which every benchmark number depends, remains unverified.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: annotation reliability. The paper's abstract repeatedly asserts clinical validation, but the methods section gives only qualitative assurances and no reproducibility evidence. This concern is not a mere matter of presentation; it determines whether the dataset's benchmark numbers have any scientific meaning. The conditional verdict is appropriate because the dataset could still be useful if the labels are in fact consistent, but the paper currently does not demonstrate that. No additional concern (e.g., about the 'first' claim or dataset availability) is as central, because even a non-first dataset would be useful if labels were reliable; the reverse is not true. My read does not change the reader's verdict.","tokens_in":9565,"tokens_out":2012,"duration_ms":22613,"concrete_test":"Select a stratified random sample of 200 images from the released dataset, spanning all 20 categories. Have three independent TCM practitioners (5+ years of experience, blinded to the original labels) re-annotate each image. Compute per-category Fleiss' kappa and per-image multi-label Jaccard agreement against the released labels. Then recompute mAP for YOLOv8s on the same test split using consensus-corrected labels as ground truth. If kappa is below 0.6 for global categories or 0.4 for local categories, or if mAP shifts by more than 3 points, the benchmark numbers and 'clinically validated' claim require substantial qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's central value proposition is that its 6,719 images carry clinically validated labels, but the annotation section ('Label Selection and Annotation') provides no quantitative evidence for label reliability. It describes a two-stage process (trained technicians, then TCM practitioners with 5+ years of experience, with consensus for critical cases), yet it reports no inter-annotator agreement statistic, no number of annotators per image, no blinded re-annotation study, and no operational definition of 'clinically validated.' The benchmark results in Table 2 (YOLOv5/v7/v8, SSD, MobileNetV2) are presented as demonstrating the dataset's utility; if label noise is high or categories are inconsistently applied, those mAP numbers are not a property of the tongue images but of a single, unrepeatable annotation process. The paper even notes that detectors confuse red versus purple tongues and miss cracked/dentate tongues; this could reflect model limitations, but it is equally consistent with ambiguous or noisy labels. Because the 'first, standardized, clinically validated' claims and every downstream performance figure hinge on annotation correctness, the absence of reliability evidence is the most load-bearing gap in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents TCM-Tongue, a dataset of 6,719 tongue images with multi-label object-detection annotations for 20 TCM symptom categories, captured with a custom standardized acquisition system and reviewed by TCM practitioners. The authors describe the hardware, the 20-class taxonomy, dataset statistics (an average of 2.54 labels per image), train/validation/test splits, and benchmark results for nine detectors (YOLOv5/7/8 variants, SSD, MobileNetV2), reporting precision, recall, mAP@0.5, and mAP@0.5:0.95. The dataset, multi-format annotations, and code are publicly available. The central claims are that this is the first publicly available, standardized, clinically validated TCM tongue dataset suitable for AI development.","tokens_in":9921,"tokens_out":5586,"duration_ms":57642,"significance":"The resource addresses a real gap: publicly available, AI-ready TCM tongue datasets with standardized annotations are scarce, and the paper provides multi-format annotations, a structured taxonomy, example images, and baseline code. If the label-reliability and acquisition-standardization claims hold, this dataset would be a useful benchmark for the TCM-AI community and for studying fine-grained medical object detection. The paper is also informative in reporting the limitations of current detectors, such as red/purple tongue confusion and misses of cracked and dentate tongues. The primary value is the dataset itself rather than the detector comparison; the moderate mAP values are honestly reported and do not undermine the dataset contribution, provided the reliability evidence is added.","major_comments":[{"comment":"The 'clinically validated' ground truth, on which both the dataset's value and every benchmark number depend, is not empirically supported. The description of two-stage review by technicians and TCM practitioners is qualitative; no inter-annotator agreement statistic (e.g., Cohen's or Fleiss' kappa), number of annotators per image, blinded re-annotation study, or operational definition of 'clinically validated' is reported. Please add per-class annotation counts, an adjudication protocol, annotator qualification details, and a label-noise analysis on a re-annotated subset. Without this, the mAP scores in Table 2 cannot be interpreted as properties of the images rather than of one unrepeatable annotation process.","section":"Label Selection and Annotation"},{"comment":"The claimed standardization of image acquisition is asserted rather than validated. The text states D65 illumination, 500-1500 lux adjustment, sub-100 micrometer resolution, and 3-8 second cycles, but no calibration certificates, color-correction measurements, focus-target tests, inter-device agreement, or repeatability data are provided. Please add a validation protocol with quantitative results such as color-checker Delta-E values, resolution-chart measurements, and within-session and between-session repeatability, and clarify how the logged metadata allows per-image auditing of these parameters. The stated real-time ResNet-50 demographic profiling of subjects also raises privacy and bias concerns; please explain what data are stored, for how long, and how bias and consent are handled.","section":"Methodological Implementation of the Tongue Diagnosis Capture System"},{"comment":"The benchmark comparison is single-run and has no error bars or statistical testing. Differences such as mAP@0.5 values of 34.57 (YOLOv5l), 34.82 (YOLOv7), and 34.95 (YOLOv8l) are likely within seed and initialization variance, so the text's conclusions about which models are 'best' are not justified. Please report means and standard deviations over at least 3-5 seeds, per-class AP, and a significance analysis or critical-difference test. In addition, Equations (1)-(5) are not displayed in the manuscript, so the precise definitions of P, R, mAP0.5, and mAP0.5:0.95 are not actually provided.","section":"Technical Validation (Table 2; Eqs. (1)-(5))"},{"comment":"Several details necessary for assessing a medical dataset are missing: institutional review board approval or an ethics statement, the consent process beyond a one-line mention, participant inclusion and exclusion criteria, demographic and diagnostic distribution, collection sites and date range, and image resolution and format. The paper also contains an internal inconsistency: Fig. 1 states an 80/10/10 split, while Section 4 reports 82.3/8.4/8.1; please reconcile or correct this. These details are needed for the 'clinically validated' and 'standardized' claims to be auditable.","section":"Data Records and Fig. 1"},{"comment":"The 'first specialized dataset' claim is not supported by a baseline survey. The related-datasets paragraph lists datasets in other imaging domains but does not systematically compare existing TCM tongue datasets or publicly available tongue-image collections in terms of size, annotation type, illumination control, and availability. Please add a comparison table and restrict the 'first' claim to what can actually be verified from the cited prior work.","section":"Background & Summary (Related Image Datasets)"}],"minor_comments":[{"comment":"The notation for the mAP variants is inconsistent (mAP0.5-0.95, mAP0.50.95, mAP@0.5); please standardize to mAP@0.5 and mAP@0.5:0.95.","section":"Technical Validation"},{"comment":"Figure 6 shows per-label counts, but the text only states the average number of labels per image; please provide the per-image label-count distribution and a per-class instance count table with the number of images containing each label.","section":"Fig. 6 and Data Records"},{"comment":"Reference [11] contains spacing errors in the title ('Tongue Di agnosi s i n Chi nese Medicine'), and the Usage Notes refers to 'SD[29]' where 'SSD' is meant.","section":"References and Usage Notes"},{"comment":"The exact image resolution, bit depth, color space, and storage format should be stated; 'high-resolution source images' is not a technical specification.","section":"Data Records"},{"comment":"Figure 1's caption lists three panel functions, but the subfigures are not individually labeled; please add panel labels and describe what each panel shows.","section":"Fig. 1"},{"comment":"The text refers to 'the first 100 rounds of training' in Figure 8; please specify whether these are epochs and report the training hyperparameters such as learning rate, batch size, and total number of epochs.","section":"Technical Validation"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is potentially valuable and the public release is a clear strength. The central risk is not the benchmark experiments but the absence of quantitative evidence for annotation reliability and acquisition standardization; a small re-annotation reliability study would address the most serious concern. Also, the missing displayed equations in the technical validation section should be fixed in the revision, as should the split inconsistency between Figure 1 and the text. This is a fixable paper rather than one requiring rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core of this paper is a new open dataset: 6,719 tongue images with 20 TCM symptom labels, multi-format annotations (COCO/XML/TXT), and a nine-model detection benchmark. That is a real contribution to a niche but active area. The label taxonomy is grounded in TCM classics and clinical practice, and the global/local box distinction is a sensible way to represent both whole-tongue states and regional features. The benchmark is routine but adequate as a sanity check that the dataset is trainable.\n\nThe main soft spot is exactly where the stress test points: label reliability. The paper claims the labels are 'clinically validated' by licensed TCM practitioners, but reports no inter-annotator agreement, no number of annotators per image, no re-annotation study, and no operational definition of 'clinically validated.' Every downstream number—the 2.54 labels per image and all mAP values—depends on that assumption. If label noise is high, the benchmark numbers reflect the annotation process as much as the tongue images. The authors even note that detectors confuse red and purple tongues and miss cracked or dentate tongues; that is consistent with ambiguous labels, not just model limits. This is a load-bearing gap, not a minor omission.\n\nSecond, the 'first specialized dataset' claim is unverified. The related-work section compares against unrelated image datasets (UNISA2020, WORD, CubiCasa5K) and never discusses prior TCM tongue datasets. There are existing tongue image corpora in the literature; a comparison—even to show they are smaller or not AI-ready—would substantiate the novelty claim.\n\nThird, the benchmark lacks error bars or repeated runs, so the ranking across YOLOv5/v7/v8 variants could be within noise. The acquisition hardware is described in detail, but the description itself is not validated: no color calibration check, no test against a reference device, no inter-device consistency data.\n\nMinor issues: Figure 1 says 'data collection by robot' while the text describes a static dual-camera device; the writing is repetitive in places. The dataset itself is open and versioned (v3.0), which is good and should be credited.\n\nWho should read this: researchers building TCM tongue diagnosis models and anyone benchmarking object detection on small, specialized medical datasets. The paper deserves a serious referee, but the peer review should require evidence of label reliability and a proper comparison with prior tongue datasets. If the authors add inter-annotator agreement and clarify the validation protocol, this becomes a solid data descriptor. As it stands, treat the clinical-validation claim as unsubstantiated and the benchmark as indicative only.","headline":"Useful new tongue dataset with sensible labels, but the 'clinically validated' claim lacks the reliability evidence the benchmark numbers depend on.","tokens_in":10268,"tokens_out":3007,"would_cite":false,"duration_ms":28960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a dataset of 6,719 standardized tongue images with 20 expert symptom labels, and benchmarks nine object-detection models on it, aiming to give AI-assisted TCM diagnosis a common, publicly available testbed.","keywords":["tongue diagnosis dataset","traditional Chinese medicine","object detection benchmark","multi-label annotation","standardized imaging protocol","deep learning medical imaging","symptom classification"],"falsifier":"Select 300 images at random from the released set, have two independent licensed TCM practitioners label each image with the same 20 categories, and compute per-category Cohen's kappa. If several categories fall below roughly 0.6, the labels are too inconsistent to count as ground truth, and the reported detector rankings lose their anchor.","tokens_in":9402,"feed_emoji":"👅","tokens_out":7098,"duration_ms":73878,"temperature":0.7,"pith_summary":"This paper introduces TCM-Tongue, a dataset of 6,719 tongue photographs taken under a standardized capture protocol and labeled by TCM practitioners with 20 symptom categories, averaging 2.54 labels per image. The authors position it as the first publicly available, AI-ready dataset for TCM tongue diagnosis, filling a gap that has blocked deep learning work on tongue examination. The dataset ships in COCO, XML, and TXT annotation formats, and the paper benchmarks nine object-detection models to show what current architectures achieve on it. If the dataset and its labels hold up, it gives the field a common testbed for automating a diagnostic tradition that has until now been subjective and hard to reproduce.","feed_headline":"6,719 tongue images become a benchmark for AI-based TCM diagnosis","feed_subtitle":"Twenty pathological symptom categories and standardized capture give computer vision a shared tongue-diagnosis testbed.","key_machinery":"The central object is the annotation framework: a dual-level label system in which global labels describe the whole tongue (for example red tongue, white coating, thin tongue) and local labels describe subregions (cracks, teeth marks, depressed or protruding organ areas), all encoded as bounding boxes with class indices 0–19. The other load-bearing component is the standardized acquisition protocol: a purpose-built dual-camera capture system with calibrated D65 lighting, distance-controlled wide-angle and telephoto imaging, and automatic quality checks, designed so that all images share similar geometry and color. Together these convert an observational, subjective clinical skill into a quantifiable object-detection task.","core_discovery":"The central claim is that expert-reviewed, standardized collection can convert TCM tongue diagnosis into a computable object-detection problem. The dataset contains 6,719 images annotated with 20 symptom categories, split into global and local features, averaging 2.54 labels per image, with labels reviewed by licensed practitioners. The authors benchmark nine detection models, report mAP@0.5 up to about 35%, and interpret the results as showing that the dataset is usable and that mid-sized detectors offer the best accuracy/compute trade-off. These results are meant to establish the dataset as a foundation for further AI-assisted TCM research.","pith_inferences":["A published per-label inter-annotator agreement study would let downstream users weight labels by their reliability and would likely be a prerequisite for clinical deployment.","The dataset invites a domain-shift experiment: train on the standardized images, then evaluate on smartphone-captured or clinic-casual tongue photos to see how far standardization transfers.","A natural follow-up is linking these 20 visual labels to TCM syndrome patterns or to modern laboratory measures, making the dataset a bridge between visual features and clinical endpoints."],"forward_implications":["With a common dataset and label taxonomy, results of different tongue-diagnosis models become directly comparable.","The multi-label structure lets a single image express coexisting signs, so models can be trained to output combination patterns rather than a single disease label.","The reported benchmarks give a numerical starting point (best mAP@0.5 around 35%) that future work can try to beat.","The standardized acquisition protocol makes dataset extension reproducible, so the collection can grow while preserving image conditions."],"supporting_citations":[{"why":"Establishes tongue diagnosis as one of the four TCM diagnostic methods, giving the dataset its clinical motivation.","marker":"[1]"},{"why":"Documents subjectivity and inconsistent imaging in current tongue diagnosis, the problem the standardized protocol addresses.","marker":"[2]"},{"why":"Classical TCM text that supplies the theoretical basis for the symptom labels chosen.","marker":"[10]"},{"why":"Clinical reference for tongue diagnosis used to select and define the 20 annotation categories.","marker":"[11]"},{"why":"Defines precision, recall, and mAP metrics used in the technical validation.","marker":"[12]"},{"why":"Defines the IoU measure on which the mAP@0.5 and mAP@0.5:0.95 benchmarks are computed.","marker":"[13]"}],"fun_headline_variants":["6,719 standardized tongue images for AI TCM diagnosis","TCM tongue dataset: 6,719 images, 20 symptoms, AI-ready","First standardized tongue image dataset for AI TCM diagnosis","AI TCM tongue dataset: standardized images, expert labels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the expert-reviewed labels are correct, consistent, and complete enough to serve as clinical ground truth, but the paper does not report any inter-annotator agreement or reliability statistics.","fun_headline_variants_meta":{"raw":{"variants":["6,719 standardized tongue images for AI TCM diagnosis","TCM tongue dataset: 6,719 images, 20 symptoms, AI-ready","First standardized tongue image dataset for AI TCM diagnosis","AI TCM tongue dataset: standardized images, expert labels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000583,"raw_usage":{"total_tokens":2697,"prompt_tokens":851,"completion_tokens":1846,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":1774}},"tokens_in":467,"tokens_out":1846,"duration_ms":12226,"temperature":1.0,"reasoning_tokens":1774,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:15:03.369619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select 300 images at random from the released set, have two independent licensed TCM practitioners label each image with the same 20 categories, and compute per-category Cohen's kappa. If several categories fall below roughly 0.6, the labels are too inconsistent to count as ground truth, and the reported detector rankings lose their anchor.","supporting_citations":[{"cited_title":"& Rivas-Echeverría, F","cited_arxiv_id":null,"evidence_quote":"Defines the IoU measure on which the mAP@0.5 and mAP@0.5:0.95 benchmarks are computed."},{"cited_title":"& Liu, H","cited_arxiv_id":null,"evidence_quote":"Establishes tongue diagnosis as one of the four TCM diagnostic methods, giving the dataset its clinical motivation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents subjectivity and inconsistent imaging in current tongue diagnosis, the problem the standardized protocol addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classical TCM text that supplies the theoretical basis for the symptom labels chosen."},{"cited_title":"Yellow Emperor’s Classic of Medicine, The-Essential Questions: Translation Of Huangdi Neijing Suwen","cited_arxiv_id":null,"evidence_quote":"Clinical reference for tongue diagnosis used to select and define the 20 annotation categories."},{"cited_title":"Tongue Di agnosi s i n Chi nese Medicine","cited_arxiv_id":null,"evidence_quote":"Defines precision, recall, and mAP metrics used in the technical validation."}],"review_version":2}