{"id":"fed55547-b8d4-4323-a5f9-87844245f609","arxiv_id":"2411.13842","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new dataset and detector suite localize human artifacts in text-to-image outputs and feed back into generation to reduce them.","lead":"This paper builds a dataset of over 37,000 AI-generated images labeled for human body artifacts, and trains detectors that spot distorted, missing, or extra body parts across different generators. The same detectors then guide diffusion model finetuning and inpainting to reduce these artifacts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim of strong generalization to unseen generators has no quantitative ground-truth evaluation: all AP50/AUC numbers are in-domain, and the out-of-domain evidence is qualitative examples plus prediction statistics that do not measure correctness.","rationale":"The paper has real strengths: a large multi-generator dataset, in-domain quantitative detection results, ablations, and a small independent user study for the finetuning claim. The reader's weakest assumption about annotation subjectivity is valid, and the absence of inter-annotator agreement does undermine the precision of all AP50/AUC numbers. However, the most load-bearing point for the paper's headline promise is that the 'strong generalization, even on images from unseen generators' claim is asserted without any quantitative ground-truth evaluation on those domains. The qualitative figures and prediction statistics in Tab. 3 are suggestive but do not measure detection correctness; a domain-shifted model with well-calibrated confidence would produce exactly the observed pattern. This is a gap in evidence, not an internal inconsistency, and it can be closed by a modest annotation effort. The reader noted this issue in the rationale ('Generalization to unseen generators is only qualitative') but selected annotation subjectivity as the weakest assumption, so I mark agreement as partial. The verdict should remain CONDITIONAL: the central claims are plausible and partially supported, but the unseen-generator generalization claim needs a quantitative test before full acceptance.","tokens_in":17815,"tokens_out":4183,"duration_ms":46686,"concrete_test":"Build a small labeled evaluation set for the unseen generators using the same annotation protocol: sample about 200 images per generator from SD1.4, PixArt-Σ, FLUX.1-dev, and Sana, have at least two annotators label all local and global artifacts, report inter-annotator agreement, and compute AP50/AUC per domain with the released HADM checkpoints. Compare these numbers with the in-domain AP50 values in Tab. 1 and Tab. 2; if the unseen-domain AP50 is near chance or substantially below in-domain performance, the strong-generalization claim is unsupported, whereas comparable AP50 would resolve the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and Sec. 1 is that HADM can 'demonstrate strong generalization, even on images from unseen generators.' Sec. 4.1 reports AP50 and AUC only on the HAD validation set (SDXL, DALLE-2, DALLE-3, Midjourney). The unseen-domain evaluation in Sec. 4.2 (SD1.4, PixArt-Σ, FLUX.1-dev, Sana, 300W) provides no ground-truth labels. The evidence is selected top predictions (Figs. 6-7) and prediction-count/confidence statistics (Tab. 3). Neither can establish the claim: a detector can output fewer, lower-confidence boxes on cleaner generators because of domain shift or confidence calibration rather than because it recognizes artifacts. Tab. 3 compares HADM with HumanRefiner on these statistics, but prediction statistics are not a correctness measure. Consequently, the advertised property that makes HADM a reusable benchmark for rapidly evolving T2I models currently rests on cherry-picked examples rather than measured performance. This concern is distinct from, and more direct than, the annotation-noise issue: even if the HAD labels were perfectly reliable, the unseen-generator claim would still be unsupported without quantitative evaluation on labeled out-of-domain data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper curates the Human Artifact Dataset (HAD), containing 37,554 images from SDXL, DALL-E 2, DALL-E 3, and Midjourney with 84,852 bounding-box annotations covering six local artifact classes (face, torso, arm, hand, leg, feet) and twelve global artifact classes (missing and extra body parts). It trains two detectors, HADM-L and HADM-G, using ViTDet with an EVA-02 backbone and a Cascade R-CNN head, reporting in-domain AP50 of 43.3 for local and 23.9 for global artifacts and an AUC-based comparison against HumanRefiner, HPS-v2, ImageReward, GPT-4o, LLaVA, and Llama on a binarized classification version of the validation set. The paper further claims strong generalization to unseen generators, including SD1.4, PixArt-Sigma, FLUX.1-dev, Sana, and the real-image dataset 300W; it presents a LoRA finetuning of SDXL that uses HADM predictions as negative-prompt feedback, and an iterative inpainting pipeline for artifact correction, supported by a 15-participant user study that prefers the finetuned model.","tokens_in":18102,"tokens_out":13725,"duration_ms":116494,"significance":"If the claims held as stated, HAD would be a valuable released benchmark for human-structure evaluation in text-to-image models, and the demonstration that a dedicated detector outperforms strong vision-language models and aesthetics metrics on in-domain artifact ranking is a genuinely useful result. The paper ships dataset and models on GitHub, includes ablations of backbone and training choices, and provides a user study as an independent check of the finetuning claim; these are real strengths. However, the headline claim of strong generalization to unseen generators currently rests on qualitative examples and prediction statistics rather than on any ground-truth-based metric, and the quantitative finetuning evaluation is partly circular because it relies on HADM itself; both issues must be resolved for the stated contribution to stand.","major_comments":[{"comment":"The claim that HADM 'demonstrate[s] strong generalization, even on images from unseen generators' is not supported by any ground-truth-based metric on out-of-domain data. All AP50 and AUC results (Tabs. 1-2, Fig. 4) are computed on the in-domain HAD validation set; for SD1.4, PixArt-Sigma, FLUX.1-dev, Sana, and 300W the evidence is a selection of top predictions (Figs. 6-7) and prediction-count/confidence statistics (Tab. 3). Neither measures correctness, and the decreasing prediction counts on FLUX.1-dev and Sana in Tab. 3 could equally reflect confidence calibration or domain shift, while the low count on 300W could reflect a false-negative pattern; the interpretation of Tab. 3 as sensitivity to artifact severity is underdetermined. The stress-test concern lands here: even with perfectly reliable labels, the unseen-generator claim would remain unsupported without a labeled out-of-domain evaluation. I recommend annotating a labeled subset of images from at least the four unseen generators and reporting AP50/AUC there, or substantially tempering the generalization claim in the abstract and the contribution list.","section":"Abstract, Sec. 1 contributions, and Sec. 4.2 'Performance on Unseen Domains'"},{"comment":"The primary quantitative evidence for artifact reduction uses HADM scores on images generated by the finetuned model, but HADM is also the detector that selected the finetuning training data (Sec. 3.3). A model trained to avoid detector-flagged regions is expected to obtain lower scores on that same detector even without genuine artifact reduction, so Figs. 12-13 and the hyperparameter analysis in Fig. 14 are partly circular. The user study in Tab. 4 is an independent check, but as reported it is small (15 participants, 200 pairs), no statistical test or confidence interval is given, and the stated preference (55% vs. 38.7%) is modest; the study design and inter-participant variability should be reported. I suggest evaluating the finetuned model on a fresh human-annotated sample outside the finetuning loop, or adding a second independent detector as an additional quantitative sanity check.","section":"Sec. 5.1 and supplementary Sec. 10 (Figs. 12-13, 14)"},{"comment":"The ground truth is based on human annotation with no reported inter-annotator agreement, and the authors themselves acknowledge missed annotations and subjective ambiguity (Fig. 5 and the discussion of 'occasional oversight by annotators' and 'corner cases'). Because these ambiguous cases are stated to concentrate in the DALLE-3 and Midjourney domains, where AP50 is lowest (Tab. 1), the per-domain performance differences may partly reflect label noise rather than detector behavior, which in turn affects the validity of the in-domain benchmark claim. I recommend reporting annotation agreement (e.g., Cohen's kappa or IoU-based agreement on a re-annotated subset) and, if feasible, re-estimating per-domain AP50 after adjudication of the ambiguous cases.","section":"Sec. 3.1.2 (Annotation Pipeline) and Sec. 4.2"}],"minor_comments":[{"comment":"Several per-category AP50 cells are computed on one or two instances (e.g., DALLE-2 torso at 100.0/1, DALLE-3 extra torso at 33.3/1, and multiple Tab. 2 cells with fewer than ten instances), so per-category values are statistically unstable; please add a caveat, report confidence intervals, or exclude categories below a minimum instance count.","section":"Tabs. 1-2"},{"comment":"The y-axis of Fig. 4 is not labeled, and the paper should state the score convention for the quality-based baselines: for HPS-v2 and ImageReward, where higher scores indicate higher quality, an AUC below 50% may simply reflect inverse correlation rather than 'poor' performance, so the sign convention and the interpretation of below-50% AUC should be clarified.","section":"Fig. 4 and Sec. 4.1"},{"comment":"The caption is ambiguous about which boxes are predictions and which constitute the reference annotation, particularly because the blue boxes are described both as detections and as context; please specify the visual encoding of red and blue boxes and what FP and FN are measured against.","section":"Fig. 5 caption"},{"comment":"The inpainting selection rule 'closest to half of the original score' is introduced as an empirical heuristic without supporting evidence or an ablation; please provide a justification or explicitly mark it as a limitation.","section":"Supplementary Sec. 7.4"},{"comment":"The inpainting pipeline is described as a 'novel application,' yet the paper itself cites iterative inpainting with confidence feedback in [65,66]; please clarify the specific novelty relative to those works and to HumanRefiner's refinement stage.","section":"Abstract and Sec. 5.2"},{"comment":"Many reference entries carry stray trailing numbers after the year (e.g., [2] '... CVPR, pages 6154-6162, 2018. 4, 1'), which appear to be citation-page artifacts and should be cleaned up in the final version.","section":"Reference list"},{"comment":"The ethical statement discourages applying the model to real-human images, but Sec. 4.2 evaluates on the real-image dataset 300W; please add one sentence clarifying that this evaluation is a controlled analysis of false-positive behavior and not an endorsement of deployment on real humans.","section":"Ethical Considerations and Sec. 4.2"}],"recommendation":"major_revision","confidential_remarks":"The dataset release is a tangible contribution and the in-domain detection results are credible; the main risk is the gap between what is measured and what is claimed in the abstract regarding unseen-generator generalization and artifact reduction. All requested revisions are feasible within the scope of a normal revision: adding a labeled out-of-domain evaluation (or tempering the claims), de-circularizing the finetuning evaluation, and reporting annotation agreement. The differentiation from HumanRefiner beyond training-data diversity should also be sharpened. The manuscript is a reasonable fit for a vision venue, and I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful dataset-and-detector paper that overclaims its out-of-domain generalization. If you need a benchmark for human artifacts in T2I images, the HAD dataset and HADM detectors are worth a look; the in-domain numbers are respectable. But the abstract says \"strong generalization, even on images from unseen generators,\" and that specific claim is not quantitatively supported.\n\nWhat's actually new: a multi-generator artifact dataset (37k images, 84k boxes) spanning SDXL, DALLE-2/3, Midjourney, with separate local/global artifact labels; trained ViTDet-based detectors (HADM-L/G); and a finetuning pipeline that uses detector predictions as negative-prompt identifiers. The dataset is the main contribution. The detector ablations show real gains from backbone capacity and real-image regularization. The AUC comparison against HumanRefiner, HPS, ImageReward, and several VLMs is a fair way to show the task is non-trivial and that aesthetics scores don't capture it.\n\nThe soft spots are exactly where the reader's report points. First, the unseen-generator evaluation is qualitative (Figs. 6-7) plus prediction statistics (Tab. 3). Prediction counts and mean confidences don't measure correctness; a detector can produce fewer, lower-confidence boxes on cleaner generators because of calibration shift. To back the central generalization claim, you need labeled ground truth on at least one held-out generator. Second, the finetuning improvement is evaluated with the same HADM that generated the feedback: Figs. 12-13 show HADM scores drop for the finetuned model. That's circular. The 15-participant user study is an independent check, and it does lean in favor of the finetuned model, but it's small and reported without statistics. Third, annotation noise is real—the authors acknowledge ambiguity and missed labels—and there's no inter-annotator agreement. That doesn't sink the detection results, but it means metrics have a ceiling that's unknown.\n\nProportionally: the core resource is solid; the paper isn't a paradigm shift, but it's a practical contribution. The gaps are fixable—add labeled OOD evaluation, run a bigger user study or an independent artifact metric, report agreement. Who this is for: anyone building evaluators or quality control for T2I. It deserves a serious referee, but the revision needs to either soften the generalization claim or prove it.","headline":"Useful multi-generator artifact dataset and detector, but the central generalization claim lacks quantitative out-of-domain evaluation.","tokens_in":18611,"tokens_out":1826,"would_cite":true,"duration_ms":17989,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that human-body artifacts in text-to-image output can be localized by a detector trained on many generators, and that feeding its predictions back as negative prompts reduces those artifacts.","keywords":["human artifact detection","text-to-image generation","diffusion models","artifact localization","image quality evaluation","generative model benchmarking","negative prompt finetuning","inpainting"],"falsifier":"Re-annotate a random sample of HAD images with multiple independent annotators and measure pairwise agreement; if agreement is near chance on images from the higher-fidelity generators, the AP50 values that support HADM's generalization are measuring label noise rather than detection ability.","tokens_in":17621,"feed_emoji":"🖐️","tokens_out":9664,"duration_ms":79050,"temperature":0.7,"pith_summary":"This paper tries to establish that the broken hands, distorted faces, and extra or missing limbs produced by text-to-image models can be treated as a detect-and-correct problem rather than an unsolved quality mystery. The authors build a dataset of over 37,000 generated images with bounding-box labels for two artifact classes: local defects in six body parts and global anatomical errors such as missing or extra parts. They train separate detectors on this data and claim the detectors find artifacts across the four generators they trained on while still working on generators never seen in training. They further claim that using detector predictions as negative-prompt feedback during diffusion-model finetuning reduces human artifacts, and that the same detector can drive an iterative inpainting loop that corrects local artifacts in arbitrary images.","feed_headline":"Detector finds bad hands and limbs in AI images, then trims them","feed_subtitle":"Trained on 37,000 annotated images, it also catches errors in generators it never saw.","key_machinery":"The load-bearing mechanism is the pairing of a labeled artifact dataset with a detection model that has seen healthy human bodies as negative examples. Local artifacts are annotated per body part and global artifacts as missing or extra parts, with a separate detector for each. Real human images with no artifact boxes are included in training so the model learns that normal anatomy is not an artifact; this is what lets it transfer to unseen generators and report fewer detections as image quality improves. The correction loop then uses detector confidence scores to decide which generations to keep: finetuning prepends artifact-type identifiers that are later used as negative prompts, and inpainting picks the candidate with the lowest artifact score.","core_discovery":"The central claim is that human artifacts in text-to-image output are localizable, and that a detector trained on diverse generator outputs generalizes better than aesthetic scorers, vision-language models, and prior artifact detectors. The authors curate the Human Artifact Dataset with 84,852 labeled instances across six local body-part classes and twelve global structural classes, train separate local and global detection models, and report that the local model performs best on generators with frequent artifacts while still catching subtle errors in higher-fidelity generators, including some not seen in training. The paper further claims that using detector-selected images with artifact-type identifiers as negative prompts during low-rank finetuning of a diffusion model reduces artifact scores, that human raters prefer the finetuned model, and that the detector selects better results in an iterative inpainting correction pipeline.","pith_inferences":["If the generalization claim holds, HADM-style detectors could serve as an automatic anatomy-specific reward for aligning future generation models, complementing coarse aesthetic preference models.","The acknowledged annotator ambiguity suggests the low scores on high-fidelity generators may be as much label noise as detector error; re-annotating a sample with adjudicated majority labels would clarify this.","A natural extension is to combine artifact boxes with pose or segmentation priors, which could resolve the ambiguous extra-limb cases the paper itself flags as corner cases."],"forward_implications":["HAD and HADM provide a reusable benchmark for ranking text-to-image models by how often they distort human bodies, not just by aesthetics.","Detector feedback can be folded into diffusion finetuning: artifact-type identifiers used as negative prompts lower both detector-measured and user-perceived artifact severity.","Because the detector is trained to treat real human images as clean, it can be pointed at new generators and will flag fewer and lower-confidence artifacts as generator quality improves.","The same detector can act as an automatic quality gate inside an inpainting loop, repeatedly fixing the most severe local artifact in an image."],"supporting_citations":[{"why":"Supplies the off-the-shelf vision-transformer detection architecture that HADM is built on.","marker":"[24]"},{"why":"Supplies the visual backbone used by the detector.","marker":"[7]"},{"why":"Provides the primary text-to-image generator used for dataset images and the generation model that receives feedback finetuning.","marker":"[42]"},{"why":"Closest prior human-artifact detector that HADM is compared against; its single-generator training motivates HADM's multi-generator design.","marker":"[6]"},{"why":"Provides the low-rank adaptation method used to finetune the diffusion model with artifact-type identifiers.","marker":"[14]"},{"why":"Contributes real-world images with no artifact boxes, used as regularizers to suppress false positives on normal anatomy.","marker":"[26]"},{"why":"Supplies an aesthetic human-preference baseline that HADM must beat to show artifact detection is not captured by aesthetics.","marker":"[59]"}],"fun_headline_variants":["Spotter for mangled AI limbs gets better with 37k examples","Detector catches body glitches in generated images, then erases them","Generalizes to unseen generators: detector finds flawed hands and fixes","AI artifact detector improves image quality by inpainting bad parts","37k-image training yields detector that fixes human glitches from any AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline assumes that 'artifact' is a well-defined ground truth: individual human annotators decide what counts as a distorted, missing, or extra body part, and the paper reports no agreement check, so subjective and missed labels are baked into both the training signal and every evaluation number.","fun_headline_variants_meta":{"raw":{"variants":["Spotter for mangled AI limbs gets better with 37k examples","Detector catches body glitches in generated images, then erases them","Generalizes to unseen generators: detector finds flawed hands and fixes","AI artifact detector improves image quality by inpainting bad parts","37k-image training yields detector that fixes human glitches from any AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1417,"prompt_tokens":925,"completion_tokens":492,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":400}},"tokens_in":541,"tokens_out":492,"duration_ms":28696,"temperature":1.0,"reasoning_tokens":400,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:48:32.519188+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of HAD images with multiple independent annotators and measure pairwise agreement; if agreement is near chance on images from the higher-fidelity generators, the AP50 values that support HADM's generalization are measuring label noise rather than detection ability.","supporting_citations":[{"cited_title":"Girshick, and Kaiming He","cited_arxiv_id":null,"evidence_quote":"Supplies the off-the-shelf vision-transformer detection architecture that HADM is built on."},{"cited_title":"EV A-02: A visual representa- tion for neon genesis","cited_arxiv_id":null,"evidence_quote":"Supplies the visual backbone used by the detector."},{"cited_title":"SDXL: Improving latent diffusion mod- els for high-resolution image synthesis","cited_arxiv_id":null,"evidence_quote":"Provides the primary text-to-image generator used for dataset images and the generation model that receives feedback finetuning."},{"cited_title":"HumanRefiner: Benchmarking abnormal human generation and refining with coarse-to-fine pose-reversible guidance","cited_arxiv_id":null,"evidence_quote":"Closest prior human-artifact detector that HADM is compared against; its single-generator training motivates HADM's multi-generator design."},{"cited_title":"Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen","cited_arxiv_id":null,"evidence_quote":"Provides the low-rank adaptation method used to finetune the diffusion model with artifact-type identifiers."},{"cited_title":"Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C","cited_arxiv_id":null,"evidence_quote":"Contributes real-world images with no artifact boxes, used as regularizers to suppress false positives on normal anatomy."}],"review_version":1}