{"id":"0c104b47-b161-46b2-91e1-b3ecde506a60","arxiv_id":"2505.00502","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"HATIE provides a large automatically scored benchmark for text-guided image editing, with five metrics weighted to match human preference judgments.","lead":"This paper introduces HATIE, a benchmark of 18,226 images and 49,840 text editing queries, together with an automated pipeline that scores edited images on five criteria instead of relying on manual user studies. It is relevant because reproducible, large-scale evaluation is a known bottleneck for text-guided image editing research.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-alignment evidence rests on six models: held-out questions are not held-out models, so the reported r may not transfer to new editing models.","rationale":"Good-faith reading: the paper builds a large, clearly documented benchmark and the held-out correlations are real evidence. The reader's concern about detector failures is valid and is explicitly acknowledged in Appendix D.2, where 16.01% failure is dismissed after viewing 30 random samples; extreme-score substitution on those failures could bias model rankings. I do not dispute that. But the most load-bearing weakness is the unit of analysis for the alignment claim. The user study produces winning rates for six models; every reported correlation in Table 2 and Fig. 8 has n=6. The total-score r=0.714 has p≈0.11, so the headline alignment for the aggregate score is not statistically distinguishable from zero at conventional levels. Component correlations are significant, but the weights were trained on the same model pool, and the test set only holds out questions. Model identity is a strong clustering variable; without leave-one-model-out evaluation, the test r can be optimistic. The proposed check directly targets generalization, which is what 'scalable benchmark' promises. I would keep the CONDITIONAL verdict, adding a condition that the authors report cross-model generalization and the inclusive OF correlation.","tokens_in":23142,"tokens_out":8297,"duration_ms":90223,"concrete_test":"Leave-one-model-out cross-validation: for each of the six description-based models, refit all weights in Eqs. (2)–(5) using only the other five models' training user-study data, then compute Pearson and Spearman correlations on the held-out model's test-set winning rates. Also report the total-score r without excluding the OF outlier in Fig. 8. If the held-out total r falls below ~0.5 or changes sign, the human-alignment claim has not been shown to generalize to unseen models, and the paper should weaken that claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that HATIE ranks editing models like humans—is supported by correlations between user and metric winning rates computed over only six description-based models (Sec. 3.4, 4.2; Fig. 8). With n=6, the total-score r=0.714 is not significant (two-sided p≈0.11), and even component r=0.82–0.93 carry wide confidence intervals (e.g., 95% CI for r=0.82 spans roughly 0.03 to 0.98). The 'held-out' split in Sec. 4.2 holds out user-study questions, not models: the weights in Eqs. (2)–(5) were fit to maximize alignment on 1,350 training queries from these same six models, so stable model-level differences seen during fitting can inflate the test correlation. A scalable benchmark must show the learned aggregation transfers to a novel model; this has not been demonstrated. The detector-failure substitution flagged in Appendix D.2 (16.01% failure, extreme scores used) is a real secondary threat, but even if it were perfectly unbiased, the alignment claim would still rest on a six-point correlation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HATIE, a large-scale benchmark and automated evaluation pipeline for text-guided image editing. The benchmark is built on GQA, yielding 18,226 images and 49,840 editing queries across object-centric and non-object-centric tasks, and it evaluates edited images along five criteria: Image Quality, Object Fidelity, Background Fidelity, Object Consistency, and Background Consistency. The criteria are computed from CLIP, LPIPS, DINO, L2, detection confidence, and rule-based size/position scores, and are aggregated by convex combination weights fitted to a user study conducted on six description-based models. The paper reports Pearson correlations between model winning rates from user study and HATIE on a held-out split of user-study questions, compares against conventional metrics, and provides benchmark tables for nine editing models. The main claim is that HATIE provides a fully automated, human-aligned evaluation that can be used to assess existing and new text-guided image editing models.","tokens_in":23549,"tokens_out":6018,"duration_ms":61361,"significance":"If the human-alignment claim were established for unseen editing models, HATIE would be a valuable contribution to the community: it provides a much larger and more diverse query set than existing benchmarks, ships code, uses a structured multi-metric aggregation, and includes extensive appendix documentation. The comparison with conventional metrics and the bootstrapped benchmark scores are strengths. However, the current evidence for human alignment is weaker than claimed: the correlations are computed on only six models, the held-out split does not hold out models, the total-score correlation is below the reported 0.8 threshold, and the object-fidelity correlation excludes an outlier. The significance of the work therefore depends on additional validation of transfer to novel models and on corrected, uncertainty-aware reporting of the alignment results.","major_comments":[{"comment":"The text states that 'Pearson's coefficient r over 0.8 throughout every criterion', but Table 2 reports Total Score r = 0.7143, which is below 0.8. In addition, Fig. 8 shows that the Object Fidelity correlation of 0.8208 is computed after excluding a marked outlier, so that coefficient is based on five rather than six model points. With only six models, the total-score correlation is not statistically significant at the 0.05 level, and even the component correlations have wide confidence intervals. Please report the exact sample size after outlier exclusion, provide confidence intervals or p-values for all reported correlations, and revise the abstract and Section 6 claims so they accurately reflect the observed total-score correlation.","section":"Section 4.2, Table 2, Fig. 8"},{"comment":"The held-out split in Section 4.2 holds out user-study questions, not editing models. The weights in Eqs. (2)-(5) are fitted to winning rates from the same six description-based models that are then used to compute the test-set correlations, so the experiment demonstrates transfer across questions for those six models but does not demonstrate transfer to a new editing model. Since the introduction and abstract position HATIE as a scalable evaluation method for 'existing and new image editing models', the paper should provide leave-one-model-out cross-validation or a separate alignment study with a held-out model. Without such evidence, the human-alignment claim must be restricted to the six models studied.","section":"Sections 3.4 and 4.2"},{"comment":"The evaluation relies on instance segmentation, but Appendix D.2 reports a 16.01% detection failure rate, and the task-specific workflows in Appendix C.1 assign extreme scores when the target object is not detected (e.g., Object Fidelity 0 for addition/replacement/resizing/attribute-change tasks, and 1 for removal tasks). Appendix D.2 argues qualitatively that most failures are due to absence of the target object, but no quantitative evidence is provided that detection failures are uncorrelated with editing quality. If models that produce less canonical or softer object appearances are more likely to trigger detection failures, the imputed extreme scores can bias model rankings independently of human judgment. Please report per-model detection failure rates and perform a sensitivity analysis that recomputes the alignment correlations and benchmark rankings with detection-failure cases excluded or with alternative imputation.","section":"Appendix D.2 and Appendix C.1"},{"comment":"The weight fitting procedure uses grid search to maximize Pearson correlation on model winning rates from only six models, and the training-set correlations in Table I reach 1.0 for Total Score and Background Consistency. This does not establish that the fitted weights are stable or that the aggregation generalizes. Please report the fitted weight values actually used, the number of grid points evaluated, and the stability of the weights under bootstrap resampling (of both questions and models). Without this information, the Total Score weights in Eq. (5) may be overfit to the six-model sample.","section":"Section 3.4 and Appendix C.3"}],"minor_comments":[{"comment":"In the paragraph following Eq. (2), the text says 'The overall Object Fidelity score σOC' but the symbol being defined is σOF; please correct this typo.","section":"Section 3.3, Eq. (2)"},{"comment":"The relation between 4,050 sampled images, 2,025 queries, 2,700 training images from 1,350 queries, and 1,350 test images from 675 queries is not transparent given that six models generate one output per query; please clarify how many model outputs per query are sampled and how the pairwise comparisons are constructed.","section":"Section 3.4"},{"comment":"The symbols ρ and τ are used without definition; please define them as Spearman's rank and Kendall's tau, respectively, in the caption or in the text before first use.","section":"Table 2"},{"comment":"For Object Fidelity, please report the correlation both with and without the marked outlier, and state which model and query produced the outlier, so the reader can judge whether its exclusion is justified.","section":"Fig. 8"},{"comment":"The thresholds r1, r2, and r3 in the size score are described as empirically set, but their values are not given; please report the actual values so that the rule-based size score is reproducible.","section":"Appendix C.2"},{"comment":"The claim that alignment is tested on an 'unseen new dataset' refers to new images from ImageNet but still uses the same six editing models; please rephrase to avoid implying generalization to novel models, and report the number of queries and models used in Fig. X.","section":"Appendix D.4 and Fig. X"}],"recommendation":"major_revision","confidential_remarks":"The central idea is strong and the dataset is a useful resource, but the paper currently overclaims human alignment. The main issues are fixable: the reported total-score correlation contradicts the 'r>0.8' statement, the held-out evaluation does not include held-out models, and the detection-failure substitution needs a sensitivity analysis. I would send the manuscript back with a clear request for these experiments and reporting changes, rather than reject it, because the core benchmark and pipeline are valuable if the alignment claim is properly scoped and supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, well-documented benchmark paper that the editing subfield will likely want to use. The dataset scale and the fully automated pipeline are real contributions beyond EditVal and TEdBench. The careful task-specific workflows, the filtering, and the release of code/data all look responsible. Credit where due: the held-out correlations for the component criteria (r 0.82–0.93) are real evidence that the weighted combination tracks human preference better than the individual off-the-shelf metrics, which Table 3 shows to be weak. That is the paper's strongest empirical point.\n\nThe soft spots are mostly about the breadth of the alignment claim. The alignment is computed on winning rates over only six description-based models. With n=6, the total-score r=0.714 is not significant (p≈0.11), and the 'held-out' split holds out questions, not models: the weights were fit to these same six models, so stable model-level differences can inflate the test correlation. I don't think this sinks the benchmark—the resource is still useful—but the sentence claiming r>0.8 for every criterion is inaccurate (total is 0.714), and the Object Fidelity outlier exclusion needs the inclusive value reported. Also, the 16% detector failure with extreme-score substitution is a plausible bias channel; the appendix's example images are reassuring that failures are often the object truly being absent, but that's not a formal bias test.\n\nBottom line: as a benchmark and evaluation pipeline, it deserves a serious referee and probably publication after revision. The authors should either add more editing models to the alignment study (even a couple of instruction-based ones) or explicitly limit the human-alignment claim to the tested model family. Report the total-score r accurately, show the detector-failure robustness analysis, and this becomes a citation-worthy resource. I'd bring it to reading group as an example of a well-executed benchmark effort.","headline":"HATIE is a genuinely useful benchmark resource, but its central human-alignment claim rests on only six models, and the paper should either add more or soften the conclusion.","tokens_in":24138,"tokens_out":1935,"would_cite":true,"duration_ms":19950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully automated benchmark for text-guided image editing, built on 49,840 queries and five weighted visual scores, reproduces human quality judgments with a total-score correlation of about 0.71 and per-criterion correlations typically…","keywords":["text-guided image editing","benchmark","human-aligned evaluation","automated evaluation","diffusion models","user study correlation","instance segmentation","image editing metrics"],"falsifier":"Compute the detection-failure rate separately for each editing model and correlate it with human preference scores on the same images; if models that produce unusual or degraded object appearances fail detection more often, the extreme-score replacement will depress their HATIE scores without depressing human ratings. A direct check is to rerun the benchmark after excluding all images where the segmenter fails and see whether model rankings change materially.","tokens_in":22892,"feed_emoji":"🖼️","tokens_out":5646,"duration_ms":49963,"temperature":0.7,"pith_summary":"This paper argues that subjective evaluation of text-guided image editing can be automated without losing human alignment. It introduces HATIE, a benchmark of 18,226 images and 49,840 edit queries covering object addition, removal, replacement, attribute change, resizing, background change, and style change. Instead of relying on any single model-based score, it combines five criteria—object fidelity, background fidelity, object consistency, background consistency, and image quality—and fits the combination weights to pairwise human preference judgments. On held-out user study questions, the resulting total score tracks human winning rates with Pearson correlation about 0.71, and each individual criterion reaches 0.82–0.93. If correct, this gives the field a reproducible, labour-free way to compare editing models and tune their hyperparameters.","feed_headline":"Edit-quality benchmark tracks human judgment at r=0.71","feed_subtitle":"HATIE scores 49,840 editing queries with five weighted metrics, replacing costly user studies for model comparison.","key_machinery":"The load-bearing object is the weighted five-criteria score: object fidelity, background fidelity, object consistency, background consistency, and image quality, aggregated as $\\sigma_{\\text{Total}}=\\sum_{x\\in X}w_x\\sigma_x$ with weights fitted to user-study winning rates. Task-specific workflows decide which criteria apply; for object-centric edits an instance segmentation model crops the target object, so fidelity is measured on the object alone and consistency compares the object or background across original and edited images. The component metrics include CLIP alignment, detection confidence, LPIPS, DINO similarity, L2 distance, and rule-based position and size scores. The argument that the whole is human-aligned rests on fitting the weights to pairwise human comparisons and then verifying correlation on set-aside questions.","core_discovery":"The paper's central claim is that a structured combination of off-the-shelf metrics, organized into five categories and weighted by human preference data, can serve as a proxy for human evaluation of text-guided image editing. The benchmark defines fidelity as whether the instructed change happened, consistency as whether unintended content was preserved, and image quality as the distributional realism of outputs; instance segmentation separates both fidelity and consistency into object and background components. Each component score is a convex combination of simpler signals such as CLIP text-image alignment, detection confidence, LPIPS and DINO similarities, L2 distance, and rule-based position and size scores. The weights are selected by grid search to maximize Pearson correlation between model winning rates computed from user studies and from the automated scores. On a held-out user study test set, the paper reports correlations of 0.82–0.93 for the four component criteria and about 0.71 for the total score, and shows that each single conventional metric correlates much worse with human judgment.","pith_inferences":["If segmenter failures are not neutral—models that produce less canonical object appearances may be missed more often—the 16.01% failure rate with extreme-score replacement could bias rankings; this is testable by comparing HATIE rankings with and without failure cases.","The weight-fitting procedure uses only six description-based models and a small participant pool, so weights optimized on that model set may not be optimal for future models with different failure modes; periodic re-fitting or uncertainty reporting would strengthen the human-alignment claim.","The same five-criteria structure could transfer to open-vocabulary editors by replacing the fixed-class segmenter with an open-vocabulary detector, extending the benchmark beyond the 76 COCO classes.","Because the user study asks binary pairwise choices by majority vote, the fitted weights inherit the granularity of those comparisons; fine-grained preference strengths, such as ratings, might change the optimal weights."],"forward_implications":["Editing models can be compared on a common 49,840-query benchmark without running new user studies, because the automated total score reproduces human preferences on held-out questions.","Model rankings can be decomposed by edit type and object class, exposing task-specific strengths such as Imagic on attribute and background changes or InstDiff on resizing.","The benchmark is sensitive enough to guide hyperparameter choice: varying edit strength produces monotone fidelity and consistency trends with a clear optimum around the fidelity-consistency trade-off.","Individual conventional metrics are insufficient: none of CLIP alignment, detection rate, LPIPS, DINO, L2, position or size scores, or FID alone matches human judgment as well as the combined score.","Because the pipeline is fully automated and does not depend on external API services, evaluations can be reproduced and extended to new models and datasets."],"supporting_citations":[{"why":"Supplies the base images and scene-graph annotations from which editable objects and feasible edit queries are generated.","marker":"[14]"},{"why":"Provides the instance segmentation model that detects target objects and separates object from background for fidelity and consistency scoring.","marker":"[32]"},{"why":"Supplies the CLIP-based text-image alignment score used as a component of object and background fidelity.","marker":"[11]"},{"why":"Provides the LPIPS perceptual distance used to compare original and edited objects and backgrounds in consistency scores.","marker":"[35]"},{"why":"Provides DINO embedding similarity used as another perceptual consistency component alongside LPIPS and L2 distance.","marker":"[23]"},{"why":"Provides the FID score used to measure overall image quality of an editing model's output set.","marker":"[12]"},{"why":"Supplies the detection-rate metric and partial automated evaluation approach that HATIE compares against as a conventional baseline.","marker":"[2]"}],"fun_headline_variants":["HATIE: Automated editing benchmark aligned with human judgment","Benchmark replaces user studies for image editing quality","Five metric scores predict human preference for editing","Text-editing eval benchmark matches humans at r=0.71"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that when the instance segmentation model fails to detect a target object in 16.01% of evaluation images, the failure is caused by the object being genuinely absent or by detector limitations that are uncorrelated with editing quality; on that basis the pipeline replaces missing detections with extreme fidelity scores (0 for most tasks, 1 for removal).","fun_headline_variants_meta":{"raw":{"variants":["HATIE: Automated editing benchmark aligned with human judgment","Benchmark replaces user studies for image editing quality","Five metric scores predict human preference for editing","Text-editing eval benchmark matches humans at r=0.71"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1390,"prompt_tokens":880,"completion_tokens":510,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":446}},"tokens_in":496,"tokens_out":510,"duration_ms":4976,"temperature":1.0,"reasoning_tokens":446,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:40:39.790730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the detection-failure rate separately for each editing model and correlate it with human preference scores on the same images; if models that produce unusual or degraded object appearances fail detection more often, the extreme-score replacement will depress their HATIE scores without depressing human ratings. A direct check is to rerun the benchmark after excluding all images where the segmenter fails and see whether model rankings change materially.","supporting_citations":[{"cited_title":"GQA: A new dataset for real-world visual reasoning and com- positional question answering","cited_arxiv_id":null,"evidence_quote":"Supplies the base images and scene-graph annotations from which editable objects and feasible edit queries are generated."},{"cited_title":"Detectron2","cited_arxiv_id":null,"evidence_quote":"Provides the instance segmentation model that detects target objects and separates object from background for fidelity and consistency scoring."},{"cited_title":"CLIPScore: A reference-free evaluation metric for image captioning”","cited_arxiv_id":null,"evidence_quote":"Supplies the CLIP-based text-image alignment score used as a component of object and background fidelity."},{"cited_title":"a photo of a crimson cat","cited_arxiv_id":null,"evidence_quote":"Provides the LPIPS perceptual distance used to compare original and edited objects and backgrounds in consistency scores."},{"cited_title":"DINOv2: Learning robust visual fea- tures without supervision","cited_arxiv_id":null,"evidence_quote":"Provides DINO embedding similarity used as another perceptual consistency component alongside LPIPS and L2 distance."},{"cited_title":"GANs trained by a two time-scale update rule con- verge to a local nash equilibrium","cited_arxiv_id":null,"evidence_quote":"Provides the FID score used to measure overall image quality of an editing model's output set."}],"review_version":1}