{"id":"947bccda-18e3-496f-b226-f0cba204e930","arxiv_id":"2502.04757","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"ELITE is a curated benchmark and a toxicity-weighted rubric evaluator for vision-language model safety, claiming better alignment with human judgments than existing automated safety checks.","lead":"This paper introduces ELITE, a benchmark and automated evaluator for measuring how safely vision-language models respond to harmful image-text prompts. The evaluator adds a toxicity score to an existing rubric, and the benchmark filters thousands of pairs to keep only those that actually provoke unsafe answers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AUC superiority claim is computed on a disagreement-enriched sample, so the headline comparison with StrongREJECT is not established; the paper's own representative 228-sample data are never used for that comparison.","rationale":"The reader's weakest_assumption (representativeness of the three filtering models) is a real concern for the benchmark's generalization, but the paper partially mitigates it with the human validation in Sec. 6.2 and the observation that many diverse VLMs show high E-ASR on the filtered benchmark. The concern I raise is more directly tied to the evaluator's central claim: the only comparative evidence that ELITE beats StrongREJECT is Fig. 4, computed on a sample deliberately enriched for disagreements between those two methods. A disagreement-conditioned ROC comparison can exaggerate the gap and is not a valid estimate of population alignment. The paper's own additional 228-sample dataset would provide the necessary check, but the paper does not run the ELITE-vs-StrongREJECT AUC on it, instead reporting only ELITE-to-human correlation. Thus the first central claim is not yet established; the required analysis is cheap and could be added without new data collection. If that analysis shows no significant advantage for ELITE, the paper's evaluator contribution would be substantially weakened. The benchmark could still be useful (Table 8 human validation is independent of StrongREJECT), but the 'better alignment' headline would not hold. Therefore the verdict remains CONDITIONAL pending this specific analysis.","tokens_in":29141,"tokens_out":14114,"duration_ms":139430,"concrete_test":"On the 228-sample dataset from Sec. 6 (or a new random sample stratified by taxonomy from the full set), compute ROC AUC of ELITE (GPT-4o) and StrongREJECT (GPT-4o) against the majority-vote human safe/unsafe labels, with a DeLong test for the difference. If ELITE's AUC is not significantly higher than StrongREJECT's (e.g., non-significant or reversed), the Fig. 4 result is an artifact of disagreement-enriched sampling and the human-alignment claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's first central claim is that ELITE aligns with human judgments better than StrongREJECT (Fig. 4, AUC 0.77 vs 0.46). That comparison is run on a 963-pair sample that, per Sec. 5.1, 'primarily collected where the evaluation results differed between the ELITE evaluator and existing evaluation methods.' Conditioning on disagreement between the two methods being compared invalidates the ROC estimates as measures of population-level alignment: the sample is enriched for cases where the two methods are anti-correlated, so the absolute AUC values and their difference can be inflated or deflated by the selection rule. The paper concedes this in Sec. 6 and collects a new 228-sample dataset (110 from ELITE, 118 filtered out) with majority-vote human safety labels, but it only reports ELITE-to-human toxicity correlations (Fig. 5) and never recomputes the ELITE vs StrongREJECT AUC on this more representative sample. Therefore the claim of superior human alignment remains unverified by the paper's own less-biased data. This is load-bearing because the evaluator is the foundation of the benchmark: if it is not actually better aligned than StrongREJECT, the filtering pipeline has no principled advantage.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ELITE, a rubric-based automated safety evaluator for vision-language models that extends StrongREJECT by multiplying the existing (1 - refused) x (specific + convincing)/2 term with a toxicity score (Eq. 2). Using this evaluator, the authors construct the ELITE benchmark of 4,587 image-text pairs by filtering existing safety benchmarks and adding generated pairs, retaining pairs for which at least two of three open-source VLMs obtain an ELITE score of at least 10 (Sec. 3.4). The paper claims that (i) the ELITE evaluator aligns better with human judgments than StrongREJECT (Fig. 4, AUC 0.77 vs 0.46) and safeguard models (Table 6), and (ii) the ELITE benchmark has higher quality and diversity than prior benchmarks (Table 4, E-ASR 63.59 vs 27.75 for LLaVa-v1.5-7B on VLGuard). A second human evaluation with 228 samples is presented in Sec. 6 to address the acknowledged bias in the first human study.","tokens_in":29399,"tokens_out":6826,"duration_ms":68872,"significance":"The problem targeted by this paper is important: existing VLM safety benchmarks rely on automated metrics that frequently misclassify vague or merely descriptive responses as harmful and miss implicit harmful content. The addition of a toxicity dimension to the StrongREJECT rubric is a natural and potentially effective extension. The ELITE benchmark, spanning 11 taxonomies and including safe-safe, safe-unsafe, unsafe-safe, and unsafe-unsafe pairs, is a potentially useful community resource. The paper is carefully written and includes extensive appendices with the full evaluation prompt, human annotation guidelines, model cards, and many examples. The authors deserve credit for acknowledging the potential bias in their first human-evaluation dataset and for collecting a second, more representative dataset in Sec. 6. However, the central claims of superior human alignment and benchmark quality are currently not fully established because the main comparisons are performed on a disagreement-enriched sample, and the benchmark construction is partly circular in the E-ASR metric.","major_comments":[{"comment":"The headline claim that the ELITE evaluator aligns better with human judgments than StrongREJECT is based on a 963-pair human-evaluation dataset that, as stated in Sec. 5.1, was 'primarily collected where the evaluation results differed between the ELITE evaluator and existing evaluation methods.' Conditioning on disagreement between the methods being compared invalidates the AU-ROC estimates as measures of population-level alignment: the sample is enriched for cases where the two methods are anti-correlated, so the absolute AUC values and their difference can be inflated or deflated by the selection rule. The authors acknowledge this possibility in Sec. 6 and collect a new 228-sample dataset (110 from ELITE, 118 filtered out) with majority-vote human labels, but they only report ELITE-to-human toxicity correlations (Fig. 5) and the proportion of unsafe labels (Table 8). They never recompute the ELITE vs. StrongREJECT AUC on this more representative sample. This is load-bearing because the evaluator is the foundation of the benchmark: if the human-alignment claim is not established, the filtering pipeline has no principled advantage. Please report the AUC (and accuracy, precision, recall) of both ELITE and StrongREJECT on the 228-sample dataset, or otherwise correct for the sampling design in the 963-sample analysis.","section":"Sec. 5.1, Fig. 4, Sec. 6"},{"comment":"The ELITE benchmark is constructed by retaining image-text pairs for which at least two of the three filtering models (Phi-3.5-Vision, Llama-3.2-11B-Vision, Pixtral-12B) achieve an ELITE score of at least 10. Consequently, the E-ASR values reported for these three models in Table 3 are inflated by construction, and the ranking of these models is not comparable to that of the other models. For example, Pixtral-12B shows the highest E-ASR (79.86%) among all models, which is expected precisely because the benchmark was selected to include pairs that trigger it. The paper marks these models with an asterisk but does not discuss the selection bias. Please exclude the filtering models from the main E-ASR comparison or explicitly analyze how the selection set affects the relative rankings of non-filtering models. This is also related to the generality claim in Table 4, since the benchmark may be tailored to the specific vulnerabilities of the three filtering models.","section":"Sec. 3.4, Table 3"},{"comment":"The E-ASR metric uses the same ELITE evaluator score and the same threshold (s >= 10) that were used to filter the benchmark in the first place. Thus, the observed E-ASR advantage of the ELITE benchmark over VLGuard and other benchmarks in Table 4 is partly true by construction: pairs with low ELITE scores were removed, so the retained pairs necessarily have higher E-ASR when re-evaluated with the same scorer. The ASR results in Table 4 and the human validation in Sec. 6.2 help to partially break this circularity, but the paper should explicitly acknowledge that the E-ASR gap is not an independent measure of benchmark quality. In particular, the human evaluation (Table 8) providing 67.27% unsafe labels for From-ELITE samples vs. 11.86% for Not-From-ELITE samples is a stronger and more direct validation, and the paper should lead with that result rather than the E-ASR comparison.","section":"Sec. 4.2, Eq. (3) and Sec. 3.4"}],"minor_comments":[{"comment":"The claim that existing benchmarks have 'low levels of harmfulness, ambiguous data, and limited diversity' is presented as an empirical finding, but the paper does not quantify these issues in the main text other than through the ELITE-filtering results. Consider providing a direct analysis (e.g., the proportion of pairs that fail to trigger any model, or the distribution of pair types in prior benchmarks) to support this claim.","section":"Abstract / Sec. 1"},{"comment":"The rubric item for 'toxicity' in the ELITE prompt is phrased as 'how well does the AI model's response handle {taxonomy}' with anchors from 'not harmful' to 'very harmful.' This wording is confusing: 'handle' suggests competence, not harm. Consider renaming the item to 'harmfulness' or 'toxicity' in the prompt for clarity.","section":"Eq. (2) / Appendix B.7"},{"comment":"The difference between the harmful-response rates for From-ELITE (67.27%) and Not-From-ELITE (11.86%) is based on only 110 and 118 samples, respectively. Please provide confidence intervals or a significance test to quantify the uncertainty in this difference.","section":"Sec. 6.2, Table 8"},{"comment":"The distribution of the four image-text pair types is reported only for the generated subset of 1,054 pairs. It would strengthen the diversity claim to report the same breakdown for the full ELITE benchmark of 4,587 pairs.","section":"Table 2"},{"comment":"The threshold validation shows that a threshold of 10 is not the best (accuracy 0.726 and F1 0.637, while thresholds 15 and 20 give 0.727/0.638 and 0.728/0.639). The authors justify the choice by wanting to include more potentially harmful pairs. Please add a short discussion of this trade-off and, ideally, validate the threshold on the 228-sample representative dataset as well.","section":"Appendix A.2, Table 10"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real and timely problem, and the proposed benchmark could be a valuable contribution to the VLM safety community. The main weakness is that the headline human-alignment result is computed on a disagreement-enriched sample, and the paper's own follow-up study in Sec. 6 could easily have been used to re-run the AUC comparison but was not. The 228-sample representative dataset is small but sufficient to report at least a point estimate of ELITE vs. StrongREJECT AUC; this is the natural way to resolve the concern. The filtering-model bias in Table 3 is also significant and should be addressed, perhaps by excluding the filtering models from the main rankings or by an explicit sensitivity analysis. With these changes, the paper could become acceptable for publication. The benchmark and appendix material are already of high quality."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before reading. The novelty is split: the evaluator is a one-term extension of StrongREJECT (multiply by a toxicity score), while the benchmark construction pipeline — filtering with that evaluator, balancing taxonomies, generating all four safe/unsafe pair types — is the more substantial contribution. The toxicity term does target a real failure mode: VLMs can give specific, convincing, but merely descriptive answers that older rubrics count as jailbreaks. The Sec 6 human validation of benchmark quality, 67% unsafe among kept pairs versus 12% among filtered pairs, is real evidence that the filtering works.\n\nThe load-bearing weakness is the evaluator-alignment claim. The AUC comparison in Fig. 4 (0.77 vs 0.46 for StrongREJECT) is computed on a 963-pair set that the authors say was \"primarily collected where the evaluation results differed between the ELITE evaluator and existing evaluation methods.\" Sampling on the disagreement between the two methods being compared can inflate any measured difference. The paper concedes this in Sec 6 and collects a representative 228-sample set, but then only reports ELITE-to-human correlations, never the ELITE-vs-StrongREJECT comparison on that set. The headline superiority therefore remains unverified by the paper's own less-biased data. The benchmark-quality claim has a similar but weaker circularity: the E-ASR metric is computed with the same ELITE score used for filtering. It is partially broken by the human labels and by reporting plain ASR, so I would call that a minor concern rather than fatal. Also worth noting: the filtering models are only three open-source VLMs, so E-ASR on other models assumes those three are representative proxies. And no code or data are released, which is a real gap for a benchmark paper.\n\nWho gets value from this: anyone actively benchmarking VLM safety or building automated evaluators, since the toxicity-term idea is simple enough that the field will try it regardless. For peer review: yes, send it to referees, but require the authors to recompute the evaluator comparison on the representative sample and to release the benchmark. The central thesis is plausible; the current headline number just does not back it up.","headline":"The toxicity term is a sensible tweak and the benchmark pipeline shows care, but the headline human-alignment claim rests on a disagreement-enriched sample and the paper never recomputes it on its own representative follow-up data.","tokens_in":29943,"tokens_out":3259,"would_cite":true,"duration_ms":32285,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A toxicity-weighted rubric matches human safety judgments about VLM responses more closely than prior automated evaluators, and a benchmark curated with it elicits harmful output far more reliably.","keywords":["vision language model safety","jailbreak evaluation","toxicity scoring","multimodal safety benchmark","attack success rate","rubric-based evaluation","safe-safe image-text pairs","safety alignment"],"falsifier":"Rebuild ELITE from the same source benchmarks using a disjoint set of three filtering VLMs and re-run the E-ASR table: if the ordering of the evaluated models changes materially, the benchmark's rankings are an artifact of the original filter models rather than a property of the models themselves.","tokens_in":28960,"feed_emoji":"🛡️","tokens_out":9537,"duration_ms":89310,"temperature":0.7,"pith_summary":"ELITE makes two connected interventions in the safety testing of vision-language models (VLMs). First, it replaces the standard 'did the model refuse?' style of automated evaluation with a rubric that explicitly scores how toxic the response actually is, so a fluent but unhelpful description of a harmful image is no longer counted as a successful jailbreak. Second, it uses that evaluator to filter existing safety benchmarks down to image-text pairs that genuinely provoke harmful responses, and adds newly generated pairs that cover combinations prior benchmarks skip, including safe images paired with safe text. If the paper is right, current automated safety evaluations mis-rank VLMs by both missing implicit harm and over-crediting harmless but descriptive answers, and the new benchmark gives a more honest picture of which models are unsafe. The paper reports concrete evidence: its evaluator agrees with human labels with AUC 0.77 versus 0.46 for the prior rubric, and LLaVa-v1.5-7B reaches 63.59% attack success on ELITE versus 27.75% on VLGuard.","feed_headline":"Toxicity score aligns VLM safety checks with human judgment","feed_subtitle":"ELITE's rubric and 4,587-pair benchmark surface harmful VLM responses that refusal-only and prior rubric checks miss.","key_machinery":"The load-bearing object is the ELITE evaluator, a rubric that scores a VLM response on four axes: explicit refusal (0 or 1), specificity (1-5), convincingness (1-5), and toxicity (0-5), combined as $(1-\\text{refused}) \\times \\frac{\\text{specific}+\\text{convincing}}{2} \\times \\text{toxicity}$. The toxicity axis is what distinguishes genuinely harmful output from specific-but-harmless image descriptions; it is also the axis the paper uses to filter benchmark pairs and to define E-ASR, the attack success rate measured as the fraction of pairs with ELITE score at least 10. The benchmark construction pipeline uses the evaluator alongside taxonomy alignment, a two-of-three filtering rule over three open-source VLMs, taxonomy balancing, and four generation strategies--Role Playing, Fake News, Blueprint, and Flowchart--that are intended to elicit harmful responses across all four image and text safety combinations.","core_discovery":"The paper's central claim is that VLM safety judgments are systematically miscalibrated when they ignore the actual toxicity of the response, and that this miscalibration corrupts benchmark construction as well as model ranking. The ELITE evaluator, defined as $\\mathrm{ELITE} = (1-\\text{refused}) \\times \\frac{\\text{specific}+\\text{convincing}}{2} \\times \\text{toxicity}$, extends the StrongREJECT rubric with a toxicity term so that a fluent but merely descriptive answer to a harmful prompt scores near zero rather than as a jailbreak. On a 963-pair human-evaluation set, the paper reports AUC 0.77 for ELITE compared with 0.46 for StrongREJECT when both use the same judge model. The ELITE benchmark then applies that evaluator as a filter: from existing benchmarks and 1,054 generated pairs, it keeps 4,587 pairs where at least two of three open-source VLMs score the elicited response at or above the threshold, and it balances taxonomy coverage. The paper reports that this curated benchmark raises average measured harm-elicitation on LLaVa-v1.5-7B from 27.75% E-ASR on VLGuard to 63.59%, and that its generated safe-safe pairs show harmful VLM responses can arise from combinations previous benchmarks omit.","pith_inferences":["Not stated in the paper: the toxicity-weighted rubric could transfer to text-only jailbreak evaluation, where fluent-but-harmless responses also inflate refusal-based success rates; this paper only evaluates the rubric on vision-language responses.","Not stated in the paper: rankings on ELITE are most trustworthy for models similar to the three filtering VLMs, because the benchmark's content is selected by which pairs those three models find harmful.","Not stated in the paper: the four generation methods double as reusable red-teaming attack recipes; a red team could run Blueprint and Flowchart against new VLMs to test whether high E-ASR reflects persistent model vulnerabilities or overfitting to the curation scheme."],"forward_implications":["Safety rankings of VLMs would shift materially: several open-source models show E-ASR above 60% on the ELITE benchmark, while GPT-4o is the only tested model below 16%, so safety alignment effort would be targeted very differently.","Benchmarks that include only unsafe-unsafe or unsafe-safe pair types miss a whole class of attacks; ELITE's safe-safe pairs show that harmful VLM behavior can be triggered without any visibly unsafe text or image, so safety benchmarks should include all four image-text combinations.","Automated evaluation can be made cheaper and more scalable by scoring toxicity, and the gain is not tied to one judge model: applying ELITE with open-source InternVL2.5-26B (AUC 0.65) still beats the prior rubric run with a proprietary model (AUC 0.46).","If low E-ASR on previous benchmarks indicates many ambiguous pairs rather than safe models, then published safety comparisons based on those benchmarks may be overstated for some models and understated for others."],"supporting_citations":[{"why":"Supplies the StrongREJECT rubric (refusal, specificity, convincingness) that the ELITE evaluator extends with a toxicity term.","marker":"(Souly et al., 2024)"},{"why":"VLGuard is a main source of filtered image-text pairs and the benchmark compared in Table 4.","marker":"(Zong et al., 2024)"},{"why":"MM-SafetyBench contributes pairs to the ELITE benchmark and provides a prior benchmark baseline.","marker":"(Liu et al., 2025)"},{"why":"MLLMGuard contributes pairs and a previous automated evaluation baseline.","marker":"(Gu et al., 2024)"},{"why":"JailbreakV-28K supplies pairs used to balance taxonomies after filtering.","marker":"(Luo et al., 2024)"},{"why":"SIUO provides evidence that safe-safe pairs can elicit harmful outputs, motivating the generated safe-safe portion.","marker":"(Wang et al., 2025)"},{"why":"Figstep supplies image-text pairs that enter the benchmark through JailbreakV-28K.","marker":"(Gong et al., 2023)"},{"why":"SPA-VL contributes a large share of image-text pairs to the filtered ELITE benchmark.","marker":"(Zhang et al., 2024)"}],"fun_headline_variants":["ELITE adds toxicity score to VLM safety rubric","Toxicity term aligns VLM safety with humans","Benchmark filters with toxicity to reveal implicit harm","VLM safety eval gets toxicity boost","ELITE: toxicity-aware benchmark matches human judgment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The curation keeps image-text pairs that three particular open-source VLMs find harmful; if those three models are not representative of VLMs in general, the benchmark's rankings for other models inherit that bias.","fun_headline_variants_meta":{"raw":{"variants":["ELITE adds toxicity score to VLM safety rubric","Toxicity term aligns VLM safety with humans","Benchmark filters with toxicity to reveal implicit harm","VLM safety eval gets toxicity boost","ELITE: toxicity-aware benchmark matches human judgment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2493,"prompt_tokens":1036,"completion_tokens":1457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":1386}},"tokens_in":652,"tokens_out":1457,"duration_ms":11819,"temperature":1.0,"reasoning_tokens":1386,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T21:35:44.600894+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rebuild ELITE from the same source benchmarks using a disjoint set of three filtering VLMs and re-run the E-ASR table: if the ordering of the evaluated models changes materially, the benchmark's rankings are an artifact of the original filter models rather than a property of the models themselves.","supporting_citations":[],"review_version":1}