{"id":"2019228f-d605-47da-a257-ff040c3f88e8","arxiv_id":"2505.14728","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MORALISE is a new benchmark of 2,481 expert-annotated real image-text pairs spanning 13 moral topics, and 19 vision-language models score far worse on identifying the violated norm than on judging whether a violation occurred.","lead":"This paper introduces MORALISE, a benchmark of 2,481 real-world image-text pairs that tests whether vision-language models can detect moral violations and name the violated norm. The authors evaluate 19 models and find that binary moral judgment is often strong, but fine-grained norm attribution is much harder, especially from images alone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Label reliability is unquantified: majority-vote annotations by ML graduate students, no inter-annotator agreement or human performance baseline, so low model scores could reflect label or taxonomy ambiguity rather than moral deficiency.","rationale":"The reader's weakest assumption identifies exactly the load-bearing point: the benchmark's ground-truth topic and modality labels are treated as correct based on an unquantified majority-vote process, with no inter-annotator agreement, external validation, or human performance baseline. My reading of the full text confirms this is the place where the central claim is least secure. The headline results are all computed relative to these gold labels, so label noise or taxonomy overlap directly threatens the conclusion that models have persistent moral limitations rather than merely that the benchmark is hard for everyone. This is not an internal inconsistency; the dataset is public, the evaluation protocol is mostly specified, and the balanced statistics in Appendix A are useful. The concern is a missing validation step, not a demonstrated flaw. Therefore the appropriate verdict remains CONDITIONAL: the resource is valuable and the observed trends are plausible, but the central claim should be accepted only after label reliability and a human baseline are reported. Since this matches the reader's conditional verdict, no change is needed.","tokens_in":26817,"tokens_out":4147,"duration_ms":42105,"concrete_test":"Recruit independent annotators (ideally with moral-psychology expertise), not including the authors, to re-annotate a stratified random sample of at least 200 MORALISE examples balanced across the 13 topics and both modality types using the published definitions. Compute Fleiss' kappa for topic labels and modality labels, and measure human performance on the S2/S3 prompts under the identical response format. If human-annotator agreement is low (kappa below 0.6) or human S3 F1 on 'Respect' is close to GPT-4o's 42.32, the labels are too ambiguous to support the claimed moral limitation; if agreement is high and human F1 is substantially higher, the central claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that MORALISE reveals persistent moral limitations and that fine-grained norm attribution is much harder than binary moral judgment—rests on treating the gold labels as ground truth. Section 3.2 describes a majority-vote protocol among graduate students in ML-related fields, but the paper reports no inter-annotator agreement, no adjudication statistics, no exclusion counts, and no human performance baseline. This matters most for multi-norm attribution (S3) because several topic definitions overlap (Justice vs Fairness vs Discrimination; Authority vs Responsibility vs Respect; Liberty vs Respect), so a model predicting a semantically adjacent label is scored as wrong solely because the annotator majority chose a different adjacent label. Reported evidence such as GPT-4o's 42.32 F1 on Respect (Table 4) and the gap between 88.28/83.55 binary judgment accuracy and 66.60/42.63 attribution hit rates (Tables 2 and 3) could therefore reflect ambiguous labels and arbitrary tie-breaking, not a genuine moral deficiency. The limitation statement in Appendix E acknowledges the pipeline is labor-intensive and not scalable but does not quantify label noise. Without a human baseline on the same task and prompt, a 'significant challenge' score cannot be distinguished from a poorly calibrated task.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MORALISE, a benchmark for evaluating moral alignment in vision-language models (VLMs). It proposes a taxonomy of 13 moral topics organized under Turiel's Domain Theory (personal, interpersonal, societal), curates 2,481 real-world image-text pairs with topic and modality (text-centric vs. image-centric) annotations, and defines two tasks: moral judgment (binary) and moral norm attribution (single- and multi-label). The authors evaluate 19 open and proprietary VLMs across three subtasks (S1 binary accuracy, S2 single-norm hit rate, S3 multi-norm F1) and report large performance drops from binary judgment to fine-grained attribution, along with analyses of model scale, family, modality sensitivity, and prediction correlation. The benchmark is publicly released on Hugging Face.","tokens_in":27022,"tokens_out":5046,"duration_ms":47046,"significance":"If the benchmark labels are reliable, MORALISE would be a valuable resource: it is one of the first VLM moral-alignment benchmarks to use real-world images rather than AI-generated ones, covers a relatively broad taxonomy, provides modality-centric annotations that enable isolating visual vs. textual moral cues, and evaluates a substantial set of models. The finding that proprietary and open-source models retain high binary judgment accuracy but drop sharply on norm attribution (e.g., proprietary average 88.28 vs. 66.60, open-source 83.55 vs. 42.63) is a potentially important and actionable observation for alignment research. The paper also commendably provides the full dataset and detailed prompts for reproducibility. However, the current manuscript does not establish the reliability of its ground-truth labels, so these claims are not yet fully supported.","major_comments":[{"comment":"","section":"§3.2, Appendix E"},{"comment":"","section":"§3.1, Tables 3–4"},{"comment":"","section":"§4.2, Tables 2–4"},{"comment":"","section":"§3.2, §5"}],"minor_comments":[{"comment":"","section":"References"},{"comment":"","section":"Tables 5–7"},{"comment":"","section":"§3.2, Appendix A"},{"comment":"","section":"Appendix B.1, Prompt τS3"},{"comment":"","section":"§4.4, Figure 6"},{"comment":"","section":"§1, §4.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is potentially valuable and the dataset release is a plus, but as a benchmark paper it cannot stand without inter-annotator agreement statistics, a human performance baseline, and a chance-level baseline, because the central 'significant challenge' claim depends on label reliability. The overlapping taxonomy definitions compound this problem for the fine-grained tasks. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection. The reviewer's take that 'MORALISE poses a significant challenge' is plausible but under-evidenced as written; the quantitative claims need error bars or significance tests as well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a useful benchmark resource, and the headline result — fine-grained norm attribution is much harder for VLMs than binary moral judgment — is credible. Second: the label reliability is unquantified, and nobody should take the exact numbers as ground truth until the authors add annotation agreement and a human baseline.\n\nWhat is actually new: MORALISE combines real-world images, a 13-topic moral taxonomy from Turiel's domain theory, and per-sample text-centric vs image-centric violation labels across 2,481 image-text pairs. That combination doesn't exist in M3oralBench or VIVA. The dataset is publicly released, the topic balance looks careful, and the authors evaluate 19 open- and proprietary models with a mostly reproducible protocol (prompts are in the appendix). The modality split supports the paper's claim that image-based moral reasoning lags text-based reasoning across all three tasks. That is a solid, useful piece of evidence.\n\nSoft spots: no inter-annotator agreement, no adjudication statistics, no exclusion counts, no human performance baseline, no chance baseline. Labels come from majority vote among ML graduate students, and the taxonomy has real overlap — Justice vs Fairness, Authority vs Responsibility, Liberty vs Respect. A model picking a semantically adjacent label gets marked wrong solely because the annotators chose a different adjacent label. So the reported gaps (88.28 vs 66.60 for proprietary models on judgment vs hit rate, or GPT-4o's 42.32 F1 on Respect) could partly reflect label ambiguity, not just model deficiency. The authors acknowledge in Appendix E that the curation pipeline is labor-intensive, but they don't quantify noise. Results are point estimates without error bars or significance tests, and no evaluation code is released.\n\nThese are fixable issues, not a fatal flaw. The central direction holds: attribution is harder than binary judgment, and visual moral reasoning is weaker than textual. I'd ask the authors for IAA, a human baseline on the same prompts, and code release; those additions would make the benchmark considerably stronger.\n\nThis paper is for people building or auditing VLMs in safety-sensitive settings. I'd bring it to a reading group and would cite it as a resource. Send it to peer review; just require the missing reliability measurements before accepting.","headline":"A useful real-image moral alignment benchmark whose headline numbers need an annotation-agreement and human-baseline check before being read as ground truth.","tokens_in":27583,"tokens_out":2583,"would_cite":true,"duration_ms":25429,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark of 2,481 real-world image-text pairs shows that vision-language models can usually tell right from wrong but often fail to name which of 13 moral norms a scenario violates.","keywords":["moral alignment","vision-language models","multimodal benchmark","moral norm attribution","moral judgment","Turiel's domain theory","real-world image-text data","modality annotation"],"falsifier":"Re-annotate the 2,481 image-text pairs with a second independent panel and compute per-topic and per-modality inter-annotator agreement, then run the same 19 models alongside a human baseline on the norm-attribution task. If agreement falls below conventional thresholds, such as Cohen's kappa under 0.6 on topic labels, or if human participants score no higher than the best models on the fine-grained task, then the reported moral deficiencies would be better explained by label ambiguity or task design than by model incapacity.","tokens_in":26602,"feed_emoji":"⚖️","tokens_out":5380,"duration_ms":46827,"temperature":0.7,"pith_summary":"This paper introduces MORALISE, a benchmark that tests whether vision-language models can tell when an image-text scenario is morally wrong and, crucially, which of 13 moral norms it violates. The authors assemble 2,481 real-world image-text pairs, each manually labeled for the violated moral topic and for whether the violation is carried by the image or the text, then evaluate 19 open-source and proprietary models. They find that models are reasonably good at binary moral judgment, averaging 88.28 accuracy for proprietary models and 83.55 for open-source models, but much worse at naming the violated norm, with average hit rates of 66.60 and 42.63 respectively. The central claim is that moral alignment in multimodal settings remains an open problem, and that fine-grained moral norm attribution—not just coarse judgment—is the binding constraint.","feed_headline":"New benchmark finds VLMs fail to name violated moral norms","feed_subtitle":"Binary judgment holds up, but naming the violated norm drops to 67% for proprietary and 43% for open-source models.","key_machinery":"The load-bearing mechanism is the pairing of a 13-topic moral taxonomy with modality-centric annotation. Each sample is labeled with one or more of 13 moral topics, organized into personal, interpersonal, and societal domains following Turiel's domain theory, and with a binary modality label indicating whether the moral violation is conveyed primarily through the image or through the text. This supports two evaluation tasks: moral judgment, a binary wrong/not-wrong decision, and moral norm attribution, scored as single-norm hit rate and multi-norm F1 over the 13 topics. The taxonomy gives the benchmark its fine-grained diagnostic power, while the modality label isolates where models fail.","core_discovery":"On the authors' own terms, MORALISE establishes that current vision-language models exhibit a systematic gap between recognizing that something is morally wrong and identifying which moral norm is violated. Using a taxonomy of 13 moral topics rooted in Turiel's domain theory, spanning personal, interpersonal, and societal domains, the authors curated 2,481 expert-verified real-world image-text pairs with both topic and modality labels. Across 19 models, binary moral judgment accuracy averages 88.28 for proprietary and 83.55 for open-source models, but single-norm attribution hit rates drop to 66.60 and 42.63 respectively, and multi-label F1 scores are lower still. Even the strongest evaluated model, GPT-4o, reaches only 42.32 F1 on the 'respect' topic, and the paper reads this as evidence that fine-grained moral reasoning is a distinct and largely unsolved capability.","pith_inferences":["If label noise is a real driver, then reporting human performance on the same 2,481 pairs would calibrate the benchmark: a human baseline near or below GPT-4o's 42 F1 on 'respect' would suggest the task is ambiguous rather than the model deficient.","The text/image gap points to a concrete testable extension: pipelines that first caption the image and then judge morality from the caption should recover most of the text-centric advantage, directly testing whether image understanding or moral reasoning is the bottleneck.","A fixed 13-topic taxonomy may encode culturally specific moral assumptions, so a comparative study with annotators from different cultural backgrounds could reveal whether the norm-attribution failures partly reflect a cultural mismatch rather than a general moral deficit."],"forward_implications":["Fine-grained moral norm attribution is substantially harder for vision-language models than binary moral judgment: the average gap is roughly 22 accuracy points for proprietary models and 41 for open-source models.","Text-centric moral violations are consistently easier than image-centric ones across all three subtasks, indicating that current VLMs lean on language rather than visual content when reasoning about morality.","Scaling model size helps moral performance up to about 10 billion parameters and then plateaus, so larger models alone will not close the gap without targeted moral-alignment training.","Norms that are common in social discourse, such as harm, justice, and integrity, are handled relatively well, while abstract norms such as liberty, respect, and reciprocity remain persistent weak points, especially in multi-label attribution."],"supporting_citations":[{"why":"Supplies Turiel's Domain Theory, the psychological foundation for the paper's three moral domains and 13-topic taxonomy.","marker":"[44]"},{"why":"ETHICS is the text-only moral benchmark whose task framing the paper extends to multimodal settings.","marker":"[12]"},{"why":"M3oralBench is the main multimodal moral benchmark being contrasted, since it relies on AI-generated images rather than real-world ones.","marker":"[46]"},{"why":"VIVA is a real-world moral VLM benchmark that lacks the 13-topic granularity and modality-violation annotation introduced here.","marker":"[13]"},{"why":"MoralBench is a text-only LLM moral benchmark used in the positioning table to show the gap in VLM-specific evaluation.","marker":"[14]"}],"fun_headline_variants":["VLMs judge moral wrongs but can't name the norm","New benchmark: VLMs see wrong but not why","Moral blind spot: VLMs fail to name violated norms","Benchmark shows VLMs lack fine-grained moral reasoning","MORALISE: VLMs judge but can't explain moral violations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ground-truth labels, which say which scenarios are morally wrong, which of 13 topics they violate, and whether the violation is image-centric or text-centric, are taken as correct on the strength of a majority vote among graduate student annotators, with no reported inter-annotator agreement, external validation of the taxonomy, or human performance baseline.","fun_headline_variants_meta":{"raw":{"variants":["VLMs judge moral wrongs but can't name the norm","New benchmark: VLMs see wrong but not why","Moral blind spot: VLMs fail to name violated norms","Benchmark shows VLMs lack fine-grained moral reasoning","MORALISE: VLMs judge but can't explain moral violations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1588,"prompt_tokens":1042,"completion_tokens":546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":462}},"tokens_in":658,"tokens_out":546,"duration_ms":5449,"temperature":1.0,"reasoning_tokens":462,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:09:13.632919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the 2,481 image-text pairs with a second independent panel and compute per-topic and per-modality inter-annotator agreement, then run the same 19 models alongside a human baseline on the norm-attribution task. If agreement falls below conventional thresholds, such as Cohen's kappa under 0.6 on topic labels, or if human participants score no higher than the best models on the fine-grained task, then the reported moral deficiencies would be better explained by label ambiguity or task design than by model incapacity.","supporting_citations":[{"cited_title":"Cambridge University Press, 1983","cited_arxiv_id":null,"evidence_quote":"Supplies Turiel's Domain Theory, the psychological foundation for the paper's three moral domains and 13-topic taxonomy."},{"cited_title":"VIV A: A benchmark for vision-grounded decision- making with human values","cited_arxiv_id":null,"evidence_quote":"VIVA is a real-world moral VLM benchmark that lacks the 13-topic granularity and modality-violation annotation introduced here."}],"review_version":1}