REVIEW 2 major objections 6 minor 1 cited by
JETHICS: Japanese Ethics Understanding Evaluation Dataset
T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read JETHICS, a 77,896-example Japanese ethics benchmark, finds GPT-4o reaches only 0.713 and the best Japanese LLM 0.497.
desk verdict Useful first Japanese ethics benchmark, but the Utilitarianism label noise and unreleased data mean the headline scores should be read as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the ETHICS-style construction pipeline: crowdworkers generate Japanese sentences and candidate labels, then three or four independent crowdworkers vote on each example-label pair, with split votes removed. On top of this, the evaluation uses category-specific Japanese prompts with eight in-context examples, and for multi-sentence categories a model is correct only if it classifies every continuation sentence correctly. The normative-theory categories — utilitarianism, deontology split into role and request, virtue ethics, and justice split into impartiality and desert — supply the task definitions that make the benchmark a test of moral theory understanding rather than general language ability.
What would settle it
Take a random sample of JETHICS examples, re-annotate them with a larger and demographically diverse panel of Japanese speakers, and compare labels and model rankings; if agreement drops substantially or the ranking of models changes, the benchmark's claimed validity would be undercut. A cheaper check targets the utilitarianism category: a second annotation round with the same guidelines should reproduce near-zero kappa and unstable labels, which would disqualify that sub-benchmark.
Extended reading notes
Core claim
The central claim is that JETHICS is a working Japanese counterpart to ETHICS: it follows the same construction recipe, covers the same normative-theory categories, and produces annotator agreement (average kappa 0.61) that the authors read as acceptable. Under this benchmark, current models fall short: GPT-4o averages 0.713, with a weak 0.445 on virtue ethics, and the strongest evaluated Japanese LLM averages 0.497. The paper also finds that larger models and extra Japanese instruction tuning improve scores, interpreting this as evidence that Japanese-specific training helps but that substantial work remains before models reliably understand Japanese moral norms. The authors flag the utilitarianism category's low annotator agreement (kappa 0.18) as a caveat while arguing that excluding split votes preserves dataset quality.
Load-bearing premise
The load-bearing premise is that the majority labels of three or four crowdworkers represent Japanese moral values, including in the utilitarianism category where those workers barely agreed (kappa 0.18); if those labels are noisy or unrepresentative, the model scores do not measure ethics understanding.
Editorial extensions
If this is right
- GPT-4o's 0.713 average and its 0.445 on virtue ethics indicate advanced models have substantial room to improve in Japanese moral understanding.
- The best evaluated Japanese LLM, llm-jp-3-13b, reaches only 0.497, so no current model comes close to saturating the benchmark.
- Scaling from llm-jp-3-3.7b to llm-jp-3-13b raised average accuracy by 0.171, and adding Japanese pre-training and instruction tuning (LlamaELYZA8b vs MetaLlama8b) raised it by 0.112, pointing to model size and Japanese-specific training as levers.
- Cultural items such as the Anpanman graduation song show that Japanese-specific norms are part of what the benchmark measures.
- JETHICS provides a non-Western resource for studying moral understanding, complementing English-centric datasets.
Reading between the lines
- If the benchmark is accepted, it becomes a test bed for cross-cultural ethics: the same prompts translated into other languages could expose where model moral judgments are Western-biased.
- The low kappa (0.18) in the utilitarianism category suggests that category should be interpreted cautiously; a model's score there may reflect which subjective notion of well-being its training data happened to encode rather than a stable moral fact.
- A natural next step is to use JETHICS for fine-tuning or preference tuning of Japanese LLMs and then re-test on a held-out split; the paper reports evaluation only, not training.
- The examples from Table 4 could be turned into a targeted diagnostic of culturally specific norms if each item were tagged with the cultural rule it relies on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents JETHICS, a Japanese-language dataset for evaluating the moral understanding of large language models. The dataset contains 77,896 examples across five categories (Utilitarianism, Deontology with Role and Request subcategories, Justice with Desert and Impartiality subcategories, Virtue Ethics, and Commonsense Morality), constructed by following the methodology of the English ETHICS dataset. Crowdworkers generated example-label pairs, and a separate set of crowdworkers validated them via majority vote. The authors evaluate four non-proprietary LLMs (llm-jp 3.7B/13B, Meta-Llama-3-8B, and Llama-3-ELYZA-JP-8B) and two GPT-4o variants in an 8-shot setting, reporting that GPT-4o achieves an average accuracy of 0.713 and the best Japanese model 0.497, which they interpret as evidence that current models have substantial room for improvement in Japanese ethics understanding.
Significance. If the dataset's labels are reliable, JETHICS fills a clear gap: most existing ethics benchmarks are English-centric, and the paper provides a large-scale, non-Western resource with a datasheet and a public-release plan. The evaluation includes a meaningful comparison between a Japanese-tuned model and its base, and the paper is transparent about reporting inter-annotator agreement. The main significance is therefore conditional on the validity of the ground truth, which is questionable for the Utilitarianism category.
major comments (2)
- [Section 2.1, Table 2, Section 5] The reported inter-annotator kappa of 0.18 for the Utilitarianism category (19,529 examples, 25% of the dataset) is a load-bearing validity concern, and the paper's defense in Section 5 is insufficient. Please clarify whether the kappa in Table 2 is computed on all step-2 annotations or on the filtered set after excluding split evaluations. If it is computed after exclusion, the kappa is inflated by construction because disagreements are removed; if it is computed before, the filtering removes only exact ties, leaving many examples decided by a 2-1 or 3-1 majority whose reliability is questionable given the low kappa. For a binary task with kappa 0.18, the implied per-annotator agreement is about 0.59, and a majority of three such annotators is only about 0.63 accurate. The paper needs to provide independent validation of the Utilitarianism labels, for example a second annotation pass with a larger panel or expert adjudication, or the category should be reported separately rather than being averaged into the headline results.
- [Section 5 and Ethics Statement] The claim that 'since examples with significant disagreement were excluded from JETHICS, the overall quality of the dataset is not compromised' is not supported by the evidence. Excluding split votes does not address the underlying ambiguity that causes low kappa; it merely removes the most extreme disagreements. The Ethics Statement itself acknowledges that 'it is unclear whether the labeled actions are actually morally right since only three or four people annotated a example,' which directly undermines the dataset's validity claim. To make the central claim that JETHICS is a valid Japanese ethics benchmark, the authors should either re-annotate the Utilitarianism category with a larger number of workers, provide external evidence of label validity, or exclude this category from the overall average with a clear caveat.
minor comments (6)
- [Section 2.1] The number of annotators per example is not reported; please state the distribution of three versus four annotators and define precisely what 'split evaluations' means for each case (e.g., 1-1-1 with three annotators, 2-2 with four annotators).
- [Section 5] The sentence 'the overall quality of the dataset is not compromised' is a claim that needs direct evidence; please provide an analysis of the filtered set, such as the kappa computed on the remaining examples or a comparison of model scores on high-agreement versus low-agreement subsets.
- [Datasheet, Distribution section] The datasheet states 'The dataset will be released after the paper is accepted,' while the abstract and introduction state that the dataset is 'released'; please align these statements so that the availability status is unambiguous.
- [Appendix B, Table 5] The prompt format appears to contain a duplicated '応答(Response):' line; please verify the actual prompt used in the experiments to avoid confusion.
- [Section 2.2, footnote 5] The Commonsense Morality category is based on the authors' earlier Takeshita et al. (2023) dataset; this should be stated in the main text of Section 2.2 so that readers can assess the degree of independence of this subcategory from prior work.
- [Table 3] The random baseline for the multi-sentence categories (0.063) is not derived in the text; please specify the number of following sentences per category (e.g., four binary sentences yield 1/16) so that the baseline is verifiable.
Circularity Check
No significant circularity: JETHICS labels are fixed crowdsourced inputs and model scores are independent measurements against those inputs.
full rationale
JETHICS's construction chain is not circular. The dataset is built by crowdsourcing sentence-label pairs and validating them by majority vote, and model accuracies are then computed against those fixed labels. Nothing in the paper fits a parameter to a subset of the evaluation data and then reports a closely related quantity as a prediction. The only self-referential element is the footnote that the Commonsense Morality category is based on Takeshita et al. (2023), but this is a data provenance statement, not a derivation: the category's labels predate the evaluation and are not constructed from model outputs. Similarly, following Hendrycks et al. (2021) for prompt format and evaluation protocol is methodological borrowing, not circularity. The low utilitarianism kappa (0.18) and the exclusion of split votes raise label-quality and validity concerns, but those are correctness risks rather than circular steps: the majority-vote labels are inputs, and the model scores are independent measurements against those inputs. No equation, fitted parameter, or self-citation chain forces the reported averages, so no significant circularity is present.
Assumptions & free parameters
free parameters (2)
- Evaluation sample size =
1,000 examples per category
- Few-shot count =
8
assumptions (4)
- domain assumption Crowdworker majority labels (three or four annotators, split votes removed) are a valid ground truth for Japanese moral judgments.
- domain assumption The hand-written category instructions and 8-shot prompts measure the intended ethical construct rather than test-format artifacts.
- domain assumption Randomly selecting 1,000 examples per category gives a representative sample of the full 77,896-example dataset for model comparison.
- domain assumption The crowdsourced templates faithfully operationalize utilitarianism, deontology, virtue ethics, and justice as described in Section 2.2.
Cite this review
Pith. "Pith review of JETHICS: Japanese Ethics Understanding Evaluation Dataset." pith.science (2026). https://pith.science/paper/6D2SRJRX
@misc{pith2026250616187,
author = {Pith},
title = {Pith review of: JETHICS: Japanese Ethics Understanding Evaluation Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/6D2SRJRX}},
note = {Machine review of arXiv:2506.16187}
}
read the original abstract
In this work, we propose JETHICS, a Japanese dataset for evaluating ethics understanding of AI models. JETHICS contains 78K examples and is built by following the construction methods of the existing English ETHICS dataset. It includes four categories based normative theories and concepts from ethics and political philosophy; and one representing commonsense morality. Our evaluation experiments on non-proprietary large language models (LLMs) and on GPT-4o reveal that even GPT-4o achieves only an average score of about 0.7, while the best-performing Japanese LLM attains around 0.5, indicating a relatively large room for improvement in current LLMs.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses
CEDAR is a 7-language, 2-modality benchmark of 10,962 culturally divergent emotion scenarios; 17 LLMs perform poorly, and prompt-language matching does not fix cultural misalignment.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Larry Alexander and Michael Moore. 2021. Deontological Ethics . In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy , W inter 2021 edition. Metaphysics Research Lab, Stanford University
work page 2021
-
[4]
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-Fran c ois Bonnefon, and Iyad Rahwan. 2018. The moral machine experiment. Nature, 563(7729):59--64
work page 2018
-
[5]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, ...
arXiv 2022
-
[6]
David Bourget and David J Chalmers. 2023. Philosophers on philosophy: The 2020 philpapers survey. Philosophers' Imprint, 23
work page 2023
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
-
[8]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD : Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 862--872
work page 2021
Show all 25 references
-
[9]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The L lama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[10]
Hwang, Maxwell Forbes, and Yejin Choi
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.54 Moral S tories: Situated reasoning about norms, intents, actions, and their consequences . In Proceedings of the 2021 Conference on Empirical Method...
2021 doi
-
[11]
Fred Feldman and Brad Skow. 2020. Desert . In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy , W inter 2020 edition. Metaphysics Research Lab, Stanford University
2020
-
[12]
Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.48 Social chemistry 101: Learning to reason about social and moral norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Languag...
2020 doi
-
[13]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.301 R eal T oxicity P rompts: Evaluating neural toxic degeneration in language models . In Findings of the Association for Computational Linguist...
2020 doi
-
[14]
Jian Guan, Ziqi Liu, and Minlie Huang. 2022. https://doi.org/10.18653/v1/2022.naacl-main.374 A corpus for understanding and generating moral stories . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...
2022 doi
-
[15]
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI with shared human values. In International Conference on Learning Representations
2021
-
[16]
Rosalind Hursthouse. 1999. On Virtue Ethics. Oxford University Press, Oxford
1999
-
[17]
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O'Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, S...
2024 arXiv
-
[18]
LLM-jp. 2024. LLM -jp: A cross-organizational project for the research and development of fully open J apanese LLM s. arXiv preprint arXiv:2407.03963
2024 arXiv
-
[19]
John Rawls. 1999. A theory of justice: Revised edition. Harvard university press
1999
-
[20]
Ines Reinig, Maria Becker, Ines Rehbein, and Simone Ponzetto. 2024. https://doi.org/10.18653/v1/2024.findings-acl.245 A survey on modelling morality for text analysis . In Findings of the Association for Computational Linguistics: ACL 2024, pages 4136--4155, Bangkok, Thailand....
2024 doi
-
[21]
Andrew E Reisner. 2013. Prima facie and pro tanto oughts. International Encyclopedia of Ethics
2013
-
[22]
Sebastin Santy, Jenny Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. https://doi.org/10.18653/v1/2023.acl-long.505 NLP ositionality: Characterizing design biases of datasets and models . In Proceedings of the 61st Annual Meeting of the Association for Computa...
2023 doi
-
[23]
Masashi Takeshita, Rafal Rzpeka, and Kenji Araki. 2023. https://www.anlp.jp/proceedings/annual_meeting/2023/pdf_dir/D2-1.pdf Jcommonsensemorality: Japanese dataset for evaluating commonsense morality understanding . In In Proceedings of The Twenty Nineth Annual Meeting of The ...
2023
-
[24]
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.naacl-main.56 On the machine learning of ethical judgments from natural language . In Proceedings of the 2022 Conference of the North America...
2022 doi
-
[25]
Christopher Woodard. 2019. Taking utilitarianism seriously. Oxford University Press
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.