Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

JETHICS: Japanese Ethics Understanding Evaluation Dataset

T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read JETHICS, a 77,896-example Japanese ethics benchmark, finds GPT-4o reaches only 0.713 and the best Japanese LLM 0.497.

desk verdict Useful first Japanese ethics benchmark, but the Utilitarianism label noise and unreleased data mean the headline scores should be read as provisional. read the letter →

arxiv 2506.16187 v1 pith:6D2SRJRX submitted 2025-06-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords JapaneseethicsdatasetmoralunderstandingLLMevaluationnormativecommonsensemoralityculturalrelativityAIalignmentbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces JETHICS, a Japanese-language dataset of 77,896 moral examples built to mirror the English ETHICS benchmark. Its categories cover utilitarianism, deontology (role and request), virtue ethics, justice (impartiality and desert), and commonsense morality, with labels validated by crowdworker majority vote. The paper evaluates six LLMs and reports that GPT-4o reaches an average of 0.713, while the best Japanese LLM, llm-jp-3-13b, reaches 0.497, leaving substantial room for improvement. These results matter because existing ethics datasets are mostly Western, and cultural differences in moral judgments require non-Western benchmarks to test whether models understand values beyond the English-speaking context.

What carries the argument

The load-bearing object is the ETHICS-style construction pipeline: crowdworkers generate Japanese sentences and candidate labels, then three or four independent crowdworkers vote on each example-label pair, with split votes removed. On top of this, the evaluation uses category-specific Japanese prompts with eight in-context examples, and for multi-sentence categories a model is correct only if it classifies every continuation sentence correctly. The normative-theory categories — utilitarianism, deontology split into role and request, virtue ethics, and justice split into impartiality and desert — supply the task definitions that make the benchmark a test of moral theory understanding rather than general language ability.

What would settle it

Take a random sample of JETHICS examples, re-annotate them with a larger and demographically diverse panel of Japanese speakers, and compare labels and model rankings; if agreement drops substantially or the ranking of models changes, the benchmark's claimed validity would be undercut. A cheaper check targets the utilitarianism category: a second annotation round with the same guidelines should reproduce near-zero kappa and unstable labels, which would disqualify that sub-benchmark.

Watch

Extended reading notes

Core claim

The central claim is that JETHICS is a working Japanese counterpart to ETHICS: it follows the same construction recipe, covers the same normative-theory categories, and produces annotator agreement (average kappa 0.61) that the authors read as acceptable. Under this benchmark, current models fall short: GPT-4o averages 0.713, with a weak 0.445 on virtue ethics, and the strongest evaluated Japanese LLM averages 0.497. The paper also finds that larger models and extra Japanese instruction tuning improve scores, interpreting this as evidence that Japanese-specific training helps but that substantial work remains before models reliably understand Japanese moral norms. The authors flag the utilitarianism category's low annotator agreement (kappa 0.18) as a caveat while arguing that excluding split votes preserves dataset quality.

Load-bearing premise

The load-bearing premise is that the majority labels of three or four crowdworkers represent Japanese moral values, including in the utilitarianism category where those workers barely agreed (kappa 0.18); if those labels are noisy or unrepresentative, the model scores do not measure ethics understanding.

Editorial extensions

If this is right

  • GPT-4o's 0.713 average and its 0.445 on virtue ethics indicate advanced models have substantial room to improve in Japanese moral understanding.
  • The best evaluated Japanese LLM, llm-jp-3-13b, reaches only 0.497, so no current model comes close to saturating the benchmark.
  • Scaling from llm-jp-3-3.7b to llm-jp-3-13b raised average accuracy by 0.171, and adding Japanese pre-training and instruction tuning (LlamaELYZA8b vs MetaLlama8b) raised it by 0.112, pointing to model size and Japanese-specific training as levers.
  • Cultural items such as the Anpanman graduation song show that Japanese-specific norms are part of what the benchmark measures.
  • JETHICS provides a non-Western resource for studying moral understanding, complementing English-centric datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the benchmark is accepted, it becomes a test bed for cross-cultural ethics: the same prompts translated into other languages could expose where model moral judgments are Western-biased.
  • The low kappa (0.18) in the utilitarianism category suggests that category should be interpreted cautiously; a model's score there may reflect which subjective notion of well-being its training data happened to encode rather than a stable moral fact.
  • A natural next step is to use JETHICS for fine-tuning or preference tuning of Japanese LLMs and then re-test on a held-out split; the paper reports evaluation only, not training.
  • The examples from Table 4 could be turned into a targeted diagnostic of culturally specific norms if each item were tagged with the cultural rule it relies on.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper presents JETHICS, a Japanese-language dataset for evaluating the moral understanding of large language models. The dataset contains 77,896 examples across five categories (Utilitarianism, Deontology with Role and Request subcategories, Justice with Desert and Impartiality subcategories, Virtue Ethics, and Commonsense Morality), constructed by following the methodology of the English ETHICS dataset. Crowdworkers generated example-label pairs, and a separate set of crowdworkers validated them via majority vote. The authors evaluate four non-proprietary LLMs (llm-jp 3.7B/13B, Meta-Llama-3-8B, and Llama-3-ELYZA-JP-8B) and two GPT-4o variants in an 8-shot setting, reporting that GPT-4o achieves an average accuracy of 0.713 and the best Japanese model 0.497, which they interpret as evidence that current models have substantial room for improvement in Japanese ethics understanding.

Significance. If the dataset's labels are reliable, JETHICS fills a clear gap: most existing ethics benchmarks are English-centric, and the paper provides a large-scale, non-Western resource with a datasheet and a public-release plan. The evaluation includes a meaningful comparison between a Japanese-tuned model and its base, and the paper is transparent about reporting inter-annotator agreement. The main significance is therefore conditional on the validity of the ground truth, which is questionable for the Utilitarianism category.

major comments (2)
  1. [Section 2.1, Table 2, Section 5] The reported inter-annotator kappa of 0.18 for the Utilitarianism category (19,529 examples, 25% of the dataset) is a load-bearing validity concern, and the paper's defense in Section 5 is insufficient. Please clarify whether the kappa in Table 2 is computed on all step-2 annotations or on the filtered set after excluding split evaluations. If it is computed after exclusion, the kappa is inflated by construction because disagreements are removed; if it is computed before, the filtering removes only exact ties, leaving many examples decided by a 2-1 or 3-1 majority whose reliability is questionable given the low kappa. For a binary task with kappa 0.18, the implied per-annotator agreement is about 0.59, and a majority of three such annotators is only about 0.63 accurate. The paper needs to provide independent validation of the Utilitarianism labels, for example a second annotation pass with a larger panel or expert adjudication, or the category should be reported separately rather than being averaged into the headline results.
  2. [Section 5 and Ethics Statement] The claim that 'since examples with significant disagreement were excluded from JETHICS, the overall quality of the dataset is not compromised' is not supported by the evidence. Excluding split votes does not address the underlying ambiguity that causes low kappa; it merely removes the most extreme disagreements. The Ethics Statement itself acknowledges that 'it is unclear whether the labeled actions are actually morally right since only three or four people annotated a example,' which directly undermines the dataset's validity claim. To make the central claim that JETHICS is a valid Japanese ethics benchmark, the authors should either re-annotate the Utilitarianism category with a larger number of workers, provide external evidence of label validity, or exclude this category from the overall average with a clear caveat.
minor comments (6)
  1. [Section 2.1] The number of annotators per example is not reported; please state the distribution of three versus four annotators and define precisely what 'split evaluations' means for each case (e.g., 1-1-1 with three annotators, 2-2 with four annotators).
  2. [Section 5] The sentence 'the overall quality of the dataset is not compromised' is a claim that needs direct evidence; please provide an analysis of the filtered set, such as the kappa computed on the remaining examples or a comparison of model scores on high-agreement versus low-agreement subsets.
  3. [Datasheet, Distribution section] The datasheet states 'The dataset will be released after the paper is accepted,' while the abstract and introduction state that the dataset is 'released'; please align these statements so that the availability status is unambiguous.
  4. [Appendix B, Table 5] The prompt format appears to contain a duplicated '応答(Response):' line; please verify the actual prompt used in the experiments to avoid confusion.
  5. [Section 2.2, footnote 5] The Commonsense Morality category is based on the authors' earlier Takeshita et al. (2023) dataset; this should be stated in the main text of Section 2.2 so that readers can assess the degree of independence of this subcategory from prior work.
  6. [Table 3] The random baseline for the multi-sentence categories (0.063) is not derived in the text; please specify the number of following sentences per category (e.g., four binary sentences yield 1/16) so that the baseline is verifiable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: JETHICS labels are fixed crowdsourced inputs and model scores are independent measurements against those inputs.

full rationale

JETHICS's construction chain is not circular. The dataset is built by crowdsourcing sentence-label pairs and validating them by majority vote, and model accuracies are then computed against those fixed labels. Nothing in the paper fits a parameter to a subset of the evaluation data and then reports a closely related quantity as a prediction. The only self-referential element is the footnote that the Commonsense Morality category is based on Takeshita et al. (2023), but this is a data provenance statement, not a derivation: the category's labels predate the evaluation and are not constructed from model outputs. Similarly, following Hendrycks et al. (2021) for prompt format and evaluation protocol is methodological borrowing, not circularity. The low utilitarianism kappa (0.18) and the exclusion of split votes raise label-quality and validity concerns, but those are correctness risks rather than circular steps: the majority-vote labels are inputs, and the model scores are independent measurements against those inputs. No equation, fitted parameter, or self-citation chain forces the reported averages, so no significant circularity is present.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central benchmark rests on assumptions about label validity and sampling, not on fitted parameters. The most important assumption is that majority votes by a few crowdworkers constitute a usable ground truth for Japanese morality. No new physical or conceptual entities are introduced; the paper contributes a corpus, not a mechanism.

free parameters (2)
  • Evaluation sample size = 1,000 examples per category
    Section 3 randomly selects 1,000 examples from each category to compute accuracy; no seed or confidence interval is reported, so sampling variability is uncontrolled.
  • Few-shot count = 8
    Section 3 uses an 8-shot prompt following Hendrycks et al. (2021); the choice is not varied, so sensitivity of the scores to the number of shots is unknown.
assumptions (4)
  • domain assumption Crowdworker majority labels (three or four annotators, split votes removed) are a valid ground truth for Japanese moral judgments.
    Section 2.1 step 2 defines labels by majority vote; Section 2.3 reports kappa scores, with utilitarianism at 0.18. If these labels are noisy, all model scores are compromised.
  • domain assumption The hand-written category instructions and 8-shot prompts measure the intended ethical construct rather than test-format artifacts.
    Section 3 and Appendix B specify the prompts; there is no validation that requiring exactly one character does not bias answers or suppress reasoning.
  • domain assumption Randomly selecting 1,000 examples per category gives a representative sample of the full 77,896-example dataset for model comparison.
    Section 3 describes random selection without reporting a seed or confidence intervals, so sampling variability is not quantified.
  • domain assumption The crowdsourced templates faithfully operationalize utilitarianism, deontology, virtue ethics, and justice as described in Section 2.2.
    No expert philosophical validation is reported; category construction follows the authors' reading of the theories.

how reviews work

0 comments
Cite this review

Pith. "Pith review of JETHICS: Japanese Ethics Understanding Evaluation Dataset." pith.science (2026). https://pith.science/paper/6D2SRJRX

@misc{pith2026250616187,
  author       = {Pith},
  title        = {Pith review of: JETHICS: Japanese Ethics Understanding Evaluation Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6D2SRJRX}},
  note         = {Machine review of arXiv:2506.16187}
}
read the original abstract

In this work, we propose JETHICS, a Japanese dataset for evaluating ethics understanding of AI models. JETHICS contains 78K examples and is built by following the construction methods of the existing English ETHICS dataset. It includes four categories based normative theories and concepts from ethics and political philosophy; and one representing commonsense morality. Our evaluation experiments on non-proprietary large language models (LLMs) and on GPT-4o reveal that even GPT-4o achieves only an average score of about 0.7, while the best-performing Japanese LLM attains around 0.5, indicating a relatively large room for improvement in current LLMs.

Figures

Figures reproduced from arXiv: 2506.16187 by the authors.

Figure 1
Figure 1. Original annotation guideline of the commonsense morality category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Original annotation guideline of the virtue category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Original annotation guideline of the utilitarianism category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Original annotation guideline of the justice: desert category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Original annotation guideline of the justice: impartiality category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Original annotation guideline of the deontology: request category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Original annotation guideline of the deontology: role category (url redacted for anonymity) [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses

    cs.CL 2026-01 conditional novelty 6.0 of 10

    CEDAR is a 7-language, 2-modality benchmark of 10,962 culturally divergent emotion scenarios; 17 LLMs perform poorly, and prompt-language matching does not fix cultural misalignment.

Reference graph

Works this paper leans on

25 extracted references · 13 canonical work pages · cited by 1 Pith paper

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Larry Alexander and Michael Moore. 2021. Deontological Ethics . In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy , W inter 2021 edition. Metaphysics Research Lab, Stanford University

  4. [4]

    Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-Fran c ois Bonnefon, and Iyad Rahwan. 2018. The moral machine experiment. Nature, 563(7729):59--64

  5. [5]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, ...

  6. [6]

    David Bourget and David J Chalmers. 2023. Philosophers on philosophy: The 2020 philpapers survey. Philosophers' Imprint, 23

  7. [7]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901

  8. [8]

    Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD : Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 862--872

Show all 25 references
  1. [9]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The L lama 3 herd of models. arXiv preprint arXiv:2407.21783

  2. [10]

    Hwang, Maxwell Forbes, and Yejin Choi

    Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.54 Moral S tories: Situated reasoning about norms, intents, actions, and their consequences . In Proceedings of the 2021 Conference on Empirical Method...

  3. [11]

    Fred Feldman and Brad Skow. 2020. Desert . In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy , W inter 2020 edition. Metaphysics Research Lab, Stanford University

  4. [12]

    Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi

    Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.48 Social chemistry 101: Learning to reason about social and moral norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Languag...

  5. [13]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.301 R eal T oxicity P rompts: Evaluating neural toxic degeneration in language models . In Findings of the Association for Computational Linguist...

  6. [14]

    Jian Guan, Ziqi Liu, and Minlie Huang. 2022. https://doi.org/10.18653/v1/2022.naacl-main.374 A corpus for understanding and generating moral stories . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...

  7. [15]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI with shared human values. In International Conference on Learning Representations

  8. [16]

    Rosalind Hursthouse. 1999. On Virtue Ethics. Oxford University Press, Oxford

  9. [17]

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O'Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, S...

  10. [18]

    LLM-jp. 2024. LLM -jp: A cross-organizational project for the research and development of fully open J apanese LLM s. arXiv preprint arXiv:2407.03963

  11. [19]

    John Rawls. 1999. A theory of justice: Revised edition. Harvard university press

  12. [20]

    Ines Reinig, Maria Becker, Ines Rehbein, and Simone Ponzetto. 2024. https://doi.org/10.18653/v1/2024.findings-acl.245 A survey on modelling morality for text analysis . In Findings of the Association for Computational Linguistics: ACL 2024, pages 4136--4155, Bangkok, Thailand....

  13. [21]

    Andrew E Reisner. 2013. Prima facie and pro tanto oughts. International Encyclopedia of Ethics

  14. [22]

    Sebastin Santy, Jenny Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. https://doi.org/10.18653/v1/2023.acl-long.505 NLP ositionality: Characterizing design biases of datasets and models . In Proceedings of the 61st Annual Meeting of the Association for Computa...

  15. [23]

    Masashi Takeshita, Rafal Rzpeka, and Kenji Araki. 2023. https://www.anlp.jp/proceedings/annual_meeting/2023/pdf_dir/D2-1.pdf Jcommonsensemorality: Japanese dataset for evaluating commonsense morality understanding . In In Proceedings of The Twenty Nineth Annual Meeting of The ...

  16. [24]

    Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.naacl-main.56 On the machine learning of ethical judgments from natural language . In Proceedings of the 2022 Conference of the North America...

  17. [25]

    Christopher Woodard. 2019. Taking utilitarianism seriously. Oxford University Press

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.