REVIEW 8 cited by
CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the rapid evolution of large language models (LLMs), there is a growing concern that they may pose risks or have negative social impacts. Therefore, evaluation of human values alignment is becoming increasingly important. Previous work mainly focuses on assessing the performance of LLMs on certain knowledge and reasoning abilities, while neglecting the alignment to human values, especially in a Chinese context. In this paper, we present CValues, the first Chinese human values evaluation benchmark to measure the alignment ability of LLMs in terms of both safety and responsibility criteria. As a result, we have manually collected adversarial safety prompts across 10 scenarios and induced responsibility prompts from 8 domains by professional experts. To provide a comprehensive values evaluation of Chinese LLMs, we not only conduct human evaluation for reliable comparison, but also construct multi-choice prompts for automatic evaluation. Our findings suggest that while most Chinese LLMs perform well in terms of safety, there is considerable room for improvement in terms of responsibility. Moreover, both the automatic and human evaluation are important for assessing the human values alignment in different aspects. The benchmark and code is available on ModelScope and Github.
Forward citations
Cited by 8 Pith papers
-
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
The authors release a large multimodal benchmark showing that current LMMs struggle to detect toxicity that emerges only from combining image and text, and that many-shot toxic demonstrations further reduce their accuracy.
-
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.
-
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
MM-Eval is a new Modern Mongolian benchmark showing LLMs perform best at syntax, worse at semantics and knowledge, and worst at reasoning.
-
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
Value Compass Benchmarks is a live, self-evolving platform that scores 33 LLMs across 27 value dimensions from four value systems, aiming to reveal true behavioral alignment with human values.
-
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
A 2,000-question Chinese safety factuality benchmark shows most LLMs are inaccurate on safety knowledge, with retrieval helping more than self-reflection.
-
Aligning LLMs with Domain Invariant Reward Models
Applying Wasserstein-distance domain adaptation, a known technique, to reward models lets preference signals learned on labeled source data transfer to unlabeled target domains, with consistent but modest gains across...
-
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
A multi-stage LLM-based attack/defense dataset pipeline improves reported safety scores of Llama-3.2-1B after SFT, but the evaluation is partly circular and lacks statistical baselines.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
Discussion (0). Continue with ORCID to comment.