Pith. sign in

REVIEW 5 cited by

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18058 v2 pith:ESGZ4K6W submitted 2024-03-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords chinesecoig-cqiadatasetsinstructionmodelstuningdatasetllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remarkable progress on English instruction tuning has facilitated the efficacy and reliability of large language models (LLMs). However, there remains a noticeable gap in instruction tuning for Chinese, where the complex linguistic features pose significant challenges. Existing datasets, generally distilled from English-centric LLMs, are not well-aligned with Chinese users' interaction patterns. To bridge this gap, we introduce COIG-CQIA, a new Chinese instruction tuning dataset derived from various real-world resources and undergoing rigorous human verification. We conduct extensive experiments on COIG-CQIA, and compare them with strong baseline models and datasets. The experimental results show that models trained on COIG-CQIA achieve highly competitive performance in diverse benchmarks. Additionally, our findings offer several insights for designing effective Chinese instruction-tuning datasets and data-mixing strategies. Our dataset are available at https://huggingface.co/datasets/m-a-p/COIG-CQIA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A large synthetic instruction corpus with guidelines, preference rules, and format variants improves LLM performance on five NLU benchmarks by an average of 3.1%.

  2. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.5 of 10

    A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.

  3. CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    A new 35 TB bilingual pretraining dataset with 4.5 billion chain-of-thought templates is described, but the evidence for its benefits is marginal, confounded, and contradicted by the paper's own tables.

  4. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Reinforcing the two experts most correlated with thinking tokens improves reasoning accuracy and efficiency in MoE large reasoning models, with gains of up to 10 points on AIME benchmarks.

  5. Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    A survey of LLM hallucination research that formalizes hallucination types and argues, via incompleteness and undecidability arguments, that hallucinations cannot be fully eliminated.

Pith tools