REVIEW 5 cited by
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/CCI3-Data), developed using a novel two-stage hybrid filtering pipeline that significantly enhances data quality. To evaluate its effectiveness, we trained a 0.5B parameter model from scratch on 100B tokens across various datasets, achieving superior performance on 10 benchmarks in a zero-shot setting compared to CCI3.0, SkyPile, and WanjuanV1. The high-quality filtering process effectively distills the capabilities of the Qwen2-72B-instruct model into a compact 0.5B model, attaining optimal F1 scores for Chinese web data classification. We believe this open-access dataset will facilitate broader access to high-quality language models.
Forward citations
Cited by 5 Pith papers
-
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
Sampled-BPE uses a small sample and BPE token statistics to estimate token-level pollution in web-scale Chinese corpora, revealing heavy and shifting adult content in Common Crawl.
-
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
The paper introduces SeedBench, an expert-validated 2,264-question benchmark for LLMs in seed science, and reports that the best models average around 62 to 63 points, well below expert-level performance.
-
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
Ultra-FineWeb is a fastText-filtered pretraining corpus whose seed samples were chosen by a cheap 'efficient verification' step, and 1.2B models trained on it outperform models trained on FineWeb and FineWeb-edu on av...
-
MiniCPM4: Ultra-Efficient LLMs on End Devices
MiniCPM4-8B reportedly matches Qwen3-8B on standard benchmarks while using about 22% of the training tokens, and achieves large long-context speedups on edge devices.
-
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
OpenCSG released four open Chinese LLM training datasets, and 2B-scale tests report improved C-Eval, CMMLU, and Alignbench scores.
Discussion (0). Continue with ORCID to comment.