{"id":"8e33924b-ec10-4caa-ab9e-cc6903af7169","arxiv_id":"2412.04947","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"C2LEVA is a bilingual, multi-task LLM benchmark that combines passive test-data renewal with active data watermarking to reduce contamination risk, and ranks 15 models.","lead":"This paper presents C2LEVA, a bilingual LLM benchmark designed to resist data contamination by continuously renewing test data and embedding watermarks to deter cheating. It evaluates 15 leading models across dozens of tasks and compares the resulting ranking against Chatbot Arena.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The contamination-free guarantee is not empirically supported: the Min-K% filter is unvalidated and the Arena correlation does not test contamination.","rationale":"I agree with the reader that the weakest assumption is the unvalidated Min-K% filter applied at construction time. My reading adds that the paper's own validation in §4.3 cannot support the contamination-free claim, since rank correlation with Chatbot Arena only speaks to benchmark coverage and model ordering, not to whether individual test instances appear in training corpora. I also note an internal inconsistency: §3.3 promises a maximum watermarking-induced performance loss of 5%, while Table 3 reports an average loss of 11.59% across models, suggesting the active-prevention layer distorts scores more than designed. Both issues are addressable with a direct contamination audit, so the reader's conditional verdict remains appropriate; no change in verdict is needed.","tokens_in":39969,"tokens_out":4021,"duration_ms":43749,"concrete_test":"Perform a contamination audit on the open models in C2LEVA: take text verifiably present in the pretraining data of an open model (e.g., memorized training extracts from Llama-3-8B), build a synthetic set mixing these known-contaminated instances with uncontaminated instances from the released C2LEVA set, run the paper's Min-K% pipeline with the same detector model and threshold as §3.3, and measure the false-negative rate. Repeat with a second detector (e.g., Qwen2-7B) to test transferability. If contaminated instances pass the filter at a non-negligible rate, the passive prevention layer does not substantiate the 'contamination-free' claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that C2LEVA provides a contamination-free, trustworthy assessment—depends on the passive contamination filter in §3.3: every crawled or LLM-generated test instance is screened with Min-K% scores from Llama-3-8B, a single model chosen because it was 'trained on 15T tokens' and therefore representative of other web-trained models. This assumption has two unverified parts: (i) that Min-K% detection transfers across architectures and pretraining distributions, and (ii) that the detector is calibrated at the chosen threshold. No threshold value, no contamination detection recall or precision, and no cross-model comparison are reported. The only validation offered in §4.3 (Fig. 6) is a high Spearman correlation between C2LEVA mean win rates and Chatbot Arena Elo, but correlation with a human-preference leaderboard does not establish absence of contamination; a contaminated benchmark can still reproduce the same ranking if contamination is diffuse or memorized similarly across models. Additionally, §3.3 claims watermarking is designed for 'a maximum performance loss of 5%', yet Table 3 reports an average loss of 11.59% across models, with individual decreases up to 38.39% (Claude-3.5 in Chinese), contradicting the design guarantee and further undermining the trustworthiness of watermarked task scores. As a result, the 'contamination-free' claim is asserted rather than demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces C2LEVA, a bilingual (English/Chinese) LLM benchmark with 22 tasks spanning application assessment and ability evaluation. Its central contribution is a systematic contamination-prevention pipeline: passive prevention via crawling, rule-based synthesis, LLM assistants, contamination detection (Min-K%), and data augmentation; plus an active prevention layer using watermarking, licensing, and encryption. The authors evaluate 15 open-source and proprietary LLMs, report a leaderboard, and validate the benchmark by correlating mean win rates with Chatbot Arena Elo (Spearman rank correlation 0.948). The paper claims that C2LEVA provides a 'contamination-free' and trustworthy assessment.","tokens_in":40275,"tokens_out":4447,"duration_ms":43489,"significance":"If the contamination-free claim were rigorously supported, C2LEVA would be a valuable contribution: it covers tasks often missing from dynamic benchmarks (harms, knowledge, language), provides bilingual coverage, uses multiple prompt templates, and combines passive and active prevention in a principled framework. The large-scale evaluation of 15 models and the public leaderboard are also useful. However, the load-bearing assertion that the benchmark is contamination-free is not directly demonstrated. The contamination-detection filter is unvalidated, the watermarking design goal contradicts the measured distortion, and the Arena correlation does not test contamination. These gaps make the central claim currently overreaching.","major_comments":[{"comment":"The Min-K% filter is the only passive safeguard for crawled and LLM-generated data, but the paper reports no threshold value, no detection recall or precision, and no cross-model transfer experiments. Using Llama-3-8B as a 'representative model' is an assumption that needs empirical support. Please add a validation experiment with known contaminated samples (e.g., documents from pretraining corpora or simulated membership) and report the ROC/AUC of Min-K% for several of the evaluated models at the chosen threshold.","section":"§3.3 (contamination detection)"},{"comment":"The text states that watermarked test cases are 'designed to ensure a maximum performance loss of 5%', but Table 3 shows an average loss of 11.59% across models and a per-model Chinese loss up to 38.39% for Claude-3.5. This is an internal contradiction. The design guarantee must either be revised to match the measured distortion, or the watermarking strength must be recalibrated to meet the 5% target. As written, the active-prevention component itself introduces nontrivial evaluation distortion.","section":"§3.3 vs. Table 3"},{"comment":"The Spearman correlation of 0.948 with Chatbot Arena Elo is high, but this only shows that C2LEVA ranks models similarly to a human-preference leaderboard. A contaminated benchmark can also achieve high rank correlation if contamination is diffuse or correlated with general capability. The sentence 'This supports the conclusion that C2LEVA is comprehensive and mitigates data contamination' is not justified by the evidence; the correlation is a consistency check, not a contamination test. Please either remove this claim or add a direct test, such as comparing model performance on instances flagged versus not flagged by the detector, or using a known contaminated subset.","section":"§4.3, Fig. 6"},{"comment":"The systematic-prevention claim is not uniform across tasks. Contamination detection is not applied to reasoning-primitive or realistic-reasoning tasks, nor to copyright; active watermarking is applied to only 5% of one task (fact completion). At minimum, the paper should state plainly which tasks have which safeguards and qualify the benchmark-level claim accordingly. As written, the abstract's 'contamination-free tasks' overstates the coverage of the proposed pipeline.","section":"Table 2 and §3.3 (coverage of prevention)"}],"minor_comments":[{"comment":"These figures contain garbled placeholder strings (e.g., '/uni...' sequences) instead of readable labels and values; they must be regenerated with proper text rendering.","section":"Figures 3, 4, 6, 8, 9"},{"comment":"The human quality assessment of generated theses reports scores of 0.626 (Chinese) and 0.616 (English), but does not state the number of annotators or inter-annotator agreement; a single annotator is insufficient to establish reliability.","section":"Appendix E.3"},{"comment":"The statement 'The Spearman's rank correlation is 0.948 with p < 0.05' should include the exact p-value and the number of paired observations used in the correlation.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper would benefit from tempering the abstract's 'contamination-free' claim during revision; the current evidence supports a carefully qualified claim about systematic prevention rather than an absolute guarantee. The fit with a computational linguistics venue is appropriate, but the missing validation of the core contamination-detection mechanism is substantial and should be the focus of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: C2LEVA is a useful benchmark, and the active-prevention piece — data watermarking, licensing, and encryption as part of benchmark release — is genuinely new as far as I know. The paper deserves a real referee, but it should not get away with the headline claim that the benchmark is contamination-free: that claim is asserted, not demonstrated.\n\nWhat's good: the benchmark is bilingual (English/Simplified Chinese), covers 22 tasks as claimed across application and ability categories, uses automated renewal (crawling, rule-based synthesis, LLM-assisted generation), and evaluates 15 models. The authors are unusually honest that watermarking distorts results, and they limit it to 5% of fact-completion data. They ship code and data, which is concrete and reusable.\n\nThe soft spots are real. The central 'contamination-free' guarantee rests on Min-K% detection with Llama-3-8B as a representative model, but no threshold, no detection precision/recall, and no cross-model validation are reported. That is a load-bearing gap: every crawled or generated instance is screened by a filter whose accuracy is unmeasured. The Arena correlation (rho=0.948) is a nice sanity check for ranking quality, but it cannot establish absence of contamination — a uniformly contaminated or scale-correlated benchmark would still rank models this way.\n\nThere is also a direct contradiction: Section 3.3 says watermarked test cases are designed for a maximum performance loss of 5%, but Table 3 shows an average loss of 11.59% and individual losses up to 38.39%. The paper needs to either fix the design target or explain the discrepancy.\n\nMinor: the paper says 22 tasks, but counting the taxonomy in Section 3.2 and Appendix A gives 19 (4 reasoning primitive + 7 realistic reasoning + 3 application + 2 language + 1 knowledge + 2 harms). Likely a labeling issue, but it should be corrected.\n\nBottom line: the benchmark and the active-prevention framework are worth engaging with, but the core claim needs validation or substantial softening. If this comes to you for review, send it out, and push hard on the contamination detection evidence and the watermark numbers.","headline":"Genuinely new active-prevention angle on a solid bilingual benchmark, but the 'contamination-free' claim outruns the evidence in the paper.","tokens_in":40766,"tokens_out":4009,"would_cite":true,"duration_ms":43313,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"C2LEVA claims that combining test-set renewal, contamination filtering, and data watermarking produces LLM rankings that are not inflated by memorized test data.","keywords":["data contamination","LLM evaluation","benchmark","contamination detection","data watermarking","bilingual evaluation","test set renewal"],"falsifier":"Take the released C2LEVA test set, compute Min-K% scores with Llama-3-8B as the paper does, and then check a sample of instances against the training corpora or internal memorization behavior of several newly released LLMs; if a non-trivial fraction of instances that the filter scored as clean are memorized, the contamination-free guarantee fails. A cheaper version uses a deliberately contaminated instance that Min-K% scores as clean, trains a small model on it, and shows that the model then answers it correctly.","tokens_in":39817,"feed_emoji":"🛡️","tokens_out":5199,"duration_ms":48408,"temperature":0.7,"pith_summary":"C2LEVA is a bilingual (English and Simplified Chinese) benchmark that the authors present as both comprehensive and contamination-free, covering 22 tasks across application assessment and ability evaluation (language, knowledge, reasoning, harms). The paper's central claim is that its systematic prevention strategy, which automates test-data renewal and adds data protection, yields trustworthy LLM rankings that are not inflated by models having memorized the test set. The authors support this by evaluating 15 open and proprietary models and by showing that C2LEVA's mean win rates track Chatbot Arena Elo with a Spearman correlation of 0.948. A careful reader would care because benchmark contamination currently undermines the reliability of LLM leaderboards, and this paper attempts a defense that combines passive renewal with active protection. The paper also documents that its active protection (data watermarking) degrades measured performance in most cases, which is a cost that future benchmarks will have to manage.","feed_headline":"New benchmark renews and watermarks tests to stop LLM data leakage","feed_subtitle":"C2LEVA ranks 15 models on 22 bilingual tasks and checks results against a human-vote leaderboard.","key_machinery":"The load-bearing mechanism is the two-part contamination prevention pipeline. Passive prevention automates test-set construction from fresh web content, rule-based generators, and LLM assistants, then filters candidate test instances with Min-K% (a per-instance contamination risk score computed from token probabilities of a representative model, Llama-3-8B) and augments scarce data with synonym substitution. Active prevention applies data-protection techniques: a CC BY-NC-ND license, ZipCrypto encryption, and sparse random-sequence watermarking designed to allow provable membership inference with a stated maximum 5% performance loss and p-value near 0.05. The paper's validation metric is the mean win rate across tasks, and its key external check is the Spearman correlation (0.948) between C2LEVA rankings and Chatbot Arena Elo.","core_discovery":"In the paper's own terms, the discovery is that contamination prevention for LLM evaluation can be made systematic by pairing passive prevention with active prevention. Passive prevention continuously renews test data through crawling, rule-based synthesis, and LLM-assisted generation, and filters risky instances using Min-K% token-probability contamination detection; active prevention makes the released data harder to misuse by licensing it, encrypting the archive, and watermarking a subset of test inputs so that unauthorized memorization can be proven. Applied across 22 tasks in two languages, this framework yields a benchmark whose model ranking correlates strongly with an independent, human-vote-based leaderboard, which the paper treats as evidence that the benchmark is both comprehensive and not contaminated. A secondary, cautionary finding is that watermarking measurably distorts evaluation results, with an average performance loss of about 11.59% across models in the fact completion task.","pith_inferences":["The contamination-free claim is only as strong as the Min-K% detector's ability to generalize: if a future model was trained on data that the Llama-3-8B-based filter scored as clean, that model's C2LEVA score could still be inflated, so the benchmark should publish the detector's operating characteristics.","A natural stress test is to train a small model deliberately on a held-out portion of C2LEVA and see whether Min-K% flags those instances before release; if it does not, the renewal pipeline needs a stronger filter.","The watermarking distortion finding suggests that active prevention may be viable only for tasks where small performance shifts do not change ranking conclusions, or where stronger watermarks can be developed that preserve task semantics.","The two-language design invites extension to more languages and modalities, where contamination risk from web-scale training data is at least as severe."],"forward_implications":["If C2LEVA's contamination prevention works as claimed, current and future LLM rankings from the benchmark are not inflated by test-set memorization.","The benchmark provides a reusable 22-task, bilingual template for evaluating application skills and four ability dimensions without relying on stale test data.","The documented watermarking distortion implies that contamination-free evaluation carries a measurable accuracy cost that must be traded off against protection strength.","The strong correlation with Chatbot Arena Elo suggests that benchmark rankings derived from renewed, filtered test data can reproduce independent human-preference rankings.","Because the framework is automated, the leaderboard can be continuously updated as new models and new data appear, without rebuilding the benchmark from scratch."],"supporting_citations":[{"why":"Supplies the Min-K% contamination detection used to filter test instances.","marker":"(Shi et al., 2023b)"},{"why":"Supplies the random-sequence data watermarking used for active prevention.","marker":"(Wei et al., 2024)"},{"why":"Provides the Chatbot Arena Elo used as an independent ground-truth ranking for validation.","marker":"(Chiang et al., 2024)"},{"why":"Provides HELM's task taxonomy and reasoning-primitive definitions that C2LEVA adopts.","marker":"(Liang et al., 2022)"},{"why":"Supplies DyVal's rule-based reasoning task synthesis used for realistic reasoning tasks.","marker":"(Zhu et al., 2023a)"},{"why":"Supplies the crawling approach and typo-fixing task design for fresh web data.","marker":"(White et al., 2024)"},{"why":"Supplies the overall task taxonomy and the fact-completion methodology reused in C2LEVA.","marker":"(Li et al., 2023)"},{"why":"Motivates licensing and encryption as active prevention against cooperative developers.","marker":"(Jacovi et al., 2023)"}],"fun_headline_variants":["22-task bilingual benchmark stops LLM leakage with renewal and watermark","C²LEVA renews and watermarks tests to keep LLM evaluation contamination-free","New benchmark combats LLM contamination with passive and active defenses","Renewing and watermarking tests: C²LEVA's recipe for contamination-free LLM eval"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guarantee that every test instance is unseen by any evaluated model rests on a single filter, Min-K% scores computed with Llama-3-8B, catching all instances that appear in any model's training data, plus the assumption that newly crawled web text has not already been absorbed into those training corpora.","fun_headline_variants_meta":{"raw":{"variants":["22-task bilingual benchmark stops LLM leakage with renewal and watermark","C²LEVA renews and watermarks tests to keep LLM evaluation contamination-free","New benchmark combats LLM contamination with passive and active defenses","Renewing and watermarking tests: C²LEVA's recipe for contamination-free LLM eval"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000896,"raw_usage":{"total_tokens":3814,"prompt_tokens":850,"completion_tokens":2964,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":2880}},"tokens_in":466,"tokens_out":2964,"duration_ms":24585,"temperature":1.0,"reasoning_tokens":2880,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:06:25.430494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released C2LEVA test set, compute Min-K% scores with Llama-3-8B as the paper does, and then check a sample of instances against the training corpora or internal memorization behavior of several newly released LLMs; if a non-trivial fraction of instances that the filter scored as clean are memorized, the contamination-free guarantee fails. A cheaper version uses a deliberately contaminated instance that Min-K% scores as clean, trains a small model on it, and shows that the model then answers it correctly.","supporting_citations":[],"review_version":1}