{"id":"603e7a18-4442-4a83-b064-50232a9f458a","arxiv_id":"2502.03472","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM products should be certified on curated, domain-specific datasets rather than regulated through compute thresholds or general benchmarks.","lead":"LLM regulation currently keys off training compute and general benchmark scores. This paper argues that such proxies miss the actual user-facing product, and proposes certification built on curated, domain-specific test datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Certification validity is the load-bearing assumption: passing a hidden expert-curated prompt-response set must track real-world consumer-protection outcomes, but the paper offers no evidence and does not address expert disagreement, distribution shift, or contamination.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: certification based on expert-curated prompt-response test sets is assumed to be a valid, robust proxy for real-world LLM-based user experiences. My stress-test sharpens this by pointing to specific technical conditions that would have to hold: well-defined ground truth despite expert disagreement, robustness to distribution shift, and leak-resistance. None of these conditions is demonstrated in the paper. The paper is a position statement rather than an empirical study, so the lack of evidence alone is not disqualifying; the internal logic of the argument is coherent, and the critique of compute thresholds is well supported. However, the proposal's practical value depends entirely on the certification proxy being valid, and the paper does not even acknowledge the most basic measurement-theoretic challenges. A pilot study as described in the concrete test would settle whether the central mechanism works. The verdict should remain CONDITIONAL: the proposal is potentially valuable but requires validation before it can be treated as a basis for regulation. The reader's assessment already captures this, so no change to the verdict is needed.","tokens_in":6621,"tokens_out":2734,"duration_ms":31241,"concrete_test":"Run a pilot in one narrow high-risk domain, e.g., mental-health triage. Build the certification set as specified in §3.4 with multiple expert annotators; measure inter-annotator agreement (e.g., Cohen's kappa) on ground-truth responses and rubric scores. Then have an independent panel of clinicians evaluate a sample of real user interactions from systems that passed and failed certification, blinded to certification status. If inter-annotator reliability is low, or if pass/fail status does not predict independent expert ratings of real interactions, the proxy that the proposal depends on fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that domain-specific datasets should be the engine that powers LLM regulation for consumer protection—stands or falls on whether certification outcomes from the §3.2/§3.4 workflow are a valid proxy for real-world safety and effectiveness. The paper asserts this proxy without supporting evidence. Three concrete failure modes are not addressed. First, for open-ended high-risk use cases (mental-health coaching, medical advice), there is no unique ground-truth response; expert disagreement is likely, so the 'ground truth' in §3.2 is not well-defined. Second, a hidden test set is a static sample of prompts; real deployments face distribution shift, adversarial inputs, and multi-turn dynamics that the sample cannot capture. Third, leakage and contamination: sharing a training set in the same domain makes keeping the hidden set truly hidden and continually refreshed difficult, and the paper does not discuss Goodhart-style overfitting to certification criteria. Therefore passing certification may not imply consumer protection, and the central regulatory benefit remains unestablished. This is a validation gap rather than an internal logical contradiction, but for a regulatory proposal it is the load-bearing weakness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that compute-level thresholds and generalized model evaluations are insufficient for consumer protection in LLM-based systems, and proposes that domain-specific, expert-curated prompt-response datasets should be the central mechanism of LLM regulation. The paper critiques existing compute-based proposals, outlines a potential certification workflow involving ground-truth dataset construction, public training sets, hidden test sets, and phased auditing, and claims benefits for consumers, businesses, and regulators. The paper is explicitly non-prescriptive about governance details and acknowledges that dataset curation is labor-intensive, but it offers no empirical or formal validation of the proposed certification mechanism.","tokens_in":6722,"tokens_out":3977,"duration_ms":44058,"significance":"The paper makes a timely and useful argument by shifting the regulatory conversation from model-level proxies to user-facing, domain-specific evaluation, and it correctly identifies the brittleness of compute thresholds. It also gives a concrete workflow that could, in principle, be pilot-tested, and it is honest about the labor cost of dataset curation. However, the entire regulatory proposal rests on an unvalidated assumption: that passing a hidden, expert-curated prompt-response test set is a valid proxy for real-world safety and effectiveness. That assumption is load-bearing and, as the paper stands, unsupported. The paper would be a meaningful contribution if it either narrowed its claims to domains with well-defined ground truth or provided evidence from analogous certification regimes, a pilot study, or a structured argument addressing validity threats.","major_comments":[{"comment":"The central claim depends on the assertion that expert-curated prompt-response datasets with 'ground truth' responses validly measure the safety and effectiveness of real-world LLM-based experiences. This assertion is not defended. In the paper's own high-risk examples, such as mental-health coaching or medical advice, there is often no unique correct response, so 'ground truth' is not well-defined. The paper does not discuss inter-annotator agreement, adjudication of expert disagreement, or how a rubric resolves ambiguity. Without such a protocol, a certification outcome would not be reproducible or legally defensible, which undermines the proposed mechanism for consumer protection.","section":"§3.2"},{"comment":"The proposed workflow's hidden test set is effectively a static sample of prompts evaluated at certification time, but real deployments face distribution shift, adversarial inputs, and multi-turn dynamics that a static sample cannot capture. The paper stipulates that the test set should be 'continually curated' but does not specify how drift is detected, how often refresh occurs, or how certification remains valid between recertifications. Appendix A says that recertification timing and update triggers 'should be explicitly defined' as part of the process, but it does not define them. This gap matters because a certification that is stale at deployment time cannot deliver the promised consumer protection.","section":"§3.4, steps 4 and 5"},{"comment":"The paper does not adequately address contamination and Goodhart-style overfitting to the certification. The one-sentence stipulation that the test set should be 'continually curated as a means to prevent the evaluation from directly leaking into model training' is not a contamination-resistance strategy. Moreover, step 2 permits LLM-as-judge evaluation, despite citing Bavaresco et al. [29], which empirically documents cases where LLM judges diverge from human judgments. The paper does not reconcile this evidence with the use of LLM-as-judge in high-risk certification, nor does it specify how the hidden test set remains secret and fresh in a domain where a public training set is shared. For a regulatory test, a concrete contamination-resistance and validation protocol is required.","section":"§3.4, steps 2 and 4"},{"comment":"The claimed benefits — increased consumer confidence, enabling businesses to demonstrate credibility, and expansion into previously risky domains — are asserted rather than demonstrated. No evidence from analogous certification schemes (e.g., FDA approvals, professional licensing, or third-party seals) is provided, and no cost-benefit or stakeholder analysis is offered. Since these benefits are the stated motivation for the regulatory shift, the paper needs at least a structured argument from existing certification regimes or a small case study to make the proposal persuasive.","section":"§3.1"}],"minor_comments":[{"comment":"There is a typo: 'The EU AI Act [11] uses also uses a fixed threshold' should read 'also uses a fixed threshold.'","section":"§1"},{"comment":"The paper refers to the 'Whitehouse' in the text and reference [30]; the standard spelling is 'White House,' and the M-24-10 memorandum should be cited with the specific appendix or section numbers relevant to the claim.","section":"§1 and references"},{"comment":"The choice of the M-24-10 Appendix 1 lists as the starting point for prioritization is not justified relative to other frameworks, such as the EU AI Act's high-risk categories or the Colorado AI Act. A brief comparative rationale would strengthen the prioritization discussion.","section":"§3.3"},{"comment":"The paper acknowledges that dataset creation and maintenance are labor-intensive but does not give even rough estimates of cost, time, or required expertise. Since the proposal's feasibility depends on these resources, a brief discussion of scale would be helpful.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style position paper with no empirical validation. The core thesis is plausible and timely, and the paper does a service by moving the debate toward use-case-specific evaluation. However, the certification proxy is the load-bearing component, and the manuscript does not currently address the standard validity threats. I would not reject, because the gap is addressable by narrowing the scope, adding a validation protocol, and including evidence from analogous certification regimes. If the journal's editorial standard requires empirical evaluation of proposed regulatory mechanisms, rejection could be justified; my recommendation assumes the journal accepts position papers with rigorous argumentation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a position paper, not a study. The critique of compute thresholds is solid but not new—Hooker's point is the foundation, and the paper builds on it. The real contribution is the concrete five-step certification workflow in Section 3.4, which is specific enough to argue with. That alone makes it worth a referee's time.\n\nThe paper does well in structuring the argument. It names the three limitations of compute-level regulation cleanly, uses the drug-ingredient analogy effectively, and honestly flags the labor cost of dataset curation. The citation pattern is fine: Hooker, domain-specific evaluation work, NIST guidance, and the EU AI Act are all appropriately placed. No self-citation, no fitted parameters, no circularity.\n\nThe soft spot is the one the stress-test flags: the load-bearing assumption is that passing an expert-curated hidden prompt-response set tracks real-world consumer protection. The paper asserts this proxy without evidence. It doesn't address that for high-risk open-ended tasks like mental-health coaching there is no unique ground-truth response, so expert disagreement is likely and the \"ground truth\" in Section 3.2 is not well-defined. It also doesn't address distribution shift after certification, or the leakage and Goodhart dynamics that arise when a training set is shared publicly. Passing certification may not imply consumer safety, and the central regulatory benefit is therefore unestablished.\n\nThis is a validation gap, not an internal contradiction. The internal logic is coherent. But the paper should be clearer in Section 3.1 that the benefits are hypotheses, not measured outcomes. The claims about consumer confidence and business credibility are restatements of the proposal.\n\nWho is this for? People working on AI governance, especially regulators and workshop attendees. It is not a technical contribution and it offers no empirical evidence, so it should not be judged as one. For a position paper it does its job: it gives the community a concrete proposal to critique.\n\nMy recommendation: send it to peer review for a workshop venue like Regulatable ML, with the expectation that the authors add a section on validity threats—expert disagreement, distribution shift, contamination—and ideally a small case study. If it's being considered for a journal, it needs substantially more. As a workshop position paper, it deserves engagement.","headline":"A coherent, well-cited position paper that consolidates known critiques of compute thresholds into a concrete dataset-certification workflow, but the load-bearing assumption that certification tracks real-world consumer protection is asserted, not argued.","tokens_in":7307,"tokens_out":1440,"would_cite":false,"duration_ms":15799,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that LLM regulation should be built on domain-specific evaluation datasets rather than compute thresholds.","keywords":["LLM regulation","compute thresholds","domain-specific datasets","AI certification","consumer protection","LLM-as-judge","dataset curation","AI policy"],"falsifier":"An observational study comparing certified and uncertified LLM products in the same high-risk domain, using expert review of live user interactions, would settle the claim: if certified products show the same rate of harmful or incorrect outputs as uncertified ones, then passing a domain-specific test set does not protect consumers as proposed.","tokens_in":6331,"feed_emoji":"⚖️","tokens_out":5552,"duration_ms":51585,"temperature":0.7,"pith_summary":"The paper proposes shifting consumer-protection regulation of LLM products away from compute thresholds and generic model benchmarks, which are indirect proxies, and toward certification of specific user-facing experiences using expert-curated domain-specific datasets. It argues that such datasets, carefully maintained sets of prompts with ground-truth responses, can test whether a particular LLM-based product performs at an acceptable level in the actual context where consumers encounter it. If the approach works, regulators could target high-risk use cases directly, businesses would have a clear path to demonstrate credibility, and consumers would get a meaningful signal about the products they use.","feed_headline":"Test LLM products on domain datasets, not compute thresholds","feed_subtitle":"A proposal to certify LLM-based experiences through expert-curated prompt-response tests for consumer protection.","key_machinery":"The carrying mechanism is a dynamic, expert-curated prompt-response dataset: pairs of realistic user inputs and ground-truth responses for a specific use case, kept current over time. Around this dataset the paper builds a certification procedure with a rubric and scoring system, tiered passing criteria, a public training set and a hidden test set, and a phased audit plan that starts with expert manual review and later adds LLM-as-judge automation. The dataset, not the model's compute budget, is the measuring instrument, which is why it must be domain-specific and continually maintained.","core_discovery":"On its own terms, the paper's central claim is that compute thresholds and generalized model evaluations are not sufficient measures of the risk of a specific LLM-based consumer experience. The paper asserts that domain-specific datasets should be the engine of LLM regulation for consumer protection, and it sketches a certification workflow centered on these datasets. The workflow pairs expert-curated prompts with ground-truth responses, uses a rubric and tiered passing criteria, keeps a hidden test set, and moves from manual expert review to LLM-as-judge automation for scalability.","pith_inferences":["A testable extension the paper leaves implicit: certifying an experience and then tracking real-world consumer complaints against it would calibrate whether passing thresholds are set too leniently.","If domain-specific test sets become commercially valuable, dataset leakage and overfitting become the central failure mode; an independent adversarial audit of hidden test questions would be needed to keep certification meaningful.","The same logic could extend beyond LLMs to any probabilistic generative system where the user-facing behavior, not the model card, determines harm.","Synthetic data could make dataset curation cheaper, but the paper's reliance on expert ground truth implies synthetic data would need validation against human expert judgment before it can replace it."],"forward_implications":["Regulators would stop treating all LLM-based systems as one uniform entity and instead certify specific user-facing experiences for specific use cases.","Smaller, specialized models that fall below compute thresholds would still be covered if their specific use case is high-risk, since coverage follows the experience rather than the compute bill.","Businesses gain a standardized path to prove product safety and effectiveness, which could unlock offerings in areas like healthcare that are currently considered too risky.","Certification would need to be dynamic, with continually curated test sets and explicit re-certification triggers as systems and marketplaces evolve.","Evaluation would scale through a phased process: expert manual review first, then LLM-as-judge automation with manual auditing retained."],"supporting_citations":[{"why":"Supplies the argument that compute thresholds are limited because compute and model performance are not perfectly correlated, undercutting compute-level regulation.","marker":"[1]"},{"why":"Provides the method of domain-specific evaluation sets and LLM-as-judge that the proposed certification workflow adopts.","marker":"[3]"},{"why":"Offers a rubric-based approach to evaluating domain-specific human-AI conversations, informing the rubric step of the certification process.","marker":"[2]"},{"why":"Explains that regulators lack access to customer use cases, motivating the data-centric shift proposed by the paper.","marker":"[5]"},{"why":"Supplies the risk-tier framework (the EU AI Act) that the paper uses to prioritize which use cases should receive certification.","marker":"[11]"},{"why":"Describes the existing compute-threshold regulation (Executive Order 14110) and the consumer benefits that the proposal builds on.","marker":"[6]"},{"why":"Illustrates generalized model evaluation (HELM), which the paper argues is insufficient for specific use cases.","marker":"[22]"},{"why":"References the need for certification programs relevant to specific industry and context, supporting the proposal's direction.","marker":"[4]"}],"fun_headline_variants":["Regulate LLMs with domain data, not compute limits","Switch LLM oversight from compute specs to user tests","Data-centric certification: the new way to govern LLMs","Expert-curated data should drive LLM product certification","Move LLM regulation from threshold numbers to real experiences"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proposal stands on the assumption that a hidden, expert-curated set of sample prompts and ideal answers, continuously updated, is a reliable measure of whether a real LLM-based product is safe and effective for consumers.","fun_headline_variants_meta":{"raw":{"variants":["Regulate LLMs with domain data, not compute limits","Switch LLM oversight from compute specs to user tests","Data-centric certification: the new way to govern LLMs","Expert-curated data should drive LLM product certification","Move LLM regulation from threshold numbers to real experiences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000132,"raw_usage":{"total_tokens":1077,"prompt_tokens":835,"completion_tokens":242,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":164}},"tokens_in":451,"tokens_out":242,"duration_ms":2930,"temperature":1.0,"reasoning_tokens":164,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:49:05.502515+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An observational study comparing certified and uncertified LLM products in the same high-risk domain, using expert review of live user interactions, would settle the claim: if certified products show the same rate of harmful or incorrect outputs as uncertified ones, then passing a domain-specific test set does not protect consumers as proposed.","supporting_citations":[{"cited_title":"Rubicon: Rubric-based evaluation of domain-speciﬁc human ai conversations,","cited_arxiv_id":null,"evidence_quote":"Offers a rubric-based approach to evaluating domain-specific human-AI conversations, informing the rubric step of the certification process."},{"cited_title":"Grounding ai policy towards researcher ac cess to ai usage data,","cited_arxiv_id":null,"evidence_quote":"Explains that regulators lack access to customer use cases, motivating the data-centric shift proposed by the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the risk-tier framework (the EU AI Act) that the paper uses to prioritize which use cases should receive certification."},{"cited_title":"Executive order on the safe, secure, an d trustworthy development and use of artiﬁcial intelligence,","cited_arxiv_id":null,"evidence_quote":"Describes the existing compute-threshold regulation (Executive Order 14110) and the consumer benefits that the proposal builds on."},{"cited_title":"Artiﬁcial in- telligence risk management framework: Generative artiﬁci al intelligence proﬁle,","cited_arxiv_id":null,"evidence_quote":"References the need for certification programs relevant to specific industry and context, supporting the proposal's direction."}],"review_version":1}