{"id":"a2e11645-11d5-4bb1-975a-9099b6d461b9","arxiv_id":"2412.09990","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SuperNUGGETS uses small language models and a refined 100-example test set to pick instruction data, matching the quality of the LLM-based NUGGETS at far lower cost.","lead":"This paper presents SuperNUGGETS, a faster variant of the NUGGETS method for selecting high-quality instruction data to fine-tune large language models. It replaces the large data-prospecting model with a small one and shrinks the test set from 1,000 random examples to 100 curated ones, reporting comparable fine-tuning quality at much lower compute.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline parity/efficiency claim is not actually tested: no original NUGGETS baseline or runtime measurement appears, and the table values do not consistently support a 1-2% drop.","rationale":"The paper's value proposition is not 'golden score works' per se, but that the SLM-based variant matches NUGGETS at 1/58 the cost. If that quantified claim is unsupported, the paper reduces to an ablation of predefined-task-set refinement, which is still interesting but not the claimed contribution. The reader's weakest-assumption (golden-score transfer from SLM to LLM) is a valid scientific risk, and it would matter if the parity claim were established; but it is downstream of a more immediate problem: the experiment needed to measure parity and efficiency has not been run in a way that permits verification. The tables actually undercut the 1-2% statement at the top-5% ratio that Section 3.2 emphasizes, so this is not merely a missing-number issue; the evidence as presented is inconsistent with the claim. I therefore recommend the verdict remain CONDITIONAL, with the condition that a direct NUGGETS comparison and a measured efficiency number be supplied. I do not join the reader's identification of the golden-score proxy as the single weakest point, because even a perfect proxy would not rescue an unmeasured headline comparison.","tokens_in":7699,"tokens_out":5458,"duration_ms":55011,"concrete_test":"Run a controlled head-to-head on Alpaca: original NUGGETS (Llama2-7B, 1000 random predefined tasks) vs SuperNUGGETS (Opt-350m, refined 100 predefined tasks), both screening the same 52k pool and selecting top 5%, with identical fine-tuning (LR 2e-5, batch 16, 3 epochs) and Alpaca-Eval. Measure wall-clock time or GPU-hours for the full screening pipeline and report top-5% win rates over 3 seeds. The central claim stands only if the Opt-350m win rate is within 1-2% (as the paper states) of the Llama2-7B/1000-random baseline and the measured speedup is approximately 58x.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and Section 5, is that SuperNUGGETS is 1-2% less performant than NUGGETS while being 58x more efficient. No experiment in the paper measures either half of this claim against the original method. The experimental section compares filtered-data fine-tuning to full-data fine-tuning, and Table 2 compares refined-100 with random-100 and random-1000 predefined task sets, but there is no row labeled 'NUGGETS' under identical evaluation, and no wall-clock/GPU-hour/FLOP measurement appears anywhere. The 58x figure is not derived from data; the paper's stated mechanism (100 vs 1000 tasks and 7B vs 125m/350m models) would naively suggest a much larger ratio, so the actual number is unexplained. Moreover, the available numbers do not consistently support a 1-2% decrease: at the headline top-5% ratio, Table 2 gives Llama2-7B/1000-random 21.49 vs Opt-350m/100-refined 23.98 and Opt-125m/100-refined 22.11, i.e. SuperNUGGETS is better, not 1-2% worse; at top 1% the Opt-350m refined result is 15.65 vs 18.63 for random-1000. Thus the central claim is currently an assertion, not a result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SuperNUGGETS, a variant of the NUGGETS instruction-data prospecting method. SuperNUGGETS replaces the 7B-parameter data prospector with Opt-125m or Opt-350m small language models and replaces the original 1000-example random predefined task set with a refined 100-example set constructed via reward-model scoring and diversity-preserving clustering. The authors evaluate the filtered data by fine-tuning Llama2-7B on Alpaca subsets and measuring Alpaca-Eval win rates, reporting that top-5% selected data outperforms fine-tuning on the full 52k dataset, and claiming in the abstract and conclusion that SuperNUGGETS is only 1-2% less performant than NUGGETS while being 58 times more efficient.","tokens_in":7961,"tokens_out":5423,"duration_ms":58201,"significance":"If the central claims were fully supported, this would be a practically useful result: using a small model as a data prospector for a large model would substantially lower the cost of instruction-data selection while preserving quality. The observation that top-5% subsets selected by Opt-125m and Opt-350m outperform full-data fine-tuning is interesting and worth investigating. However, the headline parity and efficiency claims are currently not measured against the original NUGGETS pipeline, and the reported numbers do not consistently support the 1-2% figure. The paper does provide useful ablations showing that the refined 100-example task set outperforms a randomly sampled 100-example set and roughly matches a random 1000-example set.","major_comments":[{"comment":"The claim that SuperNUGGETS is only 1-2% less performant than NUGGETS is not supported by the reported numbers. Table 2 contains no row explicitly labeled \"NUGGETS\"; if the Llama2-7B / 1000-random row is intended as the original NUGGETS baseline, then the results vary widely by selection ratio. At top 5%, the Opt-350m and Opt-125m SuperNUGGETS rows are better than the baseline (23.98 and 22.11 vs. 21.49), while at top 10% the drops are about 7.6% and 10.2% (21.99 and 21.37 vs. 23.79), and at top 1% the drops are about 16% and 13% (15.65 and 16.15 vs. 18.63). A 1-2% decrease is therefore not a consistent reading of the table, and the paper should state explicitly which comparison supports the abstract's claim.","section":"Abstract and Section 5"},{"comment":"The 58x efficiency factor is asserted without measurement or a clear derivation. The paper's own numbers in Section 2 imply about 52,054,002 inference passes for the original method and about 5,252,202 for the 100-example refined set, a reduction of roughly 10x in inference count; combining this with the parameter ratios of 20x (Opt-350m vs. Llama2-7B) or 56x (Opt-125m vs. Llama2-7B) would suggest efficiency gains of roughly 200x or 560x, not 58x. No wall-clock, GPU-hour, or FLOP measurement appears anywhere in the paper. The efficiency claim needs either actual runtime data or a clearly stated calculation that explains the factor of 58.","section":"Section 2 and Section 5"},{"comment":"The golden score in Eq. (5) is the load-bearing mechanism of the method: it assumes that the fraction of predefined tasks for which one-shot perplexity improves over zero-shot perplexity ranks instruction examples by downstream fine-tuning value, and that this ranking transfers from the 7B prospector to the 125m/350m prospectors. The paper currently validates this only indirectly: Table 3 shows overlap between the top-30% sets selected by different prospectors, and Table 1 shows that top-5% subsets beat full-data fine-tuning. A concrete test of the transfer assumption would be to report fine-tuning win rates for subsets selected by each prospector at the same ratios (which Table 2 partially does) and, more directly, to compute the rank correlation between golden scores and downstream fine-tuning gains for a sample of examples. Without such a test, the parity claim for the small-model prospector remains an assumption rather than a demonstrated result.","section":"Section 2.2, Eq. (5)"}],"minor_comments":[{"comment":"The phrase \"high-quality quality data\" in the abstract contains a duplicated word; it should read \"high-quality data.\"","section":"Abstract"},{"comment":"The text states that filtering Alpaca requires \"inference a total of 52,002 (zero-shot) + [52,002 × 1,000] (one-shot) = 52,054,002 times\" and then says \"104 million times.\" The sum is about 52 million, not 104 million; this arithmetic inconsistency should be corrected.","section":"Section 2, Motivation"},{"comment":"The description \"encodes the first 20-10,000 data\" and later \"selects 80 examples from 20-1,000 data\" is unclear: it should specify the exact index ranges (for example, ranks 20 through 10,000, and ranks 20 through 1,000) in a consistent notation.","section":"Section 2.1"},{"comment":"The sentence \"we use an Adam optimiser with a learning rate of 2 × 10−5, a learning rate of 2e-5\" repeats the learning rate value; this should be presented once.","section":"Section 3.1"},{"comment":"The text says \"davincici -003 model,\" which contains a typo; it should be \"text-davinci-003.\"","section":"Section 3.1"},{"comment":"The phrase \"The above experimental results illustrate the validity letter of our refinement\" appears to contain an extra word; it should likely read \"the validity of our refinement.\"","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper shares a coauthor with the original NUGGETS paper, yet it does not include a direct comparison against the original NUGGETS implementation under identical evaluation conditions. Since the headline claims are about parity with NUGGETS and a 58x efficiency gain, the editor should ask the authors for either a clear baseline row or a detailed explanation of how the 58x figure was obtained. The method itself is plausible and the top-5% versus full-data result is interesting, so the issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. The core idea is reasonable: swap the 7B prospector for a 125M/350M one and shrink the predefined task set from 1000 random examples to 100 refined ones. The second is that the headline numbers in the abstract—1-2% performance drop and 58x efficiency gain—are not established by the experiments that follow.\n\nWhat is genuinely new is the two-step refinement of the task set: reward-model scoring to keep high-quality examples, then k-center greedy clustering for diversity. The ablation in Table 2 is the paper's real contribution. For every prospector, the refined-100 task set beats a random-100 set and is comparable to or better than the random-1000 set used in the original NUGGETS. That is a clean, useful result. The agreement between SLM and LLM selections (Table 3) also supports the intuition that small models can serve as proxies here. The evaluation is independent (Alpaca-Eval), so there is no circularity problem.\n\nWhere it falls short: the paper never runs the original NUGGETS under identical conditions. The closest row, Llama2-7B with 1000 random tasks, appears in Table 2 but is not labeled as NUGGETS and is not used to substantiate the 1-2% claim. Looking at the numbers yourself, the claim does not hold up. At top 5%, Opt-350m/refined-100 scores 23.98 vs Llama2-7B/random-1000's 21.49—SuperNUGGETS is better. At top 10% it is 21.99 vs 23.79, about a 1.8-point drop, and at top 1% it is 15.65 vs 18.63, a 3-point drop. That is not a uniform 1-2% decrease; it ranges from better to roughly 16% relative worse. The 58x efficiency figure is also unexplained. The paper's own arithmetic gives ~10x fewer inference calls (1000 vs 100 tasks) and 20-56x smaller models, which naively multiplies to 200-560x, not 58x. No wall-clock time, GPU hours, or FLOPs are reported. The number just appears in the abstract and conclusion.\n\nThere are also small errors: the text says 104 million inference runs when it is actually about 52 million, and the ablation section says \"the validity letter of our refinement.\" No error bars or multiple seeds are reported anywhere, which matters for fine-tuning comparisons.\n\nWho should read it: anyone working on instruction-data selection. The task-set refinement result is worth citing even if the efficiency claim is ignored. But the central comparison to NUGGETS needs to be redone with a proper baseline and a measured runtime or FLOP count before the 1-2% / 58x claims can be taken seriously. I would send it to peer review—the idea is solid enough and the ablation is informative—but with a strong request for major revision. The authors need to run NUGGETS as a real baseline, report actual efficiency metrics, and restate the performance comparison based on the table rather than the abstract.","headline":"Plausible efficiency-oriented variant of NUGGETS with a useful task-set refinement ablation, but the headline 1-2% parity and 58x efficiency claims are unsupported by the experiments as presented.","tokens_in":8574,"tokens_out":4980,"would_cite":true,"duration_ms":48604,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A small language model can rank instruction data as well as a 7B model while using 58 times less compute.","keywords":["instruction data selection","one-shot learning","small language model","data prospecting","golden score","instruction fine-tuning","perplexity","data efficiency"],"falsifier":"Take the same candidate pool, score it with Opt-125m and with Llama2-7B, then fine-tune identical copies of a fixed base model on the 5% each scorer picks and on a random 5% baseline. If the Opt-125m subset does not beat the random subset on the held-out benchmark, or if the two scorers' rankings have no better-than-chance overlap, the claim that small-model Golden Scores transfer to large-model fine-tuning would be refuted.","tokens_in":7451,"feed_emoji":"⛏️","tokens_out":6763,"duration_ms":70100,"temperature":0.7,"pith_summary":"Instruction-tuning quality depends less on data volume than on which examples you keep. The paper proposes SuperNUGGETS, a data-filtering method that scores each candidate instruction by how much it improves a small model's one-shot perplexity on a curated set of 100 test tasks. It claims this lets a 350M-parameter model (and even a 125M-parameter model) pick training data almost as well as the 7B-parameter scorer used by the original NUGGETS, losing only 1-2% in final win rate while cutting compute by a factor of 58. If true, high-quality data selection for LLM fine-tuning becomes a cheap preprocessing step rather than a large-model job. The paper also reports that fine-tuning on just the top 5% of Alpaca selected this way beats fine-tuning on all 52,002 examples.","feed_headline":"Small model picks winning training data 58x faster","feed_subtitle":"A 350M-parameter scorer matches a 7B scorer within 1-2% while cutting compute.","key_machinery":"The Golden Score (Eq. 5) is the ranking object: for a candidate instruction $z_k$, $\\mathrm{GS}(z_k)=\\frac{1}{m}\\sum_{i=1}^{m}\\mathbb{I}[s_{\\mathrm{one}}^i(z_k)>s_{\\mathrm{zero}}^i]$, the fraction of predefined test tasks whose one-shot answer perplexity improves over zero-shot when $z_k$ is prepended. The paper's second load-bearing mechanism is the refined test set: 100 tasks assembled by taking the top 20 reward-model-scored instructions and 80 diverse examples from a greedy k-center clustering, which replaces 1,000 randomly sampled tasks and makes the scoring pass ten times cheaper.","core_discovery":"The central claim is that the information needed to prospect good instruction data lives in the one-shot behavior of a much smaller model, not in the size of the scorer. Given a pool of instruction examples and a small set of predefined tasks, SuperNUGGETS computes, for every candidate example, the fraction of tasks whose answer perplexity improves when that example is prepended as a one-shot prompt (the Golden Score). It then keeps the top n% of examples by that score. The paper shows this selection transfers: fine-tuning Llama2-7B on the top 5% scored by Opt-350m reaches a win rate of 23.98 versus 18.51 for the full dataset and 24.47 for the top 5% scored by Llama2-7B itself, and the overlap between the 7B and 125M rankings is 65% at the top 30% cutoff. The method's efficiency comes from replacing the 1,000 random test tasks with a 100-task set built by reward-model scoring and diversity clustering, which cuts the number of scoring passes by another factor of ten.","pith_inferences":["Because the scoring pass is cheap, a natural extension the paper does not run is to scale the same filter to millions of candidate instructions, where an LLM scorer would be prohibitively expensive; nothing in the method prevents that scale.","The paper fine-tunes only Llama2-7B, so it does not establish that the same top-5% subset is optimal for larger or different base models; the transfer claim would be stronger if tested on a 13B or 70B fine-tune.","The success of a 100-task refined set suggests the predefined tasks can be optimized rather than sampled, and one could search over task sets to maximize agreement with an LLM's ranking, a direction the paper does not explore."],"forward_implications":["Fine-tuning on SuperNUGGETS's top 5% of Alpaca (about 2,600 examples) outperforms fine-tuning on the full 52,002-example dataset across all three prospector sizes.","A 350M-parameter prospector (20 times smaller than Llama2-7B) and a 125M-parameter prospector (56 times smaller) both recover most of the data-selection signal, with final win rates within 1-2% of the 7B prospector.","Reducing the predefined task set from 1,000 random tasks to 100 refined tasks cuts scoring computations tenfold while matching the selection quality of the larger random set.","The rankings are stable across model sizes: the top 30% selected by Opt-125m and Opt-350m overlap with Llama2-7B's selection at 65% and 70%, respectively, so a small model can stand in for a large one."],"supporting_citations":[{"why":"Proposes the original NUGGETS one-shot data-prospecting method and defines the zero/one-shot scoring baseline that SuperNUGGETS inherits.","marker":"Li et al., 2023c"},{"why":"Supplies the theoretical motivation that in-context learning approximates implicit fine-tuning, which justifies scoring by one-shot perplexity.","marker":"Dai et al., 2022"},{"why":"Provides the automatic instruction-following evaluation benchmark and the win_rate metric used in the experiments.","marker":"Li et al., 2023b"},{"why":"Describes Self-Instruct, the technique used to create the Alpaca instruction dataset being filtered.","marker":"Wang et al., 2022a"},{"why":"Supports the premise that small, carefully selected instruction sets can beat larger ones, motivating data selection.","marker":"Zhou et al., 2023"}],"fun_headline_variants":["Small model finds top data 58x faster, within 2% of big model","Data prospecting: 58x speedup with a 350M-parameter scout","Tiny model prospects data, big model reaps rewards 58x faster","Mini language model matches 7B data picker, 58x faster","58x efficiency boost: use a small model to mine instruction data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire method assumes that a small model's one-shot perplexity improvement on 100 test tasks ranks instruction examples by their downstream fine-tuning value for a much larger model.","fun_headline_variants_meta":{"raw":{"variants":["Small model finds top data 58x faster, within 2% of big model","Data prospecting: 58x speedup with a 350M-parameter scout","Tiny model prospects data, big model reaps rewards 58x faster","Mini language model matches 7B data picker, 58x faster","58x efficiency boost: use a small model to mine instruction data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000782,"raw_usage":{"total_tokens":3467,"prompt_tokens":974,"completion_tokens":2493,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":2390}},"tokens_in":590,"tokens_out":2493,"duration_ms":19143,"temperature":1.0,"reasoning_tokens":2390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:28:01.353062+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same candidate pool, score it with Opt-125m and with Llama2-7B, then fine-tune identical copies of a fixed base model on the 5% each scorer picks and on a random 5% baseline. If the Opt-125m subset does not beat the random subset on the held-out benchmark, or if the two scorers' rankings have no better-than-chance overlap, the claim that small-model Golden Scores transfer to large-model fine-tuning would be refuted.","supporting_citations":[],"review_version":1}