{"id":"72cd4081-d429-4d1a-92af-7a183aa1b1fe","arxiv_id":"2605.30288","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MIRA discovers source-specific rubrics then distills them into scalable scorers, matching full-corpus mid-training performance on code tasks while using half the tokens.","lead":"MIRA introduces a data selection method for LLM mid-training that discovers evaluation rubrics tailored to each data source group and turns them into fast student models for filtering. A smart generalist might care because mid-training is a major cost center in building capable models, and better selection could cut token usage without losing downstream performance.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest assumption matches the only plausible point of fragility (rubric stability after distillation). No additional load-bearing concern surfaces from the given material, so the provisional UNVERDICTED stance is retained.","tokens_in":1691,"tokens_out":241,"duration_ms":17181,"concrete_test":"Re-run the nine code benchmarks with the exact MIRA-selected 50% subset versus a random 50% subset of the same 21 sources; if the random subset matches or exceeds MIRA, the source-aware rubric component is not load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on MIRA's rubric discovery plus distillation producing source-aware filters that yield measurable gains over baselines while matching full-corpus performance at half the tokens. The provided abstract and reader's summary give no internal inconsistency, circularity, or unstated assumption that would falsify this on its own terms; the method description is consistent with the reported outcome. Because the full manuscript is referenced but not reproduced here, no technical flaw in equations, experimental controls, or generalization argument can be isolated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes MIRA, a source-aware data selection framework for LLM mid-training. It performs self-anchored rubric discovery per source group to define evaluation criteria, then distills these rubrics into scalable student scorers for corpus-wide filtering. On a code mid-training setup using 21 sources in 5 groups, MIRA is reported to outperform selection baselines across nine code benchmarks while matching full-corpus performance using only half the tokens.","tokens_in":1755,"tokens_out":372,"duration_ms":16921,"significance":"If the empirical results hold under rigorous controls, MIRA would address a practical gap in mid-training curation by combining scalable model-based filtering with source-adaptive semantic criteria, enabling more efficient data use without sacrificing downstream capability gains.","major_comments":[{"comment":"Abstract: the central performance claims (outperformance on nine benchmarks, parity with full corpus at half tokens) are stated without any description of the experimental setup, baselines, controls, statistical tests, or variance estimates, rendering the primary result unevaluable from the provided text.","section":"Abstract"}],"minor_comments":[{"comment":"Clarify the exact definitions of the 5 source groups and the 21 sources, including any formatting or role differences that motivate the source-aware approach.","section":null},{"comment":"Provide the list of the nine code benchmarks and the specific selection baselines against which MIRA is compared.","section":null}],"recommendation":"uncertain","confidential_remarks":"The absence of any experimental details in the abstract is unusual for a methods paper making quantitative claims; if the full manuscript does not contain a complete experimental section with controls and ablations, the submission may be premature for this venue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed review. The single major comment concerns the level of detail in the abstract; we address it directly below.","responses":[{"response":"We agree that the abstract is written at a high level and omits explicit mention of the experimental setup (21 sources in 5 groups, code mid-training), the specific baselines, controls, or variance reporting. This is a deliberate choice to keep the abstract under typical length limits while still conveying the core contribution. All of those details appear in Sections 3 (method) and 4 (experiments), including the nine code benchmarks, token budgets, and comparison to the full-corpus baseline. We are happy to revise the abstract to add one concise sentence referencing the setup and the fact that results are averaged over multiple runs if the editor prefers.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the central performance claims (outperformance on nine benchmarks, parity with full corpus at half tokens) are stated without any description of the experimental setup, baselines, controls, statistical tests, or variance estimates, rendering the primary result unevaluable from the provided text."}],"tokens_in":1220,"tokens_out":258,"duration_ms":11675,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"MIRA discovers rubrics tailored to each source group in mid-training data, then distills those into student scorers for large-scale filtering. On a code setup with 21 sources split into 5 groups, the method beats standard selection baselines on nine benchmarks and reaches similar results to the full corpus while using half the tokens.\n\nThe framing is the clearest strength. Model-based filters scale but stay implicit on quality, while fixed-rubric methods struggle with heterogeneous sources and formats. Treating rubric construction as an explicit first step that then gets compressed into cheap scorers fits the mid-training regime where data comes from many places with different roles.\n\nThe main limitation is that the abstract states the outcome without showing how the rubrics are found, what the student models look like, which baselines were used, or any controls and variance numbers. The key assumption—that the per-group rubrics pick up stable, generalizable quality signals—remains untested in the provided text. If the full paper has solid ablations and reproducible details, that assumption can be checked; right now it is just asserted.\n\nThis is aimed at groups running large LLM mid-training runs who need to cut tokens without losing downstream code performance. It is not a foundational advance but a targeted engineering fix. The work shows clear thinking about the problem constraints and deserves peer review so the experiments can be examined directly.","headline":"MIRA's rubric discovery plus distillation approach for source-aware mid-training data selection looks practically useful on code data but the abstract supplies no experimental details to back the performance claims.","tokens_in":2277,"tokens_out":356,"would_cite":false,"duration_ms":17103,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"MIRA discovers source-specific rubrics then distills them into student scorers to filter mid-training data, matching full-corpus results on code benchmarks while using half the tokens.","keywords":["data selection","mid-training","LLM training","rubric discovery","source-aware filtering","code data","model distillation","benchmark evaluation"],"falsifier":"Running the distilled scorers on the same 21 sources but with a different downstream task family (for example, math rather than code) and observing that the selected half-corpus no longer matches full-corpus performance on the new tasks would falsify the central claim.","tokens_in":2601,"feed_emoji":"","tokens_out":738,"duration_ms":17523,"temperature":0.7,"pith_summary":"The paper aims to solve data selection for mid-training, where large heterogeneous mixtures must be curated under a pretraining objective yet aimed at downstream capabilities. Existing approaches either scale via implicit model signals or rely on fixed rubrics that do not adapt to varying source formats. MIRA instead treats rubric construction as part of selection: it first identifies what quality criteria matter for each source group, then compresses those judgments into lightweight scorers that can label the entire corpus. A sympathetic reader would care because this produces source-adaptive semantic filtering without assuming standardized data or permanent rubrics, and the reported experiments show the resulting subset matches the performance of the full mixture on nine code benchmarks.","feed_headline":"MIRA matches full mid-training results with half the tokens","feed_subtitle":"Source-group rubric discovery lets the method filter 21 heterogeneous code sources into a half-size mixture that still hits nine benchmarks.","key_machinery":"self-anchored rubric discovery: the process of first determining source-group-specific evaluation criteria and then distilling those criteria into scalable student scorers for corpus-wide filtering.","core_discovery":"MIRA is a source-aware filtering framework that performs self-anchored rubric discovery: for each of five source groups drawn from 21 code sources, it identifies the evaluation criteria that matter for that group, then distills those judgments into student scorers that label the full corpus at scale. In code-oriented mid-training experiments this procedure yields a filtered mixture that outperforms prior selection baselines across nine benchmarks while matching the full-corpus run with only half the tokens.","pith_inferences":["The method could be tested on non-code domains by repeating the rubric-discovery step on math or general-text sources to check whether the same token-reduction benefit appears.","If the distilled scorers remain effective after further compression, the approach might reduce the teacher-model cost of data selection itself.","Applying the same anchoring step at the pretraining stage rather than only mid-training would test whether source-aware rubrics help earlier in the pipeline."],"forward_implications":["Mid-training mixtures can be reduced to half their original token count without loss of downstream code performance.","Source-group rubrics allow semantic filtering to scale to heterogeneous data without assuming fixed evaluation criteria.","Distilled student scorers provide explicit quality signals that outperform both purely model-based and fixed-rubric baselines on nine code benchmarks.","The same two-stage discovery-plus-distillation pipeline can be applied whenever mid-training data come from multiple sources with distinct formats and roles."],"fun_headline_variants":["MIRA halves mid-training tokens while matching full benchmarks","Rubric discovery lets MIRA filter 21 sources to half size","MIRA's source-aware rubrics match full results at half scale","Self-anchored rubrics enable MIRA to halve code mid-training data"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The rubrics found for each source group capture stable quality signals that survive distillation into student scorers and continue to work for the full corpus.","fun_headline_variants_meta":{"raw":{"variants":["MIRA halves mid-training tokens while matching full benchmarks","Rubric discovery lets MIRA filter 21 sources to half size","MIRA's source-aware rubrics match full results at half scale","Self-anchored rubrics enable MIRA to halve code mid-training data"]},"model":"grok-4.3","cost_usd":0.003797,"raw_usage":{"total_tokens":1956,"prompt_tokens":659,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":37974500,"prompt_tokens_details":{"text_tokens":659,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1226,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":659,"tokens_out":71,"duration_ms":9737,"temperature":1.0,"reasoning_tokens":1226,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:06:50.308909+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the distilled scorers on the same 21 sources but with a different downstream task family (for example, math rather than code) and observing that the selected half-corpus no longer matches full-corpus performance on the new tasks would falsify the central claim.","supporting_citations":[],"review_version":1}