{"id":"824b7974-8b54-4d16-9953-c7313243885d","arxiv_id":"2607.28109","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Book-level organization of retrieval-grounded synthetic textbooks improves mid-training by about +1 point over content-, length-, and rephrase-matched controls.","lead":"Organizing synthetic textbook text into coherent books, not just rewriting passages, improves language-model mid-training. Controlled swaps show packaging and planned structure beat matched content, length, and local rephrasing.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Single-run mid-training leaves the ~1-point Full gaps without run-level error bars; that is still the load-bearing soft spot.","rationale":"The paper’s strongest empirical move is the matched control suite (Split, RandomConcat, Rephrase) plus a second-architecture check, not a single Natural-Books replacement. Those designs make content, length, and local rewriting unlikely sole drivers. The remaining load-bearing gap is exactly what the reader flags: each condition is trained once, so ~1-point mean effects lack run-level uncertainty, which the authors acknowledge. I do not find a deeper internal inconsistency in the packaging logic or a stronger confound that overturns the controls; proprietary corpus and incomplete public artifacts hurt reproducibility but do not by themselves falsify the reported ordering. Hence the reader’s CONDITIONAL verdict stands—accept-shaped evidence contingent on multi-seed confirmation—without upgrade or downgrade from this pass.","tokens_in":28005,"tokens_out":579,"duration_ms":29862,"concrete_test":"Re-run Full and Split (primary MoE recipe, same 200B mix and 16B book slice) for ≥3 independent random seeds each, holding all other knobs fixed; report run-mean ± run-SD of the 28-benchmark mean and of Full−Split. If the mean gap stays ≥~0.5 and the run-level intervals separate in Full’s favor, the claim holds; if the gap collapses inside run noise or reverses on ≥1 seed, the organization attribution weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attributes ~+1 mean gains (Full−Split +1.02, Full−RandomConcat ~+1.01, Full−Rephrase +1.17, Full−Natural +1.09; Llama3-8B same order) to book-level organization. The controls are content-, length-, and retrieval-pool-matched and are carefully specified (§4–5, Fig. 2). What is least secure is causal attribution under n=1 training run per condition: optimizer, schedule, data order, and the 8% book slice are fixed, but §8 and Appendix G explicitly state training-run variance is unmeasured and that bootstrap/sign tests only resample benchmarks. With effects on the order of 1 point on a 28-benchmark mean (and large single-benchmark swings, e.g. GSM8K Full 88.02 vs Split 81.64), seed/initialization/packing-order noise is a live alternative explanation. Transfer to Llama3-8B reduces architecture-specificity risk but does not supply run-level replication. If run noise is comparable to the reported deltas, the packaging/organization story is not yet isolated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that book-level organization is a distinct, useful axis for synthetic textbook data in mid-training, beyond content quality or local rewriting. It presents a retrieval-grounded pipeline (taxonomy-guided retrieval, clustering, hierarchical TOC planning with a quality gate, and source-grounded section assembly) that yields 686K textbooks (32B tokens) across 15,000+ disciplines. In a fixed mid-training mix, replacing natural books with this Full corpus improves a 28-benchmark mean by +1.09. Matched controls isolate packaging (content-identical Split: Full +1.02), length (RandomConcat remains below Full), and structured synthesis (retrieval-pool-matched Rephrase: Full +1.17). The Full > RandomConcat > Natural Books ordering also holds on Llama3-8B. Component ablations attribute complementary roles to hierarchical generation, retrieval grounding, and cluster-informed TOCs.","tokens_in":28320,"tokens_out":1313,"duration_ms":33509,"significance":"If the result holds, the paper cleanly elevates document organization from an incidental packaging choice to a first-class design variable for synthetic pre-training data, with practical implications for mid-training mixtures. Strengths include carefully constructed content-, length-, and retrieval-pool-matched controls (Fig. 2; §4–5), transfer across a 3B-active MoE and Llama3-8B, decontamination reporting (~0.07% flagged tokens), fixed mixture budgets and intra-document masking, and planned release of synthesis/control code plus a research-licensed corpus subset. The control taxonomy (Full/Split/RandomConcat/Rephrase) is a reusable experimental template for future data-design work.","major_comments":[{"comment":"§5 Table 3 and §8 / Appendix G: each condition is a single mid-training run with fixed order and optimizer; training-run variance is explicitly unmeasured, and bootstrap/sign tests only resample benchmarks. Reported overall gaps are ~1 point (Full−Split +1.02, Full−RandomConcat ~+1.01, Full−Rephrase +1.17), while individual swings are large (e.g., GSM8K 88.02 vs 81.64). Without at least a second seed, multi-checkpoint comparison, or a quantitative noise bound, causal attribution of these gaps to book organization versus run noise remains incompletely secured. This is the main load-bearing soft spot for the central claim; either limited replication or a stronger, explicit sensitivity analysis should be added before treating the deltas as definitive.","section":"§5 Table 3; §8; Appendix G"},{"comment":"Table 3, Reasoning block: the Full−Split reasoning mean gain (+1.35) is dominated by GSM8K (+6.38); the other four reasoning tasks are mixed (−0.07, +1.23, −0.30, −0.50). Category and overall means are unweighted, so one outlier can materially shape the narrative that packaging helps “reasoning.” Please report leave-one-benchmark-out overall/category deltas (or median deltas) and discuss whether the packaging story remains intact without GSM8K, so the claim is not over-read from a single task.","section":"§5 Table 3 (Reasoning / GSM8K)"},{"comment":"§4 Rephrase control: the comparison is central to “structured synthesis beyond local rewriting,” but the manuscript does not fully specify how Rephrase documents are length-budgeted, packed, and sampled to the same 16B-token book slice, nor how audience×style is applied at document (vs book) granularity while matching the retrieval pool. Small mismatches in effective document-length distribution or style coverage could confound Full’s +1.17. A short protocol paragraph (and, if possible, length-distribution overlay vs Full) would make this control as airtight as Split/RandomConcat.","section":"§4 (Rephrase); Figure 2"}],"minor_comments":[{"comment":"Figure 3 caption and §4: clarify whether the 50K-token TOC target is training tokens after tokenization of the final English text, and how that maps to the reported character counts in the running example (~257K characters).","section":"§4; Figure 3"},{"comment":"Table 2: high-school Wiki pass rate (60.2%) is much lower than other cells; a one-sentence note on whether failed high-school plans are simply dropped (reducing that audience’s share) would help interpret corpus composition.","section":"§4 Table 2"},{"comment":"§6.2 / Table 9: intrinsic ablations use Gemini-3.1-Pro on 100 books; state whether the judge saw condition labels or was blinded, to bound demand-characteristic risk in diagnostic scores.","section":"§6.2; Appendix H"},{"comment":"Related Work: ACER, LiteLong, and Cosmopedia v2 are well cited; a tighter sentence contrasting “post-hoc long-context packing of related docs” (SPLiCe / In-Context Pretraining) with “generation-time TOC-planned books” would sharpen the positioning already sketched in §2.","section":"§2"},{"comment":"Minor polish: abstract and intro repeat the control deltas almost verbatim; condensing one occurrence would improve flow. Also fix “Maxm Pan” if that is a typographical error in the author list.","section":"Title page; Abstract"}],"recommendation":"minor_revision","confidential_remarks":"Solid empirical systems paper with unusually careful controls for this area; I would not reject on the single-run issue alone given community norms, but the editor should expect the authors to address GSM8K concentration and Rephrase protocol clarity. Fit is good for a data-centric ML / LM training venue. No integrity red flags; limitations are disclosed in good faith."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the control design, not the pipeline scale. They hold generated text fixed (Full vs Split), length fixed (RandomConcat), and retrieval pool plus audience×style fixed (Rephrase), then show Full wins by about a point on a 28-benchmark mean, with the same ordering on Llama3-8B. That isolates document packaging and structured synthesis better than most synthetic-data papers I’ve seen.\n\nWhat’s actually new is treating book-level organization as a first-class axis and measuring it under matched mid-training, not inventing synthetic textbooks or TOC prompts. Phi/Cosmopedia/WRAP/ACER/LiteLong/SPLiCe are cited and used as backdrop. The Full/Split/RandomConcat/Rephrase taxonomy is the contribution. Pipeline is competent: retrieval → cluster → TOC gate → grounded sections, 686K books, decontamination, intra-document masking, fixed 8% book slice. Component ablations are diagnostic and labeled as such. Limitations section owns the single-run issue.\n\nSoft spot is real but proportionate: one run per condition, no seed variance, ~1-point mean effects with big single-benchmark swings (GSM8K). Bootstrap/sign tests only resample benchmarks. Transfer to Llama reduces architecture worry; it does not replace multi-seed. Free parameters (K, k, book length, mix fraction) are fixed from pilots—fine if you treat this as a recipe study, not a universal law. Reproducibility is limited by proprietary corpus; they promise code and a subset.\n\nMath is not the load-bearing part; empirical design is. Citations look honest. No circularity on the main claim.\n\nThis is for people who build mid-training mixes and synthetic corpora. I’d bring it to reading group as a methods paper on data packaging. It deserves peer review; I’d want multi-seed confirmation and clearer public artifacts, not a desk reject. Engage if you care about how documents are assembled, not only what text is written.","headline":"Clean controlled evidence that book packaging and planned adjacency matter for synthetic textbooks—real contribution, single-run ~1-pt gaps are the main caveat.","tokens_in":28977,"tokens_out":503,"would_cite":true,"duration_ms":13950,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Organizing synthetic textbooks as coherent books—not just rewriting the same content—improves mid-training by about one point on a 28-benchmark mean.","keywords":["synthetic textbooks","mid-training","book-level organization","retrieval-grounded generation","TOC planning","document boundaries","data curation","language model pre-training"],"falsifier":"Retrain Full, Split, RandomConcat, and Rephrase several times with different seeds under the same mix and show that the ~1-point Full advantage over Split/RandomConcat/Rephrase disappears or reverses once run variance is estimated.","tokens_in":28845,"feed_emoji":"📚","tokens_out":943,"duration_ms":20591,"temperature":0.7,"pith_summary":"This paper argues that the value of synthetic textbook data is not only what is written, but how related material is packaged into book-length documents. The authors build a retrieval-grounded pipeline that clusters corpus passages by topic, plans hierarchical tables of contents, and writes source-grounded sections into complete books—686K textbooks and 32B tokens across more than 15,000 disciplines. Replacing the natural-book slice of a mid-training mix with this corpus raises average downstream scores by +1.09. Matched controls then show why: keeping the same text but splitting each section into its own document costs about a point; randomly concatenating sections to match book length does not recover the gain; and independently rephrasing the same retrieved pool without clustering or TOC planning also falls short. The same Full-over-RandomConcat and Full-over-Natural ordering appears on Llama3-8B, so the claim is that book-level organization is a real design axis for synthetic pre-training data.","feed_headline":"Book-shaped synthetic textbooks beat rewritten snippets","feed_subtitle":"Same content, split into sections, loses ~1 point; random long packs and plain rephrases do too","key_machinery":"The Full setting: a five-stage pipeline that retrieves corpus material, clusters it into topical units, plans hierarchical audience×style TOCs with a quality gate, generates source-grounded sections, and assembles them into complete training documents—so planned adjacent sections share continuous positions and intra-document attention.","core_discovery":"Book-level organization adds value beyond generated content, document length, and local rewriting. Holding text and tokens fixed, training on complete TOC-planned books (Full) beats treating each section as an independent document (Split) by +1.02 mean; a length-matched scramble of sections from different books (RandomConcat) stays near Split and below Full; and a retrieval-pool-matched per-document rewrite (Rephrase) trails Full by +1.17. Replacing natural books with Full yields +1.09 overall, with transfer to Llama3-8B.","pith_inferences":["If planned adjacency is the active ingredient, other long-form synthetic genres (lab manuals, legal treatises, codebases) may gain from the same cluster–TOC–assemble pattern.","Sequence-packing research that only concatenates related web docs may understate the benefit of generating the long document under one global plan rather than packing after the fact.","A cheap practical test for other labs: keep generation fixed and only toggle whether sections are one document or many—if the gap vanishes, packaging was not the driver in that stack."],"forward_implications":["Synthetic data design should treat document packaging and planned section order as first-class knobs, not only style or factual content.","Mid-training book slices can be improved by replacing natural books with TOC-planned, retrieval-grounded textbooks at fixed token budget.","Length-matched packing of unrelated sections is not a substitute for keeping planned book boundaries during training.","Local rewriting of retrieved documents is insufficient; clustering plus hierarchical TOC planning is needed for the reported gains.","The organization effect is expected to transfer across architectures when content and length controls are held fixed."],"fun_headline_variants":["Book-level packing beats same-text section splits by +1.02","TOC-planned synthetic books top length-matched random concats","Full books beat retrieval-matched rephrases by +1.17 mean","Organization, not just rewrite style, lifts mid-training gains","Replacing natural books with Full synthetic yields +1.09"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That one mid-training run per condition is enough to credit roughly one-point mean gaps to book organization rather than ordinary training-run noise, which the paper does not measure.","fun_headline_variants_meta":{"raw":{"variants":["Book-level packing beats same-text section splits by +1.02","TOC-planned synthetic books top length-matched random concats","Full books beat retrieval-matched rephrases by +1.17 mean","Organization, not just rewrite style, lifts mid-training gains","Replacing natural books with Full synthetic yields +1.09"]},"model":"grok-4.5","effort":"low","cost_usd":0.002342,"raw_usage":{"total_tokens":1000,"prompt_tokens":860,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":23424000,"prompt_tokens_details":{"text_tokens":860,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":66,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":860,"tokens_out":74,"duration_ms":2327,"temperature":1.0,"reasoning_tokens":66,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T17:35:21.266866+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain Full, Split, RandomConcat, and Rephrase several times with different seeds under the same mix and show that the ~1-point Full advantage over Split/RandomConcat/Rephrase disappears or reverses once run variance is estimated.","supporting_citations":[],"review_version":1}