Pith. sign in

REVIEW 3 major objections 3 minor 103 references

Soofi S, a fully open German–English model activating 3.2 of its 31.6B parameters per token, claims the top aggregates among fully open models in its comparison and matches dense 14–27B rivals while serving long contexts 8–9x faster.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:36 UTC pith:UHY7DV6U

load-bearing objection An unusually honest, well-disclosed German-English pretraining report whose headline benchmark claims are conditional on a contamination audit that covers only the QA-base slice and never screens the web or MT-German channels most likely to hide eval items. the 3 major comments →

arxiv 2607.09424 v3 pith:UHY7DV6U submitted 2026-07-10 cs.CL cs.AIcs.LG

A Sovereign, Open-Source Foundation Model for German and English

classification cs.CL cs.AIcs.LG
keywords mixture of expertshybrid Mamba-TransformerGerman-English modelsovereign AIopen-source LLMpretraining data transparencylong-context inferencebenchmark contamination
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that a sovereign, genuinely open foundation model can sit on the same capability-per-active-parameter frontier as the strongest international releases. Soofi S 30B-A3B activates only ~3.2 of its ~31.6 billion parameters per token, and because 23 of its 52 layers are recurrent Mamba-2 layers rather than attention, its per-sequence cache stays near-constant as context grows — which the authors measure as an 8–9x decode-throughput advantage over dense 14–27B models at 40K context. Trained on ~27 trillion tokens with German deliberately up-weighted (7.2% in the diverse phase, 15.3% in the annealing phase), it posts the highest English and German aggregate scores among fully open models in the comparison and matches or beats every European sovereign baseline on its German suite. The paper also discloses that GPQA evaluation items leaked into the training data, removes that benchmark from all aggregates for every model, and documents the corrective safeguards. If the results hold, other language communities gain a fully documented, rebuildable template for capable and efficient models beyond English.

Core claim

The paper claims that Soofi S 30B-A3B — a fully open, sovereign German–English base model trained end-to-end on a German HPC cloud — is the strongest fully open model in its comparison on both English and German benchmarks, matches dense 14–27B international models on aggregate performance (English 77.3, German 85.3, both excluding the leaked GPQA benchmark) while activating only 3.2 of its 31.6 billion parameters per token, and matches or outperforms every European sovereign baseline in the comparison on every German benchmark in its suite. The near-constant inference cache is the mechanism: 23 Mamba-2 layers carry most sequence mixing with a fixed-size state, so only 6 attention layers acc

What carries the argument

The carrying mechanism is the hybrid Mamba–Transformer MoE stack: 52 layers interleaving 23 Mamba-2 sequence-mixing layers (fixed-size recurrent state), 23 sparse MoE layers with 128 routed and 2 shared experts (6 active per token), and 6 Grouped-Query-Attention layers, which are the only layers maintaining a key–value cache. This yields ~3.2B active parameters per token and an incremental cache footprint of ~6 KB per token per sequence — 11–53x smaller than dense comparators — which is what keeps decode throughput flat as context grows. Around this sits a three-phase Warmup–Stable–Decay curriculum: ~20T tokens of diverse quality-tiered pretraining, ~6.6T of high-quality annealing in which G

Load-bearing premise

The capability claims rest on the completeness of the contamination audit in Section 4.3: if any evaluation item from a reported benchmark — especially a machine-translated -DE variant — remains in the ~27T-token mixture, the headline aggregates measure memorization rather than capability; the audit was conducted by the training team only after outsiders discovered the GPQA leak, and the n-gram screening safeguard was added for future runs, not applied to this one.

What would settle it

Have an independent team screen the released per-source data accounting and corrected QA-base dataset against the evaluation items of every benchmark in the reported suites — English and German, including the machine-translated -DE variants — using n-gram and paraphrase-resistant overlap. A single remaining training-data copy of any reported benchmark's test item would falsify the aggregates as capability measurements. Separately replicate the serving protocol (batch 32, TP=1, single B200, latency-subtraction formula, 4K–256K contexts): failure to reproduce the ~4.8k aggregate decode TPS/GPU a

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • German capability can be bought with data allocation: raising German to 15.3% of the annealing mixture lifts the German aggregate by 4.6 points over the architecture-identical reference while the English aggregate rises 0.6 points, showing bilingual depth need not trade away English.
  • Fully open releases — weights, per-source data accounting, hyperparameters, training and evaluation code — can reach the capability-per-active-parameter frontier of weight-only international releases, giving other communities a rebuildable template rather than a checkpoint.
  • At high concurrency and long context, serving cost tracks cache size and memory bandwidth more than parameter count: the design sustains ~4.8k decode tokens/second/GPU at 40K context, with throughput essentially flat from 4K to 256K and a window extended to 1M tokens.
  • A model can be built end-to-end on sovereign European infrastructure (~253,000 GPU-hours for the ~27T-token run) without relinquishing benchmark competitiveness, addressing deployment under local data-protection rules.
  • Contamination is a measurable, correctable failure: the incident report shows name-based split selection can leak benchmark items into training, trajectory monitoring does not reveal such leaks, and removing the affected benchmark for all models symmetrically preserves the relative rankings.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The architecture-identical comparison isolates the data recipe as the transferable asset; another language community could plausibly apply the same three-phase, native-language up-weighting curriculum to its own language pair and reproduce gains of the same shape without new architecture work.
  • Because the audit was reactive — completed only after external discovery — and the n-gram screening safeguard applies to future runs, the reported aggregates are best read as provisional upper bounds on capability until an independent overlap check clears every reported suite, including the machine-translated -DE items.
  • The near-constant-cache result reframes how models should be reported: two models with equal benchmark scores can differ by an order of magnitude in serving cost, so aggregate scores alone understate the deployment value of hybrid architectures at long context.
  • The acknowledged long-context weakness — collapse on common-word extraction beyond 32K, diagnosed as a data-mixture gap rather than a backbone limit — makes a concrete next step available: adding retrieval- and aggregation-style synthetic data in the 32K–1M window should close the gap while leaving the rest of the long-context profile intact.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents Soofi S 30B-A3B, a 31.6B-parameter mixture-of-experts hybrid Mamba-Transformer base model with ~3.2B active parameters, pretrained on roughly 26.68T tokens with deliberately up-weighted German. The authors document a three-phase curriculum (20T diverse pretraining, ~6.6T high-quality annealing, ~0.1T long-context extension), release full per-source token accounting, training hyperparameters, intermediate checkpoints, and evaluation code, and evaluate the model against 15 (elsewhere 16 or 17) open and open-weight baselines. The central claims are that Soofi S is the strongest fully open model in the evaluation on English and German aggregates, matches dense 14–27B international models on aggregate performance at a fraction of the active parameter cost, achieves best-in-comparison code aggregates, and sustains 8–9× the aggregate decode throughput of dense baselines at 40K context. The paper also discloses a benchmark-contamination incident involving GPQA in the QA-base pretraining constituent and describes a remediation that removes GPQA from reported aggregates and adds forward-looking screening practices.

Significance. If the capability claims survive scrutiny, this is a significant contribution: a fully documented, sovereign, German-English pretraining run with unusually complete data accounting, an architecture-identical baseline that cleanly isolates the data recipe, and a credible serving-efficiency advantage from the hybrid Mamba-MoE design. The release of weights, selected checkpoints, exact per-source token counts, hyperparameters, training/evaluation code, and even a discarded final-annealing stage is exemplary for reproducibility. The long-context weakness on common-word extraction is disclosed honestly. However, the headline capability claims are entirely benchmark-based, and the paper's own contamination disclosure establishes that benchmark material entered the training mixture through at least one pathway. The completeness of the contamination audit is therefore load-bearing for the central claims.

major comments (3)
  1. [Section 4.3 (GPQA Contamination Disclosure)] The audit's scope is a load-bearing limitation. The re-audit explicitly covers 'all QA-base constituents' — roughly 3.3B tokens, about 0.05% of the Phase 2 pool (Section 3.3) — and only the failure mode of evaluation-only benchmarks whose sole published split is mislabeled 'train'. The remaining ~99.95% of the ~26.68T-token corpus, including the deliberately up-sampled English web tiers (Nemotron-CC, 11.6T effective tokens in Phase 1) and the 571B-token German machine translation of ClimbMix, is not screened against the evaluation suite. This matters doubly for the German benchmarks: the -DE evaluation sets are German renderings of English items, so MT-German web text containing the underlying English items is a direct near-duplicate channel to the German eval sets that carry the flagship +5.9 German aggregate margin. The paper's own Figure 15 concedes that paraphrased contamination is i
  2. [Section 3.3 and Tables 4–5] The manuscript states that QA-base contains 'paraphrased training splits of 25 standard NLP benchmarks in English and German', and that the model trained on QA-base. The evaluation suite includes benchmarks from the same families (code, math, QA, knowledge). Training on paraphrased train splits of a benchmark family can inflate downstream scores on that family even when no evaluation item is duplicated. The contamination disclosure in §4.3 addresses only evaluation-set leakage of four specific datasets, not this broader train-split exposure. The paper does not list the 25 benchmarks, nor does it analyze which of the reported English or German eval tasks have train-split overlap with QA-base. Because several of the largest reported margins are on code and math tasks (HumanEval +10.8, MBPP-DE +13.4, Minerva +24.2), this is potentially a direct confound for the 'strongest fully open model'
  3. [Section 4.3, 'Remediation'] The n-gram screening of final training mixtures against the evaluation suite is described as a forward-looking practice ('final training mixtures are screened ... before training'), not as a result for this run. Removing GPQA from the reported aggregates is a necessary correction but does not repair the possibility that other evaluation material entered through unscreened web or MT-German data, nor does it address the train-split exposure identified above. The claims in the Contributions section and Conclusion — 'strongest fully open model', 'matches dense 14–27B models', 'first European sovereign model to sit on the same capability-per-active-parameter frontier' — are therefore stronger than the evidence currently supports. A revision should either supply screening results for this run or substantially qualify these claims, e.g., by stating that they hold 'barring undetected contaminati
minor comments (3)
  1. [Abstract and Conclusion] The number of comparison models is inconsistent: the Abstract says 'among 17 open base models', Section 4 says 'against 15 open-source and open-weight base models', and the Conclusion says 'unified evaluation of 16 open base models'. Please reconcile these counts.
  2. [Section 4.4, Eq. (1)] Equation (1) subtracts t(1) from t(1024), which removes prefill cost only if the per-token decode time is approximately linear in output length. The paper should state this linearity assumption explicitly, since the 'TTFT-like' t(1) values are reported separately and the aggregation of prefill and decode into a single TPS figure may be sensitive to the chosen output-length range.
  3. [Section 2.2 and Appendix E] The long-context comparison with Nemotron 3 Nano is a strength, but the RULER CWE collapse beyond 32K is a substantial capability gap that is only visible in Appendix E. Consider foregrounding this limitation in the main text rather than only in an appendix.

Circularity Check

0 steps flagged

No circular steps: the paper's central claims are external measurements and controlled comparisons, not derivations that reduce to their own inputs; the contamination-audit scope is an evaluation-validity risk, not a construction-level circularity.

full rationale

The paper's central claims are empirical measurements on a released model, not predictions derived from fitted parameters. The architecture is adopted without modification from an external reference (Nemotron 3 Nano), and the serving-efficiency numbers are measured under the described latency-subtraction protocol (Eq. 1) against external baselines. The German-data proxy ablation in Appendix C is a design input that informed the final mixture; the final checkpoint's benchmark scores are reported as outcomes, not as predictions of the proxy model, and the proxy and final evaluation are not conflated into a single fitted quantity. The GPQA contamination section is a disclosure and re-audit rather than a derivation: the paper removes the contaminated benchmarks (GPQA-Diamond and GPQA-Diamond-DE) from all reported aggregates, so the headline comparisons are recomputed symmetrically without them. The claim that only four QA-base constituents share the split-name failure mode is a factual audit statement about data provenance, not an equation that reduces to its inputs; it could be wrong, but its wrongness would be an empirical error, not circularity. The audit's narrow scope — covering QA-base but not screening the web and MT-German corpus against the full evaluation suite for this run — is a genuine and load-bearing evaluation-validity risk for the benchmark-based claims, and should be weighed as a correctness concern, but it does not make the benchmark results equivalent to the training inputs by construction. Self-citations such as JQL, KletterMix, Propella, and Modalities are data-source or implementation references; they are not used to justify the headline capability comparisons. Therefore no circular step is established under the required evidentiary standard.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The central claims rest on unmeasured premises: (1) cross-model comparability of harness measurements; (2) completeness of the self-performed contamination audit; (3) validity of the translated German benchmarks; (4) representativeness of the single-GPU throughput protocol. Free parameters — mixture shares, epoch multipliers, phase budgets, suite membership — are chosen by hand or proxy ablation, not derived. There is no mathematical derivation to audit; the claims are empirical and inherit the uncertainty of benchmark and throughput measurements.

free parameters (4)
  • German share of pretraining mixture = 7.2% (Phase 1), 15.32% (Phase 2)
    The defining design choice; selected via 100B-token proxy ablations (Appendix C) on bits-per-byte and rank-choice accuracy. The headline German aggregates are outcomes of this ablation-chosen allocation.
  • Per-source epoch multipliers = 0-10 epochs per source (web HQ x3, Specialized x5, QA-base x10)
    Epoch multipliers set effective token counts and therefore the mixture; chosen by hand per source (Tables 6, 8). The QA-base x10 multiplier put contaminated GPQA items through ten training repeats.
  • Phase token budgets = 20T stable / 5T decay / 1.58T constant / 0.30T discarded / 0.10T long-context
    Chosen by hand following WSD practice; the discarded 0.30T final-annealing stage shows budget choices were evaluated empirically rather than derived.
  • Evaluation-suite membership for headline aggregates = 77.3 EN / 85.3 DE aggregates; LBPP excluded from code aggregates; GPQA + held-out group withdrawn
    The 'strongest fully open model' ranking is an average over an author-chosen suite; benchmark groups were removed post-hoc (Section 4.3), and small margins (e.g., +0.6 English aggregate over Nemotron) are sensitive to suite composition.
axioms (4)
  • domain assumption Benchmark scores measured with lm-evaluation-harness under identical prompts and few-shot settings are directly comparable across base models with different training corpora.
    Section 4: every comparative claim rests on this assumption; it ignores differential training-data overlap with benchmarks, which the GPQA incident (Section 4.3) shows is non-zero for this model.
  • ad hoc to paper The post-hoc contamination audit is complete: only GPQA, TruthfulQA, BLiMP, and Inverse Scaling leaked evaluation material into training.
    Section 4.3: self-performed after external community members discovered the GPQA leak; the central benchmark claims are invalid if other evaluation items are in the mixture. The n-gram screening safeguard was added for future runs, not this one.
  • domain assumption The machine-translated German benchmark variants (-DE) are valid, leakage-free measures of German capability.
    Section 4: the German aggregate (85.3) and the +5.9 margin over Apertus rest substantially on HellaSwag-DE, PIQA-DE, ARC-Challenge-DE, and code -DE variants; their construction, translation provenance, and screening against the training corpus are not documented.
  • domain assumption Equation 1's latency-subtraction protocol adequately isolates aggregate decode throughput from prefill cost on a single B200 at batch 32.
    Section 4.4: the 8-9x decode advantage over dense models is computed as B(O_long - O_short)/[t(O_long) - t(O_short)]; validity depends on prefill removal and on TP=1, fixed-256K being representative of production serving.

pith-pipeline@v1.3.0-alltime-deepseek · 52022 in / 25161 out tokens · 257100 ms · 2026-08-02T07:36:23.480411+00:00 · methodology

0 comments
read the original abstract

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.

Figures

Figures reproduced from arXiv: 2607.09424 by Abbas Goher Khan, Alexander L\"oser, Alex Jude, Andreas Hotho, Bj\"orn Pl\"uster, Daniil Gurgurov, David Fitzek, Jan Pfister, Jan Plogsties, Joachim k\"ohler, J\"org Bienert, Kristian Kersting, Lukas Helff, Markus Frey, Maurice Kraus, Maximilian Idahl, Max L\"ubbering, Mehdi Ali, Michael Fromm, Nicolas Flores-Herr, Patrick Putzky, Richard Rutmann, Ruben H\"arle, Sebastian Sztwiertnia, Sebastian von Rohrscheidt, Simon Gottschalk, Simon Ostermann, Soofi-Team: Benedikt Droste, Timm Ruland, Tom R\"ohr, Wolfgang Nejdl.

Figure 1
Figure 1. Figure 1: Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (a) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamon… view at source ↗
Figure 1
Figure 1. Figure 1: English and German aggregate performance versus decode throughput. Soofi S obtains the highest English and German aggregate scores among the fully open base models in this comparison. Soofi S and Nemotron 3 Nano, which share the same hybrid Mamba–MoE architecture, reach the highest measured decode throughput at 4.8k TPS/GPU. The left panel reports the mean of our English benchmark suite. The right panel re… view at source ↗
Figure 2
Figure 2. Figure 2: Training dynamics over the full ∼27T-token run. The quantity is plotted against the number of consumed tokens (in trillions); the pretraining-to-annealing transition occurs at ∼20T tokens. The solid line is a rolling median over 1,000 steps and the faint trace is the raw per-step signal. by existing software from the day of release, without bespoke integration work by downstream users. Second, serving effi… view at source ↗
Figure 3
Figure 3. Figure 3: Effective-token mixture across the three training phases. A single flow diagram tracing seven data categories (English Web, Academic & Wiki, SFT, Reasoning, Code, Math, and German) from left to right across the phases. Phase 1 (diverse pretraining), 23,051.13B effective tokens. Phase 2 (high-quality annealing), 6,303.0B effective tokens, showing increased density of skill-oriented and German data relative … view at source ↗
Figure 4
Figure 4. Figure 4: Evaluation overview for the open-source comparison. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗
Figure 4
Figure 4. Figure 4: Evaluation overview for the open-source comparison. [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Code generation results (pass@1) against large open-source models on English and German [PITH_FULL_IMAGE:figures/full_fig_p017_5.png] view at source ↗
Figure 5
Figure 5. Figure 5: Code generation results (pass@1) against large open-source models on English and German [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Mathematics results against large open-source models on English and German bench [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: Mathematics results against large open-source models on English and German bench [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Knowledge (left) and reasoning/science (right) benchmarks against large open-source [PITH_FULL_IMAGE:figures/full_fig_p020_7.png] view at source ↗
Figure 7
Figure 7. Figure 7: Knowledge (left) and reasoning/science (right) benchmarks against large open-source [PITH_FULL_IMAGE:figures/full_fig_p021_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: German benchmark results against large open-source models. Soofi S ranks first on the [PITH_FULL_IMAGE:figures/full_fig_p021_8.png] view at source ↗
Figure 8
Figure 8. Figure 8: German benchmark results against large open-source models. Soofi S ranks first on [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Base model evaluation overview for Soofi S against Nemotron (same architecture) and large open-weight models (Qwen, Ministral, and Gemma). Aggregates are the harness-level English and German suite means. Code EN averages HumanEval and MBPP, Code DE averages HumanEval-DE and MBPP-DE, and LBPP is reported separately. At the aggregate level, Qwen3.5 35B-A3B has the highest English, German, and held-out means … view at source ↗
Figure 9
Figure 9. Figure 9: Base model evaluation overview for Soofi S against Nemotron (same architecture) and large open-weight models (Qwen, Ministral, and Gemma). Aggregates are the harness-level English and German suite means. Code EN averages HumanEval and MBPP, Code DE averages HumanEval-DE and MBPP-DE, and LBPP is reported separately. At the aggregate level, Qwen3.5 35B-A3B has the highest English and German aggregates in thi… view at source ↗
Figure 10
Figure 10. Figure 10: Code generation results (pass@1) against Nemotron 3 Nano and large open-weight models [PITH_FULL_IMAGE:figures/full_fig_p024_10.png] view at source ↗
Figure 10
Figure 10. Figure 10: Code generation results (pass@1) against Nemotron 3 Nano and large open-weight models [PITH_FULL_IMAGE:figures/full_fig_p025_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Mathematics results against Nemotron 3 Nano and large open-weight models on English [PITH_FULL_IMAGE:figures/full_fig_p025_11.png] view at source ↗
Figure 11
Figure 11. Figure 11: Mathematics results against Nemotron 3 Nano and large open-weight models on English [PITH_FULL_IMAGE:figures/full_fig_p026_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Knowledge (left) and reasoning/science (right) benchmarks against Nemotron 3 Nano and [PITH_FULL_IMAGE:figures/full_fig_p026_12.png] view at source ↗
Figure 12
Figure 12. Figure 12: Knowledge (left) and reasoning/science (right) benchmarks against Nemotron 3 Nano [PITH_FULL_IMAGE:figures/full_fig_p027_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: German benchmark results against Nemotron 3 Nano and large open-weight models. [PITH_FULL_IMAGE:figures/full_fig_p027_13.png] view at source ↗
Figure 13
Figure 13. Figure 13: German benchmark results against Nemotron 3 Nano and large open-weight models. [PITH_FULL_IMAGE:figures/full_fig_p028_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Per-benchmark score difference between Soofi S and the architecture-identical [PITH_FULL_IMAGE:figures/full_fig_p028_14.png] view at source ↗
Figure 14
Figure 14. Figure 14: Per-benchmark score difference between Soofi S and the architecture-identical [PITH_FULL_IMAGE:figures/full_fig_p029_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Temporal distribution of the 193.1M Genios articles [ [PITH_FULL_IMAGE:figures/full_fig_p055_15.png] view at source ↗
Figure 15
Figure 15. Figure 15: Benchmark trajectories across annealing checkpoints. GPQA-Diamond rises gradually [PITH_FULL_IMAGE:figures/full_fig_p030_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: RULER accuracy averaged over the selected subtasks, by input context length (4K–1M). [PITH_FULL_IMAGE:figures/full_fig_p055_16.png] view at source ↗
Figure 16
Figure 16. Figure 16: Long-context decode-throughput scaling. Measured aggregate decode TPS/GPU as a function of input context length under the TP=1, single-B200, batch-32 latency-subtraction protocol. The same measurements also expose the prefill/TTFT side of the hybrid architecture: t(Oshort) = t(1) is the measured batch-32 latency to process the prompt and produce the first output token, i.e., the TTFT-like measurement in t… view at source ↗
Figure 17
Figure 17. Figure 17: Temporal distribution of the 193.1M Genios articles [ [PITH_FULL_IMAGE:figures/full_fig_p057_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: RULER accuracy averaged over the selected subtasks, by input context length (4K–1M). [PITH_FULL_IMAGE:figures/full_fig_p057_18.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

103 extracted references · 31 linked inside Pith

  1. [1]

    Gqa: Training generalized multi-query transformer models from multi-head checkpoints

    J.Ainslie, J.Lee-Thorp, M.DeJong, Y.Zemlyanskiy, F.Lebrón, andS.Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, 2023

  2. [2]

    M. Ali, M. Brack, M. Lübbering, E. Wendt, A. G. Khan, R. Rutmann, A. Jude, M. Kraus, A. A. Weber, F. Stollenwerk, D. Kaczér, F. Mai, L. Flek, R. Sifa, N. Flores-Herr, J. Koehler, P. Schramowski, M. Fromm, and K. Kersting. Judging quality across languages: A multilin- gual approach to pretraining data filtering with language models. In C. Christodoulopoulo...

  3. [3]

    M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. Lübbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, H. Abdelwahab, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. Suarez, M. Ostendorff, A. Jude, L. Manjunath, S. Weinbach, C. Penke, O. Filatov, F....

  4. [4]

    Almazrouei, H

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, É. Goffinet, D. Hesslow, J. Launay, Q. Malartic, D. Mazzotta, B. Noune, B. Pannier, and G. Penedo. The falcon series of open language models.arXiv preprint arXiv:2311.16867, 2023. 36

  5. [5]

    Apertus, A

    P. Apertus, A. Hernández-Cano, A. Hägele, A. H. Huang, A. Romanou, A.-J. Solergibert, B. Pasztor, B. Messmer, D. Garbaya, E. F. Ďurech, I. Hakimi, J. G. Giraldo, M. Ismayilzada, N. Foroutan, S. Moalla, T. Chen, V. Sabolčec, Y. Xu, M. Aerni, B. AlKhamissi, I. A. Mar- iñas, M. H. Amani, M. Ansaripour, I. Badanin, H. Benoit, E. Boros, N. Browning, F. Bösch, ...

  6. [6]

    Austin, A

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. Program synthesis with large language models.arXiv preprint arXiv:2108.07732, 2021

  7. [7]

    Biderman, H

    S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. Van Der Wal. Pythia: A suite for analyzing large language models across training and scaling. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors,Proceedings of...

  8. [8]

    BLOOM: A 176b-parameter open-access multilingual language model

    BigScience Workshop. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022

  9. [9]

    Black, S

    S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. Mc- Donell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach. GPT-NeoX-20B: An open-source autoregressive language model. In A. Fan, S. Ilic, T. Wolf, and M. Gallé, editors,Proceedings of BigScience Episode #5 – Workshop o...

  10. [10]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S.Gray, B.Chess, J.Clark, C.Berner, S.McCandlish, A.Radford, I.Sutskever, andD.Amodei. Langu...

  11. [11]

    Burchell, O

    L. Burchell, O. de Gibert, N. Arefyev, M. Aulamo, M. Bañón, P. Chen, M. Fedorova, L. Guillou, B. Haddow, J. Hajič, J. Helcl, E. Henriksson, M. Klimaszewski, V. Komulainen, A. Kutuzov, J. Kytöniemi, V. Laippala, P. Mæhlum, B. Malik, F. Mehryary, V. Mikhailov, N. Moghe, A. Myntti, D. O’Brien, S. Oepen, P. Pal, J. Piha, S. Pyysalo, G. Ramírez-Sánchez, D. Sam...

  12. [12]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

  13. [13]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018

  14. [14]

    Cobbe, V

    K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

  15. [15]

    D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1280–1297, 2024

  16. [16]

    Dao and A

    T. Dao and A. Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality.arXiv preprint arXiv:2405.21060, 2024

  17. [17]

    Deutsche Telekom AG. Germany’s first AI factory for industry officially goes into operation in Munich.https://www.telekom.com/en/media/media-information/archive/germany-s-f irst-ai-factory-for-industry-1101670, Feb. 2026. Industrial AI Cloud, Munich; accessed 2026-07-08

  18. [18]

    S. Diao, Y. Yang, Y. Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y. Suhara, H. Yin, et al. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training.Advances in Neural Information Processing Systems, 38, 2026

  19. [19]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

  20. [20]

    Fujii, Y

    K. Fujii, Y. Tajima, S. Mizuki, M. Kawamura, H. Shimada, T. Shiotani, K. Saito, M. Oi, T. Nakamura, T. Okamoto, S. Ishida, K. Hattori, Y. Ma, H. Takamura, R. Yokota, J. Sakuma, and N. Okazaki. Rewriting pre-training data boosts llm performance in math and code.arXiv preprint arXiv:2505.02881, 2026

  21. [21]

    L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020

  22. [22]

    L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language model evaluation harness, 07 2024

  23. [23]

    GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025

    GBI-Genios Deutsche Wirtschaftsdatenbank GmbH. GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025. Commercially licensed corpus of 38 German newspaper and trade-press archives; obtained under a license that does not permit redistribution

  24. [24]

    Gienapp, C

    L. Gienapp, C. Schröder, S. Schweter, C. Akiki, F. Schlatt, A. Zimmermann, P. Genêt, and M. Potthast. The german commons - 154 billion tokens of openly licensed text for german language models.arXiv preprint arXiv:2510.13996, 2025

  25. [25]

    Reasoning-v1-20m, 2025

    GlaiveAI. Reasoning-v1-20m, 2025. A synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama- 70B

  26. [26]

    Gonzalez-Agirre, M

    A. Gonzalez-Agirre, M. Pàmies, J. Llop, I. Baucells, S. D. Dalt, D. Tamayo, J. J. Saiz, F. Es- puña, J. Prats, J. Aula-Blasco, M. Mina, I. Pikabea, A. Rubio, A. Shvets, A. Sallés, I. La- cunza, J. Palomar, J. Falcão, L. Tormo, L. Vasquez-Reina, M. Marimon, O. Pareras, V. Ruiz- Fernández, and M. Villegas. Salamandra technical report.arXiv preprint arXiv:25...

  27. [27]

    Groeneveld, I

    D. Groeneveld, I. Beltagy, E. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, ...

  28. [28]

    Gurgurov and T

    D. Gurgurov and T. Röhr. ReasonXL: A multilingual cross-domain reasoning corpus.https: //huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual, 2026

  29. [29]

    Gurgurov, T

    D. Gurgurov, T. Röhr, and S. Ostermann. Nemotron-multilingual-reasoning: A multilingual science reasoning dataset.https://huggingface.co/datasets/DGurgurov/Nemotron-Multi lingual-Reasoning, 2025. Derived from nvidia/Llama-Nemotron-Post-Training-Dataset with machine-translated reasoning traces

  30. [30]

    Hägele, E

    A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. Von Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024

  31. [31]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding, 2021

  32. [32]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Stein- hardt. Measuring mathematical problem solving with the math dataset. In J. Vanschoren and S. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021

  33. [33]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre. An empirical analysis of compute-optimal large language model training. InAdvanc...

  34. [34]

    Hsieh, S

    C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models?, 2024

  35. [35]

    S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024

  36. [36]

    Idahl, B

    M. Idahl, B. Droste, B. Plüster, and J. P. Harries. propella-1: Multi-property document annotation for llm data curation at scale, 2026

  37. [37]

    Idahl, J

    M. Idahl, J. Tiedemann, S. Pyysalo, D. Salinas, T. Galica, S. Qian, T. N. Mateiu, Z. Li, A. Lokrantz, F. Vitiugin, A. F. T. Martins, J. Kanerva, F. Ginter, M. Lindemann, T. Isbis- ter, B. Moell, J. Lindh, J. Hajič, J. Jitsev, A. Kutuzov, S. Oepen, and G. Ramírez-Sánchez. Multisynt/mt: Trillion-token multi-parallel pre-training data translated across 36 la...

  38. [38]

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bres- sand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mistral 7b.arXiv preprint arXiv:2310.06825, 2023

  39. [39]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. Bou Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. Le Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mixtral of experts.arXiv prepri...

  40. [40]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Rad- ford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  41. [41]

    Krajewski, J

    J. Krajewski, J. Ludziejewski, K. Adamczewski, M. Pióro, M. Krutul, S. Antoniak, K. Ciebiera, K. Król, T. Odrzygóźdź, P. Sankowski, et al. Scaling laws for fine-grained mixture of experts. arXiv preprint arXiv:2402.07871, 2024

  42. [42]

    Kraus, R

    M. Kraus, R. Härle, S. Sztwiertnia, A. G. Khan, M. Ali, M. Fromm, and K. Kersting. Kletter- mix: Climbing toward high-quality german pretraining data.arXiv preprint arXiv:2606.03773, 2026

  43. [43]

    Kwiatkowski, J

    T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:452–466, 2019

  44. [44]

    Kydlíček, G

    H. Kydlíček, G. Penedo, and L. von Werra. Finepdfs.https://huggingface.co/datasets/ HuggingFaceFW/finepdfs, 2025

  45. [45]

    Lewkowycz, A

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra. Solving quantitative reasoning problems with language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Syste...

  46. [46]

    Liesenfeld, M

    A. Liesenfeld, M. Dingemanse, D. Blankvoort, N. Kalra, and A. R. Golkhandan. European open source ai definitions.https://osai-index.eu/osai-definitions, 2026. Accessed: 2026-07-06

  47. [47]

    A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, A. Sablayrolles, A. Héliou, A. You, A. Ehrenberg, A. Lo, A. Eliseev, A. Calvi, A. Sooriyarachchi, B. Bout, B. Rozière, B. D. Monicault, C. Lanfranchi, C. Barreau, C. Courtot, D. Grattarola, D. Dabert, D. de las Casas, E. Chane-Sane, F....

  48. [48]

    Z. Liu, Z. Yang, Y. Chen, C. Lee, M. Shoeybi, B. Catanzaro, and W. Ping. Acereason- nemotron 1.1: Advancing math and code reasoning through sft and rl synergy.arXiv preprint arXiv:2506.13284, 2025

  49. [49]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  50. [50]

    Lübbering, T

    M. Lübbering, T. Ruland, R. Rutmann, F. Stollenwerk, D. Fitzek, M. Fromm, A. Weber, R. Sifa, N. Flores-Herr, J. Köhler, et al. Modalities, a pytorch-native framework for large-scale llm training and research.arXiv preprint arXiv:2602.08387, 2026

  51. [51]

    R. K. Mahabadi, S. Satheesh, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc-math: A 133 billion-token-scale high quality math pretraining dataset.arXiv preprint arXiv:2508.15096, 2025

  52. [52]

    P. H. Martins, J. Alves, P. Fernandes, N. M. Guerreiro, R. Rei, A. Farajian, M. Klimaszewski, D. M. Alves, J. Pombal, N. Boizard, M. Faysse, P. Colombo, F. Yvon, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm-9b: Technical report.arXiv preprint arXiv:2506.04079, 2025

  53. [53]

    P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm: Multilingual language models for europe.arXiv preprint arXiv:2409.16235, 2024

  54. [54]

    Matton, T

    A. Matton, T. Sherborne, D. Aumiller, E. Tommasone, M. Alizadeh, J. He, R. Ma, M. Voisin, E. Gilsenan-McMahon, and M. Gallé. On leakage of code generation evaluation datasets. In 41 Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors,Findings of the Association for Com- putational Linguistics: EMNLP 2024, pages 13215–13223, Miami, Florida, USA, Nov. 2024. A...

  55. [55]

    Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025

    Mistral AI. Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025. Model release, March 2025

  56. [56]

    Morrison, S

    J. Morrison, S. Adhikesaven, A. Bhagia, M. Zaharia, N. A. Smith, and S. Min. Train separately, merge together: Modular post-training with mixture-of-experts, 2026

  57. [57]

    Muennighoff, L

    N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. Morrison, S. Min, W. Shi, E. P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Ha- jishirzi. OLMoE: Open mixture-of-experts language models. InThe Thirteenth International C...

  58. [58]

    Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025

    MultiSynt. Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025. An automatic translations into 2 languages of Glaive AI reasoning dataset

  59. [59]

    Basant, A

    NVIDIA, :, A. Basant, A. Khairnar, A. Paithankar, A. Khattar, A. Renduchintala, A. Malte, A. Bercovich, A. Hazare, A. Rico, A. Ficek, A. Kondratenko, A. Shaposhnikov, A. Bukharin, A. Taghibakhshi, A. Barton, A. S. Mahabaleshwarkar, A. Shen, A. Tao, A. Guan, A. Shors, A. Mandarwal, A. Mehta, A. Venkatesan, A. Sharabiani, A. Aithal, A. Poojary, A. Dattagupt...

  60. [60]

    Blakeman, A

    NVIDIA, :, A. Blakeman, A. Basant, A. Khattar, A. Renduchintala, A. Bercovich, A. Ficek, A. Bjorlin, A. Taghibakhshi, A. S. Deshmukh, A. S. Mahabaleshwarkar, A. Tao, A. Shors, A. Aithal, A. Poojary, A. Dattagupta, B. Buddharaju, B. Chen, B. Ginsburg, B. Wang, B. Norick, B. Butterfield, B. Catanzaro, C. del Mundo, C. Dong, C. Harvey, C. Parisien, D. Su, D....

  61. [61]

    Blakeman, A

    NVIDIA, :, A. Blakeman, A. Grattafiori, A. Basant, A. Gupta, A. Khattar, A. Renduchintala, A. Vavre, A. Shukla, A. Bercovich, A. Ficek, A. Shaposhnikov, A. Kondratenko, A. Bukharin, A. Milesi, A. Taghibakhshi, A. Liu, A. Barton, A. S. Mahabaleshwarkar, A. Klein, A. Zuker, A. Geifman, A. Shen, A. Bhiwandiwalla, A. Tao, A. Guan, A. Mandarwal, A. Mehta, A. A...

  62. [62]

    Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov

    NVIDIA. Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov. 2025. Accessed 2026-07-08

  63. [63]

    Oepen, N

    S. Oepen, N. Arefev, M. Aulamo, M. Bañón, M. Buljan, L. Burchell, L. Charpentier, P. Chen, M. Fedorova, O. de Gibert, B. Haddow, J. Hajič, J. Helcl, A. Kutuzov, V. Laippala, Z. Li, R. Luukkonen, B. Malik, V. Mikhailov, A. Myntti, D. O’Brien, L. Poláková, S. Pyysalo, G. R. Sánchez, J. Siewert, P. Stepachev, J. Tiedemann, T. Vahtola, D. Variš, F. Vitiugin, ...

  64. [64]

    Olmo, :, A

    T. Olmo, :, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A. Asai, ...

  65. [65]

    The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024

    Open Source Initiative. The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024. Accessed: 2026-07-06. 44

  66. [66]

    OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025

    OpenEuroLLM Consortium. OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025. Project website, accessed 2026-06-25

  67. [67]

    G. Penedo. Finewiki, 2025. Source: Wikimedia Enterprise Snapshot API (https://api.enterprise.wikimedia.com/v2/snapshots). Text licensed under CC BY-SA 4.0 with attribution to Wikipedia contributors

  68. [68]

    Penedo, H

    G. Penedo, H. Kydlíček, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, T. Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024

  69. [69]

    SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025

    Pleias and AI Alliance. SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025. Dataset comprising 79,648,272 text samples (over 41 billion words) derived from the synthetic amplification of 58,698 Wikipedia and Wikibooks articles. Licensed under CC-BY-SA 4.0

  70. [70]

    Poznanski, A

    J. Poznanski, A. Rangapur, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, C. Wilhelm, K. Lo, andL.Soldaini. olmocr: Unlockingtrillionsoftokensinpdfswithvisionlanguagemodels.arXiv preprint arXiv:2502.18443, 2025

  71. [71]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026

  72. [72]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019

  73. [73]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of Machine Learning Research, 21(140):1–67, 2020

  74. [74]

    M. M. Ramos, D. M. Alves, H. Gisserot-Boukhlef, J. Alves, P. H. Martins, P. Fernandes, J. Pombal, N. M. Guerreiro, R. Rei, N. Boizard, A. Farajian, M. Klimaszewski, J. G. C. de Souza, B. Haddow, F. Yvon, P. Colombo, A. Birch, and A. F. T. Martins. Eurollm-22b: Technical report.arXiv preprint arXiv:2602.05879, 2026

  75. [75]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023

  76. [76]

    Romanou, N

    A. Romanou, N. Foroutan, A. Sotnikova, S. H. Nelaturu, S. Singh, R. Maheshwary, M. Al- tomare, Z. Chen, M. Haggag, S. A, A. Amayuelas, A. H. Amirudin, D. Boiko, M. Chang, J. Chim, G. Cohen, A. K. Dalmia, A. Diress, S. Duwal, D. Dzenhaliou, D. Florez, F. Farestam, J. M. Imperial, S. Islam, P. Isotalo, M. Jabbarishiviari, B. F. Karlsson, E. Khalilov, C. Kla...

  77. [77]

    Shazeer, A

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538, 2017. 45

  78. [78]

    W. Shi, A. Bhagia, K. Farhat, N. Muennighoff, J. Morrison, E. Walsh, D. Schwenk, S. Long- pre, J. Poznanski, A. Ettinger, et al. Flexolmo: Open language models for flexible data use. Advances in Neural Information Processing Systems, 38:165943–165974, 2026

  79. [79]

    Soldaini, R

    L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnus- son, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, E. Walsh, L. Zettlemoyer, N. Smit...

  80. [80]

    D. Su, K. Kong, Y. Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-CC: Transforming Common Crawl into a refined long-horizon pre- training dataset. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

Showing first 80 references.