Pith. sign in

REVIEW 4 major objections 6 minor 2 cited by

Trillion 7B Technical Report

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A 7B-parameter model can reach competitive Korean and Japanese performance with only 10% of its 2T training tokens in non-English languages, by letting those tokens attend to packed English documents.

desk verdict A transparent and useful training recipe, but the XLDA mechanism is never ablated, so the paper's central causal claim is unsupported. read the letter →

arxiv 2504.15431 v1 pith:TOEZ7OUR submitted 2025-04-21 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords Cross-lingualDocumentAttentionXLDAmultilinguallanguagemodelKoreanLLMtoken-efficientpretrainingpackingmaskingtransfer
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a 7B-parameter language model can reach competitive Korean performance while spending only 10% of its 2 trillion training tokens on non-English data, and that the key ingredient is an attention-mask change. The proposed Cross-lingual Document Attention (XLDA) lets tokens in a Korean, Japanese, or Chinese document attend to a preceding packed English document, instead of applying the standard mask that blocks attention at every document boundary. On top of this, the model uses quality-filtered data, a two-stage annealing schedule, and a tuned tokenizer; together the recipe costs about 59.4K H100 GPU hours, roughly $148K. If the claims hold, multilingual capability can be bought with architecture and data curation rather than with massive multilingual corpora, which matters for languages that do not have enough text on the web.

What carries the argument

Cross-lingual Document Attention (XLDA) is the load-bearing mechanism: a batch-level packing rule that places documents from at least two languages contiguously in one sequence, paired with a selective attention mask that keeps full self-attention across those language blocks rather than masking document boundaries. The paper describes this as a form of in-context pretraining and synthetic code-switching, so that low-resource-language tokens are always trained in the presence of an English context. Supporting machinery includes a controlled language-sampling mixture with a temperature and upsampling factor, a two-stage warmup-stable-decay schedule with quality-filtered annealing data, multi-token prediction, and a byte-level BPE tokenizer with roughly 100K English tokens and 24,552 Korean tokens.

What would settle it

Train the identical 7B recipe twice, once with XLDA and once with standard document-boundary masking, holding data mixture, quality filters, annealing schedule, tokenizer, and compute fixed; if the standard-mask run matches or beats Trillion-7B on the Korean benchmarks (KoBEST, HAERAE, KMMLU, HRM8k, KoIFEval), the central XLDA claim is refuted.

Watch

Extended reading notes

Core claim

Trillion-7B's central claim is that cross-lingual knowledge transfer can be engineered into a pretraining run by changing how documents are packed and masked. Standard pretraining masks document boundaries so tokens cannot attend across documents; XLDA instead enforces that each packed sequence contains contiguous spans from at least two languages and leaves the attention between those language blocks unmasked. The paper argues this acts as architectural code-switching, so the model meta-learns non-English tokens inside an English linguistic context and transfers English knowledge to Korean, Japanese, and Chinese. With this mechanism plus stronger quality filtering (top 50% of multilingual data, top 10% during annealing), increased data diversity, and a Korean tokenizer sized at 24,552 tokens, the model reports competitive Korean benchmark scores and the highest English-to-Korean prediction consistency among the compared 7-9B models (77.5% of correct English predictions are also correct in Korean), despite using far fewer multilingual tokens and far less compute than the comparison models.

Load-bearing premise

The load-bearing premise is that letting Korean tokens attend to a preceding packed English document transfers useful knowledge rather than adding noise, and that this attention change—not the concurrent quality filtering, annealing, diversity, or tokenizer changes—is what drives the reported gains; the ablations reported test the other components but never isolate the mask.

Editorial extensions

If this is right

  • Korean-capable 7B models become reproducible on a small budget: the paper reports full pretraining plus post-training for 59.4K H100 GPU hours, about $148K, with only 10% of tokens in non-English languages.
  • The transfer mechanism is not Korean-specific: the same 10% multilingual budget lifts Japanese and Chinese xwinograd and Global-MMLU scores, so the recipe should apply to any low-resource language paired with English.
  • English performance does not have to be sacrificed: the annealed high-quality, composition-shifted run improves Global-MMLU in English even while English data volume is reduced, which the paper attributes to cross-lingual bridging.
  • A small proxy model (1.8B parameters, about 100B tokens) can select the final training recipe, because downstream emergence is observable at that scale; this makes the full 2T run a scaled-up confirmation rather than a blind bet.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If XLDA works by exposing low-resource tokens to English context, the same mask could be dropped into any English-centric pretraining run for another low-resource language, provided that language has enough clean, diverse text and a well-sized tokenizer; the paper's ablations suggest the mask alone is not sufficient, so the transferable unit is the whole mask-plus-curation rec
  • Editorial inference: The consistency result implies an English-anchored shared representation; a testable corollary is that XLDA gains shrink on tasks requiring local cultural knowledge or on languages with very different syntax and little shared vocabulary, where English context cannot supply the missing information.
  • Editorial inference: The English-only vision-language result suggests XLDA-style pretraining could decouple visual alignment from language-specific data collection; if so, one English visual-instruction run would transfer to Korean and other languages, which is directly testable with the released model.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. Trillion-7B is a 7B-parameter Korean-centric multilingual language model trained on 2T tokens, with the stated goal of transferring English knowledge to Korean and Japanese using a proposed Cross-lingual Document Attention (XLDA) mechanism. XLDA packs documents from different languages into the same sequence and removes the standard causal-mask document boundary for cross-lingual attention. The report describes the pretraining recipe (data filtering, two-stage annealing, tokenizer, infrastructure, context extension), post-training with a Tulu-3-style SFT/DPO/RLVR pipeline, and evaluations on 27 benchmarks in English, Korean, Japanese, and Chinese, plus cross-lingual consistency and a VLM extension experiment. The central claims are token efficiency (59.4K H100 GPU hours, about $148K), competitive performance, and that XLDA is the mechanism enabling efficient knowledge transfer.

Significance. If established, the result would be significant: it would show that a small architectural change (attention masking) plus data curation can substitute for large amounts of target-language pretraining data, and the cost accounting is a useful contribution. The paper's strengths include a broad evaluation suite, explicit training-cost reporting, and a set of ablations for data quality, annealing composition, data diversity, and vocabulary size. Its main weakness is that the causal contribution of XLDA itself is never isolated, and the headline budget and performance statements are not fully consistent with the reported numbers.

major comments (4)
  1. [§2.2 and §6] The central causal claim that XLDA enables cross-lingual transfer is not tested. Section 2.2 describes the XLDA mask, but none of the Section 6 ablations vary the attention mask: Section 6.1 varies quality filtering and annealing composition, Section 6.2 varies data diversity, and Section 6.3 varies vocabulary size. A matched comparison with standard per-document causal masking under otherwise identical data, tokenizer, training schedule, and compute is required; without it, the observed Korean and Japanese results could be driven by the simultaneous changes in data filtering, annealing upsampling, multi-token prediction, or the Korean-specific tokenizer. This omission is load-bearing because the abstract and Section 1 attribute the efficiency gains to XLDA.
  2. [§3.1 vs Abstract/§1] The multilingual token budget is internally inconsistent. The abstract and Section 1 say roughly 10% of 2T tokens are multilingual, with "less than 180B" Korean tokens; Section 3.1 states an 8.5:1:0.5 English:Korean:Other ratio, which at 2T tokens implies about 200B Korean tokens alone and about 300B total non-English tokens (15%). These statements cannot all be correct, and the 10% figure is a headline efficiency claim that must be reconciled with the actual mixture and token counts.
  3. [Table 4] The phrase "competitive performance" is not supported by the model's own summary table. Trillion-7B macro-averages 57.15 and ranks third of six models, behind Qwen2.5-7B (66.15) and EXAONE-3.5-8B (66.71). The largest deficits are in Coding (47.94 vs 66.35 and 70.33) and Math (45.02 vs 67.29 and 65.82), which are central to general-purpose LLM claims. If the intended claim is "competitive for the amount of compute or tokens used," the paper needs a compute-normalized comparison and confidence intervals; as written, the broad competitive-performance claim overreaches the reported results.
  4. [§7.2] The vision-language generalization claim is not a controlled comparison. Trillion-LLaVA is compared to Llava-1.5-Vicuna-7B and Llava-1.6-Mistral-7B after English-only visual instruction tuning, but the underlying base LLMs differ in architecture, pretraining data, and post-training. The higher Korean VLM scores could reflect properties of the base model rather than the multilingual pretraining; an adequate control would fine-tune the same base architecture with and without the Trillion-7B multilingual pretraining.
minor comments (6)
  1. [§2.1, Eq. (2)] The formula for P(l_i) contains an undefined operator "˝" and does not specify the normalization or the concrete values of α and β_l used in training; please fix the notation and give the actual mixture parameters.
  2. [§3.1 and §4] The text refers to "Qwen-72B-Instruct" in Section 3.1 and "Qwen-2.5-72B" in Section 4; please clarify whether these are the same model and which checkpoint was used for filtering and for LLM-as-a-judge scoring.
  3. [Table 10] The evaluated model is labeled "Trillion-7B-preview" in Table 10, while the abstract and title use "Trillion-7B"; please state explicitly whether these refer to the same checkpoint.
  4. [Table 4 and Appendix C] SOLAR-10.7B appears in the Appendix C tables but is absent from the headline comparison in Table 4; either include it in the main table or explain the exclusion criterion.
  5. [§3.5 and §6.3] The relationship between the apparent optimal Korean vocabulary size of 1,500–5,000 tokens in Figure 6, the scaling-law value of roughly 13,000 tokens, and the final choice of 24,552 tokens needs a clearer derivation; the current text jumps between these numbers without a defined selection rule.
  6. [Figure 3] Figure 3 is described as showing scaling curves but provides no data points, model versions, or evaluation settings; please state the source or mark the figure as an illustrative sketch.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: benchmark results are measured externally; XLDA lacks a direct ablation, but that is an evidentiary gap, not a circular reduction.

full rationale

No load-bearing step in this report reduces to its own inputs by construction. The central claims are benchmark measurements of a trained 7B model against external baselines (MMLU, KMMLU, KoBEST, GMMLU, HAERAE, etc.), not quantities fitted from those same benchmarks and then re-labeled as predictions. XLDA and the data mixture are specified in Section 2 and the reported gains are measured afterward; even though Section 6 never ablates XLDA against standard document masking, a missing control is a causal-attribution gap, not equation-level circularity. The scaling-law choices for learning rate and vocabulary are imported from external work (DeepSeek-AI et al., 2024; Tao et al., 2024; Hoffmann et al., 2022), and the tokenizer design is independently tested in Section 6.3. The only apparent self-citations are Prometheus 2 and BigGen Bench (authored in part by J. Suk and J. Shin), cited as method references for LLM-as-a-judge filtering and evaluation; the actual scorer used is Qwen-2.5-72B, and the core benchmark results are external, so these citations are not load-bearing. The internal inconsistency in the multilingual token budget (abstract says 10%, Section 1 says <220B tokens, while the 8.5:1:0.5 mixture implies roughly 15% multilingual) is a reporting inconsistency rather than a circular step. Score 2 reflects only the minor, non-load-bearing self-citations; no constructional circularity is present.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests on several domain assumptions about cross-lingual transfer, data quality scoring, proxy-model transferability, and judge reliability. No new physical or mathematical entities are introduced. The free parameters are mostly hand-picked recipe choices; several that are most central to XLDA (packing probability, sampling temperature, upsampling factors) are not even reported, which makes the mechanism under-specified and hard to reproduce.

free parameters (6)
  • Multilingual data mixture ratio (English:Korean:Other) = 8.5 : 1 : 0.5
    Chosen by hand in Section 3.1; it determines the claimed 10% multilingual budget and is not derived from any experiment reported.
  • Korean vocabulary size = 24,552 tokens
    Selected in Section 3.5 as a compromise between the scaling-law optimum (about 13,000 tokens) and inference speed; affects Korean tokenization and all downstream numbers.
  • Quality filtering thresholds = top 80% English, top 50% multilingual, top 20%/10% at annealing
    Set by hand in Sections 3.1 and 3.2; ablations show filtering helps, but the exact thresholds are not optimized or justified.
  • XLDA packing mixing probability rho = not reported
    Defined in Section 2.1 as the chance of cross-lingual document adjacency; no value or sensitivity analysis is given, yet it is central to the XLDA mechanism.
  • Sampling temperature alpha and upsampling factors beta_l in Eq. (2) = not reported
    Parameters of the batch-level language sampling distribution (Eq. 2) are never instantiated or ablated, so the actual multilingual sampling recipe is under-specified.
  • MTP loss weight alpha = 0.2 (0.1 at annealing)
    Multi-token prediction weight chosen in Section 3.4; no ablation is shown for this value.
assumptions (5)
  • domain assumption Cross-document attention from Korean to preceding English text transfers knowledge without harmful contamination.
    This is the premise of XLDA (Section 2.3) and is asserted, not demonstrated; no experiment separates it from other recipe changes.
  • domain assumption Qwen-72B quality scores are a valid gold standard for document quality, and the distilled scorer with F1=0.734 preserves this signal.
    Section 3.1 uses these scores to filter all pretraining data; correlation with GPT-4 is claimed but not quantified, and F1 of 0.734 on a binarized task is modest.
  • domain assumption A 1.8B-parameter model trained on 100B tokens can predict the relative merits of training recipes for the 7B/2T model.
    Section 3.3 justifies proxy experiments and scaling-law extrapolations; this transfer is assumed, not validated at 7B scale for each recipe.
  • domain assumption Empirical scaling laws for learning rate and vocabulary size (DeepSeek-AI 2024; Tao et al. 2024) apply to this multilingual setting.
    Section 3.3 uses these laws to set learning rate and vocabulary; the laws were derived mostly on English or monolingual corpora.
  • domain assumption LLM-as-a-judge scores (Qwen-2.5-72B, MT-Bench judges) are reliable for post-training data selection and evaluation.
    Sections 4 and 5 rely on judge scores to filter SFT/DPO data and to score instruction-following benchmarks; no human validation is reported except inherited benchmark design.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Trillion 7B Technical Report." pith.science (2026). https://pith.science/paper/TOEZ7OUR

@misc{pith2026250415431,
  author       = {Pith},
  title        = {Pith review of: Trillion 7B Technical Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TOEZ7OUR}},
  note         = {Machine review of arXiv:2504.15431}
}
abstract

We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient and effective knowledge transfer from English to target languages like Korean and Japanese. Combined with optimized data mixtures, language-specific filtering, and tailored tokenizer construction, Trillion-7B achieves competitive performance while dedicating only 10\% of its 2T training tokens to multilingual data and requiring just 59.4K H100 GPU hours (\$148K) for full training. Comprehensive evaluations across 27 benchmarks in four languages demonstrate Trillion-7B's robust multilingual performance and exceptional cross-lingual consistency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting LLM Reasoning Performance with Small Proxy Model

    cs.LG 2025-09 conditional novelty 6.0 of 10

    rBridge uses a small proxy model's confidence-weighted likelihood of a frontier model's reasoning traces to predict and rank large-model reasoning performance across scales.

  2. K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The paper introduces an automated RAG-based pipeline that generates and filters 7,555 Korean neutral-toxic sentence pairs, plus 539 English pairs, for training detoxification models.

Reference graph

Works this paper leans on

79 extracted references · 6 canonical work pages · cited by 2 Pith papers

  1. [1]

    Austin, A

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton. Program synthesis with large language models, 2021. URL https://arxiv.org/abs/2108.07732

  2. [2]

    Blakeney, M

    C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle. Does your data spark joy? performance gains from domain upsampling at the end of training, 2024. URL https://arxiv.org/abs/2406.03476

  3. [3]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amod...

  4. [4]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herb...

  5. [5]

    Chiang, Z

    W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ URL https://lmsys.org/blog/2023-03-30-vicuna/

  6. [6]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018

  7. [8]

    Cobbe, V

    K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems, 2021 b . URL https://arxiv.org/abs/2110.14168

  8. [9]

    DeepSeek-AI, :, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, H. Gao, K. Gao, W. Gao, R. Ge, K. Guan, D. Guo, J. Guo, G. Hao, Z. Hao, Y. He, W. Hu, P. Huang, E. Li, G. Li, J. Li, Y. Li, Y. K. Li, W. Liang, F. Lin, A. X. Liu, B. Liu, W. Liu, X. Liu, X. Liu, Y. Liu, H. Lu, S. Lu, F. Luo, S. Ma, X. Nie, T. Pei, Y. Piao, J...

Show all 79 references
  1. [10]

    DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, ...

  2. [11]

    Z. Du, A. Zeng, Y. Dong, and J. Tang. Understanding emergent abilities of language models from the loss perspective, 2025. URL https://arxiv.org/abs/2403.15796

  3. [12]

    T. Gao, A. Wettig, H. Yen, and D. Chen. How to train long-context language models (effectively), 2025. URL https://arxiv.org/abs/2410.02660

  4. [13]

    Gloeckle, B

    F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction, 2024. URL https://arxiv.org/abs/2404.19737

  5. [14]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...

  6. [15]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding, 2021 a . URL https://arxiv.org/abs/2009.03300

  7. [16]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset, 2021 b . URL https://arxiv.org/abs/2103.03874

  8. [17]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021 c

  9. [18]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...

  10. [19]

    S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun. Minicpm: Unveiling the potential of small language model...

  11. [20]

    Hägele, E

    A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. V. Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations, 2024. URL https://arxiv.org/abs/2405.18392

  12. [21]

    M. Jang, D. Kim, D. S. Kwon, and E. Davis. K o BEST : K orean balanced evaluation of significant tasks. In N. Calzolari, C.-R. Huang, H. Kim, J. Pustejovsky, L. Wanner, K.-S. Choi, P.-M. Ryu, H.-H. Chen, L. Donatelli, H. Ji, S. Kurohashi, P. Paggio, N. Xue, S. Kim, Y. Hahm, Z....

  13. [22]

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed. Mistral 7b, 2023. URL https://arxiv.org/abs...

  14. [23]

    J. Ju, D. Kim, S. Park, and Y. Kim. Varco-vision: Expanding frontiers in korean vision-language models, 2024. URL https://arxiv.org/abs/2411.19103

  15. [24]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models, 2020. URL https://arxiv.org/abs/2001.08361

  16. [25]

    T. Kida, S. Fukamachi, M. Takeda, A. Shinohara, T. Shinohara, and S. Arikawa. Byte pair encoding: a text compression scheme that accelerates pattern matching. 1999. URL https://api.semanticscholar.org/CorpusID:18801509

  17. [26]

    B. Kim, H. Kim, S.-W. Lee, G. Lee, D. Kwak, D. H. Jeon, S. Park, S. Kim, S. Kim, D. Seo, H. Lee, M. Jeong, S. Lee, M. Kim, S. H. Ko, S. Kim, T. Park, J. Kim, S. Kang, N.-H. Ryu, K. M. Yoo, M. Chang, S. Suh, S. In, J. Park, K. Kim, H. Kim, J. Jeong, Y. G. Yeo, D. Ham, D. Park, ...

  18. [27]

    S. Kim, J. Shin, Y. Cho, J. Jang, S. Longpre, H. Lee, S. Yun, S. Shin, S. Kim, J. Thorne, et al. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations

  19. [28]

    S. Kim, J. Suk, J. Y. Cho, S. Longpre, C. Kim, D. Yoon, G. Son, Y. Cho, S. Shafayat, J. Baek, et al. The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models. arXiv preprint arXiv:2406.05761, 2024 a

  20. [29]

    S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo. Prometheus 2: An open source language model specialized in evaluating other language models, 2024 b . URL https://arxiv.org/abs/2405.01535

  21. [30]

    H. Ko, K. Yang, M. Ryu, T. Choi, S. Yang, J. Hyun, S. Park, and K. Park. A technical report for polyglot-ko: Open-source large-scale korean language models, 2023. URL https://arxiv.org/abs/2306.02254

  22. [31]

    H. Ko, G. Son, and D. Choi. Understand, solve and translate: Bridging the multilingual mathematical reasoning gap, 2025. URL https://arxiv.org/abs/2501.02448

  23. [32]

    A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, S. ES, S. Suri, D. Glushkov, A. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. Mattick. Openassistant conversations -- democratizing large language ...

  24. [33]

    Lambert, J

    N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi. Tulu 3: Pushin...

  25. [34]

    A. K. Lampinen, S. C. Y. Chan, A. K. Singh, and M. Shanahan. The broader spectrum of in-context learning, 2024. URL https://arxiv.org/abs/2412.03782

  26. [35]

    Levine, N

    Y. Levine, N. Wies, D. Jannai, D. Navon, Y. Hoshen, and A. Shashua. The inductive bias of in-context learning: Rethinking pretraining example design, 2022. URL https://arxiv.org/abs/2110.04541

  27. [36]

    S. Lin, J. Hilton, and O. Evans. Truthfulqa: Measuring how models mimic human falsehoods, 2022. URL https://arxiv.org/abs/2109.07958

  28. [37]

    H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning, 2023. URL https://arxiv.org/abs/2304.08485

  29. [38]

    Longpre, G

    S. Longpre, G. Yauney, E. Reif, K. Lee, A. Roberts, B. Zoph, D. Zhou, J. Wei, K. Robinson, D. Mimno, and D. Ippolito. A pretrainer's guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023. URL https://arxiv.org/abs/2305.13169

  30. [39]

    Muennighoff, T

    N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel. Crosslingual generalization through multitask fine...

  31. [40]

    T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi, N. Dziri, M. Guerquin, H. Ivison, P. W. Koh, J. Liu, S. Malik, W. Merrill, L. J...

  32. [41]

    Achiam, S

    OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L....

  33. [42]

    P. A. Ortega, J. X. Wang, M. Rowland, T. Genewein, Z. Kurth-Nelson, R. Pascanu, N. Heess, J. Veness, A. Pritzel, P. Sprechmann, S. M. Jayakumar, T. McGrath, K. Miller, M. Azar, I. Osband, N. Rabinowitz, A. György, S. Chiappa, S. Osindero, Y. W. Teh, H. van Hasselt, N. de Freit...

  34. [43]

    J. Park. Logickor. 2024. doi:doi:10.57967/hf/2440. URL https://github.com/instructkr/LogicKor

  35. [44]

    Penedo, H

    G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URL https://arxiv.org/abs/2406.17557

  36. [45]

    Petty, S

    J. Petty, S. van Steenkiste, and T. Linzen. How does code pretraining affect language model task performance?, 2025. URL https://arxiv.org/abs/2409.04556

  37. [46]

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...

  38. [47]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct preference optimization: Your language model is secretly a reward model, 2024. URL https://arxiv.org/abs/2305.18290

  39. [48]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023. URL https://arxiv.org/abs/2311.12022

  40. [49]

    L. A. Research. Komt-bench. https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench, 2024

  41. [50]

    L. A. Research, :, S. An, K. Bae, E. Choi, S. J. Choi, Y. Choi, S. Hong, Y. Hong, J. Hwang, H. Jeon, G. J. Jo, H. Jo, J. Jung, Y. Jung, E. Kim, H. Kim, J. Kim, S. Kim, S. Kim, S. Kim, Y. Kim, Y. Kim, E. H. Lee, H. Lee, H. Lee, J. Lee, K. Lee, M. Lee, S. Lee, W. Lim, S. Park, S...

  42. [51]

    L. A. Research, S. An, K. Bae, E. Choi, K. Choi, S. J. Choi, S. Hong, J. Hwang, H. Jeon, G. J. Jo, H. Jo, J. Jung, Y. Jung, H. Kim, J. Kim, S. Kim, S. Kim, S. Kim, Y. Kim, Y. Kim, Y. Kim, E. H. Lee, H. Lee, H. Lee, J. Lee, K. Lee, W. Lim, S. Park, S. Park, Y. Park, S. Yang, H....

  43. [52]

    J. Seo, J. Kim, S. Byun, and H. Shin. How does a language-specific tokenizer affect llms?, 2025. URL https://arxiv.org/abs/2502.12560

  44. [53]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024. URL https://arxiv.org/abs/2402.03300

  45. [54]

    N. Shazeer. Glu variants improve transformer, 2020. URL https://arxiv.org/abs/2002.05202

  46. [55]

    Shoeybi, M

    M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020. URL https://arxiv.org/abs/1909.08053

  47. [56]

    Singh, A

    S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, P. Limkonchotiwat, K. Marchisio, W. Q. Leong, Y. Susanto, R. Ng, S. Longpre, W.-Y. Ko, M. Smith, A. Bosselut, A. Oh, A. F. T. Martins, L. Choshen, D. Ippolito, E. Ferrante, M. Fadaee, B. Ermis, and S....

  48. [57]

    G. Son, H. Lee, S. Kim, H. Kim, J. Lee, J. W. Yeom, J. Jung, J. W. Kim, and S. Kim. Hae-rae bench: Evaluation of korean knowledge in language models, 2024 a . URL https://arxiv.org/abs/2309.02706

  49. [58]

    G. Son, H. Lee, S. Kim, S. Kim, N. Muennighoff, T. Choi, C. Park, K. M. Yoo, and S. Biderman. Kmmlu: Measuring massive multitask language understanding in korean, 2024 b . URL https://arxiv.org/abs/2402.11548

  50. [59]

    J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding, 2023. URL https://arxiv.org/abs/2104.09864

  51. [60]

    Suzgun, N

    M. Suzgun, N. Scales, N. Sch \"a rli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, , and J. Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022

  52. [61]

    C. Tao, Q. Liu, L. Dou, N. Muennighoff, Z. Wan, P. Luo, M. Lin, and N. Wong. Scaling laws with vocabulary: Larger models deserve larger vocabularies, 2024. URL https://arxiv.org/abs/2407.13623

  53. [62]

    G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev,...

  54. [63]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hoss...

  55. [64]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need, 2023. URL https://arxiv.org/abs/1706.03762

  56. [65]

    Z. Wang, J. Li, H. Zhou, R. Weng, J. Wang, X. Huang, X. Han, J. Feng, C. Deng, and S. Huang. Investigating and scaling up code-switching for multilingual language model pre-training, 2025. URL https://arxiv.org/abs/2504.01801

  57. [66]

    Wendler, V

    C. Wendler, V. Veselovsky, G. Monea, and R. West. Do llamas work in english? on the latent language of multilingual transformers, 2024. URL https://arxiv.org/abs/2402.10588

  58. [67]

    Workshop, :, T

    B. Workshop, :, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A...

  59. [68]

    Wortsman, G

    M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, 2022. URL https...

  60. [69]

    Xiong, J

    W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, M. Khabsa, H. Fang, Y. Mehdad, S. Narang, K. Malik, A. Fan, S. Bhosale, S. Edunov, M. Lewis, S. Wang, and H. Ma. Effective long-context scaling of foundation models,...

  61. [70]

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. mt5: A massively multilingual pre-trained text-to-text transformer, 2021. URL https://arxiv.org/abs/2010.11934

  62. [71]

    J. Ye, X. Tao, and L. Kong. Language versatilists vs. specialists: An empirical revisiting on multilingual transfer ability, 2023. URL https://arxiv.org/abs/2306.06688

  63. [72]

    H. Yoo, C. Park, S. Yun, A. Oh, and H. Lee. Code-switching curriculum learning for multilingual transfer in llms, 2024 a . URL https://arxiv.org/abs/2411.02460

  64. [73]

    K. M. Yoo, J. Han, S. In, H. Jeon, J. Jeong, J. Kang, H. Kim, K.-M. Kim, M. Kim, S. Kim, D. Kwak, H. Kwak, S. J. Kwon, B. Lee, D. Lee, G. Lee, J. Lee, B. Park, S. Shin, J. Yu, S. Baek, S. Byeon, E. Cho, D. Choe, J. Han, Y. Jin, H. Jun, J. Jung, C. Kim, J. Kim, J. Kim, D. Lee, ...

  65. [74]

    Zellers, A

    R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830

  66. [75]

    Zhang and R

    B. Zhang and R. Sennrich. Root mean square layer normalization, 2019. URL https://arxiv.org/abs/1910.07467

  67. [76]

    Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li. Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023. URL https:...

  68. [77]

    Y. Zhao, Y. Qu, K. Staniszewski, S. Tworkowski, W. Liu, P. Miłoś, Y. Wu, and P. Minervini. Analysing the impact of sequence composition on language model pre-training. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...

  69. [78]

    Y. Zhao, W. Zhang, G. Chen, K. Kawaguchi, and L. Bing. How do large language models handle multilingualism?, 2024 b . URL https://arxiv.org/abs/2402.18815

  70. [79]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685

  71. [81]

    J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models, 2023 b . URL https://arxiv.org/abs/2311.07911

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.