REVIEW 4 major objections 6 minor 2 cited by
Trillion 7B Technical Report
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A 7B-parameter model can reach competitive Korean and Japanese performance with only 10% of its 2T training tokens in non-English languages, by letting those tokens attend to packed English documents.
desk verdict A transparent and useful training recipe, but the XLDA mechanism is never ablated, so the paper's central causal claim is unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cross-lingual Document Attention (XLDA) is the load-bearing mechanism: a batch-level packing rule that places documents from at least two languages contiguously in one sequence, paired with a selective attention mask that keeps full self-attention across those language blocks rather than masking document boundaries. The paper describes this as a form of in-context pretraining and synthetic code-switching, so that low-resource-language tokens are always trained in the presence of an English context. Supporting machinery includes a controlled language-sampling mixture with a temperature and upsampling factor, a two-stage warmup-stable-decay schedule with quality-filtered annealing data, multi-token prediction, and a byte-level BPE tokenizer with roughly 100K English tokens and 24,552 Korean tokens.
What would settle it
Train the identical 7B recipe twice, once with XLDA and once with standard document-boundary masking, holding data mixture, quality filters, annealing schedule, tokenizer, and compute fixed; if the standard-mask run matches or beats Trillion-7B on the Korean benchmarks (KoBEST, HAERAE, KMMLU, HRM8k, KoIFEval), the central XLDA claim is refuted.
Extended reading notes
Core claim
Trillion-7B's central claim is that cross-lingual knowledge transfer can be engineered into a pretraining run by changing how documents are packed and masked. Standard pretraining masks document boundaries so tokens cannot attend across documents; XLDA instead enforces that each packed sequence contains contiguous spans from at least two languages and leaves the attention between those language blocks unmasked. The paper argues this acts as architectural code-switching, so the model meta-learns non-English tokens inside an English linguistic context and transfers English knowledge to Korean, Japanese, and Chinese. With this mechanism plus stronger quality filtering (top 50% of multilingual data, top 10% during annealing), increased data diversity, and a Korean tokenizer sized at 24,552 tokens, the model reports competitive Korean benchmark scores and the highest English-to-Korean prediction consistency among the compared 7-9B models (77.5% of correct English predictions are also correct in Korean), despite using far fewer multilingual tokens and far less compute than the comparison models.
Load-bearing premise
The load-bearing premise is that letting Korean tokens attend to a preceding packed English document transfers useful knowledge rather than adding noise, and that this attention change—not the concurrent quality filtering, annealing, diversity, or tokenizer changes—is what drives the reported gains; the ablations reported test the other components but never isolate the mask.
Editorial extensions
If this is right
- Korean-capable 7B models become reproducible on a small budget: the paper reports full pretraining plus post-training for 59.4K H100 GPU hours, about $148K, with only 10% of tokens in non-English languages.
- The transfer mechanism is not Korean-specific: the same 10% multilingual budget lifts Japanese and Chinese xwinograd and Global-MMLU scores, so the recipe should apply to any low-resource language paired with English.
- English performance does not have to be sacrificed: the annealed high-quality, composition-shifted run improves Global-MMLU in English even while English data volume is reduced, which the paper attributes to cross-lingual bridging.
- A small proxy model (1.8B parameters, about 100B tokens) can select the final training recipe, because downstream emergence is observable at that scale; this makes the full 2T run a scaled-up confirmation rather than a blind bet.
Reading between the lines
- Editorial inference: If XLDA works by exposing low-resource tokens to English context, the same mask could be dropped into any English-centric pretraining run for another low-resource language, provided that language has enough clean, diverse text and a well-sized tokenizer; the paper's ablations suggest the mask alone is not sufficient, so the transferable unit is the whole mask-plus-curation rec
- Editorial inference: The consistency result implies an English-anchored shared representation; a testable corollary is that XLDA gains shrink on tasks requiring local cultural knowledge or on languages with very different syntax and little shared vocabulary, where English context cannot supply the missing information.
- Editorial inference: The English-only vision-language result suggests XLDA-style pretraining could decouple visual alignment from language-specific data collection; if so, one English visual-instruction run would transfer to Korean and other languages, which is directly testable with the released model.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Trillion-7B is a 7B-parameter Korean-centric multilingual language model trained on 2T tokens, with the stated goal of transferring English knowledge to Korean and Japanese using a proposed Cross-lingual Document Attention (XLDA) mechanism. XLDA packs documents from different languages into the same sequence and removes the standard causal-mask document boundary for cross-lingual attention. The report describes the pretraining recipe (data filtering, two-stage annealing, tokenizer, infrastructure, context extension), post-training with a Tulu-3-style SFT/DPO/RLVR pipeline, and evaluations on 27 benchmarks in English, Korean, Japanese, and Chinese, plus cross-lingual consistency and a VLM extension experiment. The central claims are token efficiency (59.4K H100 GPU hours, about $148K), competitive performance, and that XLDA is the mechanism enabling efficient knowledge transfer.
Significance. If established, the result would be significant: it would show that a small architectural change (attention masking) plus data curation can substitute for large amounts of target-language pretraining data, and the cost accounting is a useful contribution. The paper's strengths include a broad evaluation suite, explicit training-cost reporting, and a set of ablations for data quality, annealing composition, data diversity, and vocabulary size. Its main weakness is that the causal contribution of XLDA itself is never isolated, and the headline budget and performance statements are not fully consistent with the reported numbers.
major comments (4)
- [§2.2 and §6] The central causal claim that XLDA enables cross-lingual transfer is not tested. Section 2.2 describes the XLDA mask, but none of the Section 6 ablations vary the attention mask: Section 6.1 varies quality filtering and annealing composition, Section 6.2 varies data diversity, and Section 6.3 varies vocabulary size. A matched comparison with standard per-document causal masking under otherwise identical data, tokenizer, training schedule, and compute is required; without it, the observed Korean and Japanese results could be driven by the simultaneous changes in data filtering, annealing upsampling, multi-token prediction, or the Korean-specific tokenizer. This omission is load-bearing because the abstract and Section 1 attribute the efficiency gains to XLDA.
- [§3.1 vs Abstract/§1] The multilingual token budget is internally inconsistent. The abstract and Section 1 say roughly 10% of 2T tokens are multilingual, with "less than 180B" Korean tokens; Section 3.1 states an 8.5:1:0.5 English:Korean:Other ratio, which at 2T tokens implies about 200B Korean tokens alone and about 300B total non-English tokens (15%). These statements cannot all be correct, and the 10% figure is a headline efficiency claim that must be reconciled with the actual mixture and token counts.
- [Table 4] The phrase "competitive performance" is not supported by the model's own summary table. Trillion-7B macro-averages 57.15 and ranks third of six models, behind Qwen2.5-7B (66.15) and EXAONE-3.5-8B (66.71). The largest deficits are in Coding (47.94 vs 66.35 and 70.33) and Math (45.02 vs 67.29 and 65.82), which are central to general-purpose LLM claims. If the intended claim is "competitive for the amount of compute or tokens used," the paper needs a compute-normalized comparison and confidence intervals; as written, the broad competitive-performance claim overreaches the reported results.
- [§7.2] The vision-language generalization claim is not a controlled comparison. Trillion-LLaVA is compared to Llava-1.5-Vicuna-7B and Llava-1.6-Mistral-7B after English-only visual instruction tuning, but the underlying base LLMs differ in architecture, pretraining data, and post-training. The higher Korean VLM scores could reflect properties of the base model rather than the multilingual pretraining; an adequate control would fine-tune the same base architecture with and without the Trillion-7B multilingual pretraining.
minor comments (6)
- [§2.1, Eq. (2)] The formula for P(l_i) contains an undefined operator "˝" and does not specify the normalization or the concrete values of α and β_l used in training; please fix the notation and give the actual mixture parameters.
- [§3.1 and §4] The text refers to "Qwen-72B-Instruct" in Section 3.1 and "Qwen-2.5-72B" in Section 4; please clarify whether these are the same model and which checkpoint was used for filtering and for LLM-as-a-judge scoring.
- [Table 10] The evaluated model is labeled "Trillion-7B-preview" in Table 10, while the abstract and title use "Trillion-7B"; please state explicitly whether these refer to the same checkpoint.
- [Table 4 and Appendix C] SOLAR-10.7B appears in the Appendix C tables but is absent from the headline comparison in Table 4; either include it in the main table or explain the exclusion criterion.
- [§3.5 and §6.3] The relationship between the apparent optimal Korean vocabulary size of 1,500–5,000 tokens in Figure 6, the scaling-law value of roughly 13,000 tokens, and the final choice of 24,552 tokens needs a clearer derivation; the current text jumps between these numbers without a defined selection rule.
- [Figure 3] Figure 3 is described as showing scaling curves but provides no data points, model versions, or evaluation settings; please state the source or mark the figure as an illustrative sketch.
Circularity Check
No significant circularity: benchmark results are measured externally; XLDA lacks a direct ablation, but that is an evidentiary gap, not a circular reduction.
full rationale
No load-bearing step in this report reduces to its own inputs by construction. The central claims are benchmark measurements of a trained 7B model against external baselines (MMLU, KMMLU, KoBEST, GMMLU, HAERAE, etc.), not quantities fitted from those same benchmarks and then re-labeled as predictions. XLDA and the data mixture are specified in Section 2 and the reported gains are measured afterward; even though Section 6 never ablates XLDA against standard document masking, a missing control is a causal-attribution gap, not equation-level circularity. The scaling-law choices for learning rate and vocabulary are imported from external work (DeepSeek-AI et al., 2024; Tao et al., 2024; Hoffmann et al., 2022), and the tokenizer design is independently tested in Section 6.3. The only apparent self-citations are Prometheus 2 and BigGen Bench (authored in part by J. Suk and J. Shin), cited as method references for LLM-as-a-judge filtering and evaluation; the actual scorer used is Qwen-2.5-72B, and the core benchmark results are external, so these citations are not load-bearing. The internal inconsistency in the multilingual token budget (abstract says 10%, Section 1 says <220B tokens, while the 8.5:1:0.5 mixture implies roughly 15% multilingual) is a reporting inconsistency rather than a circular step. Score 2 reflects only the minor, non-load-bearing self-citations; no constructional circularity is present.
Assumptions & free parameters
free parameters (6)
- Multilingual data mixture ratio (English:Korean:Other) =
8.5 : 1 : 0.5
- Korean vocabulary size =
24,552 tokens
- Quality filtering thresholds =
top 80% English, top 50% multilingual, top 20%/10% at annealing
- XLDA packing mixing probability rho =
not reported
- Sampling temperature alpha and upsampling factors beta_l in Eq. (2) =
not reported
- MTP loss weight alpha =
0.2 (0.1 at annealing)
assumptions (5)
- domain assumption Cross-document attention from Korean to preceding English text transfers knowledge without harmful contamination.
- domain assumption Qwen-72B quality scores are a valid gold standard for document quality, and the distilled scorer with F1=0.734 preserves this signal.
- domain assumption A 1.8B-parameter model trained on 100B tokens can predict the relative merits of training recipes for the 7B/2T model.
- domain assumption Empirical scaling laws for learning rate and vocabulary size (DeepSeek-AI 2024; Tao et al. 2024) apply to this multilingual setting.
- domain assumption LLM-as-a-judge scores (Qwen-2.5-72B, MT-Bench judges) are reliable for post-training data selection and evaluation.
Cite this review
Pith. "Pith review of Trillion 7B Technical Report." pith.science (2026). https://pith.science/paper/TOEZ7OUR
@misc{pith2026250415431,
author = {Pith},
title = {Pith review of: Trillion 7B Technical Report},
year = {2026},
howpublished = {\url{https://pith.science/paper/TOEZ7OUR}},
note = {Machine review of arXiv:2504.15431}
}
abstract
We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient and effective knowledge transfer from English to target languages like Korean and Japanese. Combined with optimized data mixtures, language-specific filtering, and tailored tokenizer construction, Trillion-7B achieves competitive performance while dedicating only 10\% of its 2T training tokens to multilingual data and requiring just 59.4K H100 GPU hours (\$148K) for full training. Comprehensive evaluations across 27 benchmarks in four languages demonstrate Trillion-7B's robust multilingual performance and exceptional cross-lingual consistency.
Forward citations
Cited by 2 Pith papers
-
Predicting LLM Reasoning Performance with Small Proxy Model
rBridge uses a small proxy model's confidence-weighted likelihood of a frontier model's reasoning traces to predict and rank large-model reasoning performance across scales.
-
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
The paper introduces an automated RAG-based pipeline that generates and filters 7,555 Korean neutral-toxic sentence pairs, plus 539 English pairs, for training detoxification models.
Reference graph
Works this paper leans on
- [1]
-
[2]
C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle. Does your data spark joy? performance gains from domain upsampling at the end of training, 2024. URL https://arxiv.org/abs/2406.03476
arXiv 2024
-
[3]
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amod...
arXiv 2020
-
[4]
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herb...
arXiv 2021
-
[5]
Chiang, Z
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ URL https://lmsys.org/blog/2023-03-30-vicuna/
2023
- [6]
- [8]
-
[9]
DeepSeek-AI, :, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, H. Gao, K. Gao, W. Gao, R. Ge, K. Guan, D. Guo, J. Guo, G. Hao, Z. Hao, Y. He, W. Hu, P. Huang, E. Li, G. Li, J. Li, Y. Li, Y. K. Li, W. Liang, F. Lin, A. X. Liu, B. Liu, W. Liu, X. Liu, X. Liu, Y. Liu, H. Lu, S. Lu, F. Luo, S. Ma, X. Nie, T. Pei, Y. Piao, J...
arXiv 2024
Show all 79 references
-
[10]
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, ...
2025 arXiv
-
[11]
Z. Du, A. Zeng, Y. Dong, and J. Tang. Understanding emergent abilities of language models from the loss perspective, 2025. URL https://arxiv.org/abs/2403.15796
2025 arXiv
-
[12]
T. Gao, A. Wettig, H. Yen, and D. Chen. How to train long-context language models (effectively), 2025. URL https://arxiv.org/abs/2410.02660
2025
-
[13]
Gloeckle, B
F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction, 2024. URL https://arxiv.org/abs/2404.19737
2024 arXiv
-
[14]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...
2024 arXiv
-
[15]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding, 2021 a . URL https://arxiv.org/abs/2009.03300
2021 arXiv
-
[16]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset, 2021 b . URL https://arxiv.org/abs/2103.03874
2021 arXiv
-
[17]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021 c
2021
-
[18]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...
2022 arXiv
-
[19]
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun. Minicpm: Unveiling the potential of small language model...
2024 arXiv
-
[20]
Hägele, E
A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. V. Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations, 2024. URL https://arxiv.org/abs/2405.18392
2024 arXiv
-
[21]
M. Jang, D. Kim, D. S. Kwon, and E. Davis. K o BEST : K orean balanced evaluation of significant tasks. In N. Calzolari, C.-R. Huang, H. Kim, J. Pustejovsky, L. Wanner, K.-S. Choi, P.-M. Ryu, H.-H. Chen, L. Donatelli, H. Ji, S. Kurohashi, P. Paggio, N. Xue, S. Kim, Y. Hahm, Z....
2022
-
[22]
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed. Mistral 7b, 2023. URL https://arxiv.org/abs...
2023 arXiv
-
[23]
J. Ju, D. Kim, S. Park, and Y. Kim. Varco-vision: Expanding frontiers in korean vision-language models, 2024. URL https://arxiv.org/abs/2411.19103
2024 arXiv
-
[24]
Kaplan, S
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models, 2020. URL https://arxiv.org/abs/2001.08361
2020 arXiv
-
[25]
T. Kida, S. Fukamachi, M. Takeda, A. Shinohara, T. Shinohara, and S. Arikawa. Byte pair encoding: a text compression scheme that accelerates pattern matching. 1999. URL https://api.semanticscholar.org/CorpusID:18801509
1999
-
[26]
B. Kim, H. Kim, S.-W. Lee, G. Lee, D. Kwak, D. H. Jeon, S. Park, S. Kim, S. Kim, D. Seo, H. Lee, M. Jeong, S. Lee, M. Kim, S. H. Ko, S. Kim, T. Park, J. Kim, S. Kang, N.-H. Ryu, K. M. Yoo, M. Chang, S. Suh, S. In, J. Park, K. Kim, H. Kim, J. Jeong, Y. G. Yeo, D. Ham, D. Park, ...
2021 arXiv
-
[27]
S. Kim, J. Shin, Y. Cho, J. Jang, S. Longpre, H. Lee, S. Yun, S. Shin, S. Kim, J. Thorne, et al. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations
-
[28]
S. Kim, J. Suk, J. Y. Cho, S. Longpre, C. Kim, D. Yoon, G. Son, Y. Cho, S. Shafayat, J. Baek, et al. The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models. arXiv preprint arXiv:2406.05761, 2024 a
2024 arXiv
-
[29]
S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo. Prometheus 2: An open source language model specialized in evaluating other language models, 2024 b . URL https://arxiv.org/abs/2405.01535
2024 arXiv
-
[30]
H. Ko, K. Yang, M. Ryu, T. Choi, S. Yang, J. Hyun, S. Park, and K. Park. A technical report for polyglot-ko: Open-source large-scale korean language models, 2023. URL https://arxiv.org/abs/2306.02254
2023 arXiv
-
[31]
H. Ko, G. Son, and D. Choi. Understand, solve and translate: Bridging the multilingual mathematical reasoning gap, 2025. URL https://arxiv.org/abs/2501.02448
2025 arXiv
-
[32]
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, S. ES, S. Suri, D. Glushkov, A. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. Mattick. Openassistant conversations -- democratizing large language ...
2023 arXiv
-
[33]
Lambert, J
N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi. Tulu 3: Pushin...
2025 arXiv
-
[34]
A. K. Lampinen, S. C. Y. Chan, A. K. Singh, and M. Shanahan. The broader spectrum of in-context learning, 2024. URL https://arxiv.org/abs/2412.03782
2024 arXiv
-
[35]
Levine, N
Y. Levine, N. Wies, D. Jannai, D. Navon, Y. Hoshen, and A. Shashua. The inductive bias of in-context learning: Rethinking pretraining example design, 2022. URL https://arxiv.org/abs/2110.04541
2022 arXiv
-
[36]
S. Lin, J. Hilton, and O. Evans. Truthfulqa: Measuring how models mimic human falsehoods, 2022. URL https://arxiv.org/abs/2109.07958
2022 arXiv
-
[37]
H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning, 2023. URL https://arxiv.org/abs/2304.08485
2023 arXiv
-
[38]
Longpre, G
S. Longpre, G. Yauney, E. Reif, K. Lee, A. Roberts, B. Zoph, D. Zhou, J. Wei, K. Robinson, D. Mimno, and D. Ippolito. A pretrainer's guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023. URL https://arxiv.org/abs/2305.13169
2023 arXiv
-
[39]
Muennighoff, T
N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel. Crosslingual generalization through multitask fine...
2023 arXiv
-
[40]
T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi, N. Dziri, M. Guerquin, H. Ivison, P. W. Koh, J. Liu, S. Malik, W. Merrill, L. J...
2025 arXiv
-
[41]
Achiam, S
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L....
2024 arXiv
-
[42]
P. A. Ortega, J. X. Wang, M. Rowland, T. Genewein, Z. Kurth-Nelson, R. Pascanu, N. Heess, J. Veness, A. Pritzel, P. Sprechmann, S. M. Jayakumar, T. McGrath, K. Miller, M. Azar, I. Osband, N. Rabinowitz, A. György, S. Chiappa, S. Osindero, Y. W. Teh, H. van Hasselt, N. de Freit...
2019 arXiv
-
[43]
J. Park. Logickor. 2024. doi:doi:10.57967/hf/2440. URL https://github.com/instructkr/LogicKor
2024 doi
-
[44]
Penedo, H
G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URL https://arxiv.org/abs/2406.17557
2024 arXiv
-
[45]
Petty, S
J. Petty, S. van Steenkiste, and T. Linzen. How does code pretraining affect language model task performance?, 2025. URL https://arxiv.org/abs/2409.04556
2025 arXiv
-
[46]
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...
2025 arXiv
-
[47]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct preference optimization: Your language model is secretly a reward model, 2024. URL https://arxiv.org/abs/2305.18290
2024 arXiv
-
[48]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023. URL https://arxiv.org/abs/2311.12022
2023 arXiv
-
[49]
L. A. Research. Komt-bench. https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench, 2024
2024
-
[50]
L. A. Research, :, S. An, K. Bae, E. Choi, S. J. Choi, Y. Choi, S. Hong, Y. Hong, J. Hwang, H. Jeon, G. J. Jo, H. Jo, J. Jung, Y. Jung, E. Kim, H. Kim, J. Kim, S. Kim, S. Kim, S. Kim, Y. Kim, Y. Kim, E. H. Lee, H. Lee, H. Lee, J. Lee, K. Lee, M. Lee, S. Lee, W. Lim, S. Park, S...
2024
-
[51]
L. A. Research, S. An, K. Bae, E. Choi, K. Choi, S. J. Choi, S. Hong, J. Hwang, H. Jeon, G. J. Jo, H. Jo, J. Jung, Y. Jung, H. Kim, J. Kim, S. Kim, S. Kim, S. Kim, Y. Kim, Y. Kim, Y. Kim, E. H. Lee, H. Lee, H. Lee, J. Lee, K. Lee, W. Lim, S. Park, S. Park, Y. Park, S. Yang, H....
2024
-
[52]
J. Seo, J. Kim, S. Byun, and H. Shin. How does a language-specific tokenizer affect llms?, 2025. URL https://arxiv.org/abs/2502.12560
2025 arXiv
-
[53]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024. URL https://arxiv.org/abs/2402.03300
2024 arXiv
-
[54]
N. Shazeer. Glu variants improve transformer, 2020. URL https://arxiv.org/abs/2002.05202
2020 arXiv
-
[55]
Shoeybi, M
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020. URL https://arxiv.org/abs/1909.08053
2020 arXiv
-
[56]
Singh, A
S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, P. Limkonchotiwat, K. Marchisio, W. Q. Leong, Y. Susanto, R. Ng, S. Longpre, W.-Y. Ko, M. Smith, A. Bosselut, A. Oh, A. F. T. Martins, L. Choshen, D. Ippolito, E. Ferrante, M. Fadaee, B. Ermis, and S....
2024 arXiv
-
[57]
G. Son, H. Lee, S. Kim, H. Kim, J. Lee, J. W. Yeom, J. Jung, J. W. Kim, and S. Kim. Hae-rae bench: Evaluation of korean knowledge in language models, 2024 a . URL https://arxiv.org/abs/2309.02706
2024 arXiv
-
[58]
G. Son, H. Lee, S. Kim, S. Kim, N. Muennighoff, T. Choi, C. Park, K. M. Yoo, and S. Biderman. Kmmlu: Measuring massive multitask language understanding in korean, 2024 b . URL https://arxiv.org/abs/2402.11548
2024 arXiv
-
[59]
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding, 2023. URL https://arxiv.org/abs/2104.09864
2023 arXiv
-
[60]
Suzgun, N
M. Suzgun, N. Scales, N. Sch \"a rli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, , and J. Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022
-
[61]
C. Tao, Q. Liu, L. Dou, N. Muennighoff, Z. Wan, P. Luo, M. Lin, and N. Wong. Scaling laws with vocabulary: Larger models deserve larger vocabularies, 2024. URL https://arxiv.org/abs/2407.13623
2024 arXiv
-
[62]
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev,...
2024 arXiv
-
[63]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hoss...
2023 arXiv
-
[64]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need, 2023. URL https://arxiv.org/abs/1706.03762
2023 arXiv
-
[65]
Z. Wang, J. Li, H. Zhou, R. Weng, J. Wang, X. Huang, X. Han, J. Feng, C. Deng, and S. Huang. Investigating and scaling up code-switching for multilingual language model pre-training, 2025. URL https://arxiv.org/abs/2504.01801
2025 arXiv
-
[66]
Wendler, V
C. Wendler, V. Veselovsky, G. Monea, and R. West. Do llamas work in english? on the latent language of multilingual transformers, 2024. URL https://arxiv.org/abs/2402.10588
2024 arXiv
-
[67]
Workshop, :, T
B. Workshop, :, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A...
2023 arXiv
-
[68]
Wortsman, G
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, 2022. URL https...
2022 arXiv
-
[69]
Xiong, J
W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, M. Khabsa, H. Fang, Y. Mehdad, S. Narang, K. Malik, A. Fan, S. Bhosale, S. Edunov, M. Lewis, S. Wang, and H. Ma. Effective long-context scaling of foundation models,...
2023 arXiv
-
[70]
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. mt5: A massively multilingual pre-trained text-to-text transformer, 2021. URL https://arxiv.org/abs/2010.11934
2021 arXiv
-
[71]
J. Ye, X. Tao, and L. Kong. Language versatilists vs. specialists: An empirical revisiting on multilingual transfer ability, 2023. URL https://arxiv.org/abs/2306.06688
2023 arXiv
-
[72]
H. Yoo, C. Park, S. Yun, A. Oh, and H. Lee. Code-switching curriculum learning for multilingual transfer in llms, 2024 a . URL https://arxiv.org/abs/2411.02460
2024 arXiv
-
[73]
K. M. Yoo, J. Han, S. In, H. Jeon, J. Jeong, J. Kang, H. Kim, K.-M. Kim, M. Kim, S. Kim, D. Kwak, H. Kwak, S. J. Kwon, B. Lee, D. Lee, G. Lee, J. Lee, B. Park, S. Shin, J. Yu, S. Baek, S. Byeon, E. Cho, D. Choe, J. Han, Y. Jin, H. Jun, J. Jung, C. Kim, J. Kim, J. Kim, D. Lee, ...
2024 arXiv
-
[74]
Zellers, A
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830
2019 arXiv
-
[75]
Zhang and R
B. Zhang and R. Sennrich. Root mean square layer normalization, 2019. URL https://arxiv.org/abs/1910.07467
2019 arXiv
-
[76]
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li. Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023. URL https:...
2023 arXiv
-
[77]
Y. Zhao, Y. Qu, K. Staniszewski, S. Tworkowski, W. Liu, P. Miłoś, Y. Wu, and P. Minervini. Analysing the impact of sequence composition on language model pre-training. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...
2024 doi
-
[78]
Y. Zhao, W. Zhang, G. Chen, K. Kawaguchi, and L. Bing. How do large language models handle multilingualism?, 2024 b . URL https://arxiv.org/abs/2402.18815
2024 arXiv
-
[79]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685
2023 arXiv
-
[81]
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models, 2023 b . URL https://arxiv.org/abs/2311.07911
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.