REVIEW 4 major objections 5 minor 2 cited by
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Marco-LLM claims that a two-stage multilingual continual-pretraining and post-training recipe lifts Qwen2's average score across 29 languages from 69.1 to 75.5 at 7B scale and from 85.2 to 87.9 at 72B scale while also improving direct…
desk verdict A credible industrial recipe with large claimed multilingual gains, but the evaluation protocol is under-specified and the results are not yet independently verifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a two-stage continual pretraining schedule on a curated 300B-token multilingual corpus, followed by multilingual supervised fine-tuning and direct preference optimization. Stage-I uses 160B tokens at a peak learning rate of 1e-5 with a mixture that keeps 32% English and 17% Chinese to limit catastrophic forgetting; Stage-II uses 140B tokens at 6e-6 and raises the low-resource share from 9% to 15% to push multilingual capability. The corpus work that makes this work is heavy filtering, MinHash deduplication, and the inclusion of parallel data wrapped in diverse translation templates to create cross-lingual alignment. The authors attribute part of the efficiency to Qwen2's 150k-token vocabulary, which keeps low-resource text highly compressible.
What would settle it
Score Marco-7B, Marco-72B, and all baselines on a freshly produced parallel test set in the 19 low-resource languages using one fixed prompt, one decoder, and one tokenizer, and also search the training corpora for near-duplicates of Flores, Belebele, and MMMLU items; the central claim fails if the margins vanish under that protocol or if contamination turns up.
Extended reading notes
Core claim
The paper's central claim is that its two-stage continual pretraining and post-training recipe turns Qwen2 into a model whose low-resource language performance is substantially better than the base and than comparable open models, without sacrificing high-resource performance. The supporting evidence is an average of 75.5 for Marco-7B across 29 languages versus 69.1 for Qwen2.5-7B, and 87.9 for Marco-72B versus 85.2 for Qwen2.5-72B, plus large gains on non-English-pivot Flores translation (19.7 versus 14.6 BLEU at 7B scale). It also reports that Marco-72B beats GPT-4 on many MMMLU languages and beats Google Translate on several Flores directions.
Load-bearing premise
Every reported margin depends on the baselines being evaluated under the identical prompt, decoding, and scoring protocol, and on none of the training corpora containing test-set sentences.
Editorial extensions
If this is right
- A 7B-parameter model built with this recipe can outperform much larger general-purpose models on low-resource language benchmarks.
- English, Chinese, and other high-resource languages do not have to be traded away when extending a model to low-resource languages.
- Direct translation between non-English pairs becomes practical without an English pivot, which is relevant for language pairs that commercial systems serve poorly.
- The parallel-data ablation implies that data filtering is load-bearing at 72B scale, so small-scale pilots may not predict what large models need.
Reading between the lines
- A natural test of generalizability would be to apply the same two-stage recipe to a different base model with a smaller vocabulary; if the gains shrink, the 150k vocabulary is doing much of the work.
- The 5.9% overall data utilization rate suggests the bottleneck is selection rather than crawl size; an explicit tokens-per-language saturation curve would show where extra low-resource data stops paying.
- The authors' observed gap between high- and low-resource languages after SFT suggests that more pretraining tokens for low-resource languages, not longer SFT, is the lever for closing the remaining gap; this is an inference from their Figure 7, not a claim they test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Marco-LLM is a recipe for extending an existing multilingual base model (Qwen2) to 29 languages by (i) continual pretraining on a curated 300B-token mixture with a two-stage curriculum and a lowered learning rate, and (ii) multilingual SFT and DPO. The paper reports large average gains: Marco-7B reaches 75.5 versus 69.1 for Qwen2.5-7B across 29 languages, and Marco-72B reaches 87.9 versus 85.2 for Qwen2.5-72B; the largest margins are in low-resource languages such as Nepali and Kazakh. It also reports improvements on MMMLU, Belebele, TyDiQA, and Flores, including English-pivot and any-to-any translation.
Significance. If the numbers are trustworthy, this is a practically valuable demonstration that a comparatively small amount (300B tokens) of well-curated multilingual continual pretraining plus multilingual post-training can substantially close the low-resource gap of a strong open model. The two-stage continual-pretraining design and the parallel-data filtering ablation (Section 3.5) are useful contributions, and the data collection pipeline is described in rare detail. However, the evidence is entirely benchmark-based, and the manuscript currently does not supply enough protocol detail or contamination checks to verify the headline margins. No code, model weights, or evaluation harness is promised, so independent verification is not possible from the paper alone.
major comments (4)
- [Section 3.3.1, Tables 6 and 11] The evaluation protocol is under-specified. The paper lists datasets, splits, shots, and metrics, but not the exact prompts, answer extraction rules, decoding hyperparameters, or BLEU tokenization/normalization used for any model. Because the central comparison is across different base models with possibly different tokenizers, small protocol differences can move scores by several points. Two values in the tables suggest protocol issues: Llama3-70B obtains 0.1 BLEU on En to Ko in Table 11, and Qwen2.5-7B drops from 80.2 to 71.6 on Dutch in Table 6. Provide the exact harness or a public reference, and report at least one per-language input/output example for each benchmark.
- [Section 4.1.1 and Section 3.1.3] No contamination audit is given for the SFT and DPO corpora. The only exclusion claim concerns high-quality knowledge data (Section 3.1.3); no such guarantee is made for the SFT mixture, which explicitly includes Aya collection, MetaMathQA, MathInstruct, Belle, Orca, WMT dev sets, WikiMatrix, translated preference data, and synthetic data. At least one of these sources has been reported to contain benchmark items, and parallel dev sets can overlap with test sets used for WMT16. Run an n-gram or embedding-based overlap analysis between every training component and every evaluation benchmark (MMMLU, AGIEval, CEval, Belebele, TyDiQA, Flores-200 devtest, XCOPA, XStoryCloze, XWinograd) and report the maximum overlap per benchmark.
- [Section 4.1.5, Table 12] The any-to-any translation section is internally inconsistent. The text refers to Table??, quotes averages of 19.5 and 14.4, while Table 12 reports averages of 19.7 and 14.6. Moreover, Table 12 contains only 7B models, so the abstract's claim of substantial enhancements in any-to-any machine translation tasks is not supported for the 72B model. Fix the reference, correct the numbers, and add the 72B any-to-any results or explicitly limit the claim to the 7B model.
- [Tables 6-11] All reported scores are single-run values without error bars or significance tests. While most claimed gains are large, several comparisons are close, for example Marco-72B versus Qwen2.5-72B on Polish in Table 7 (88.2 versus 88.8) and Marco-7B versus Qwen2.5-7B on Thai in Table 6 (72.9 versus 73.7). For such entries the textual claim of consistent outperformance is not supported without repeated evaluations or a paired test over the per-language subtasks.
minor comments (5)
- [Section 4.1.5 and Appendix A.2] Model names are used inconsistently: Marco, Marco-7B, Marco-Chat-7B, Marco-72B, and Marco-Chat appear for what seem to be the same models. Please fix the terminology once and for all.
- [Throughout] There are several typos and formatting errors: re-warned should be re-warmed in Section 3.2; Averge in Figure 7; Kazakh(he) in Section 4.2.3; Macro model in Section 4; truction (existing preference dataset) in Section 4.1.5; and the ratio of digits„ in Section 3.1.2.
- [Section 3.3.1 and Section 4.1.3] The relationship between X-MMLU (13 languages, Section 3.3.1) and MMMLU (14 languages, Section 4.1.3) is unclear. Clarify which dataset is used in which table, since both appear in the evaluation suite.
- [Section 4.2.3] The multilingual MT-bench comparison reports win/loss/tie rates from GPT-4o-mini but does not state the number of prompts per language, the judge prompt, the decoding temperature, or the tie-breaking rule. Add these details so the pairwise comparison is reproducible.
- [Section 3.5, Figure 5] The ablation on parallel data filtering is described as showing significant improvements, but Figure 5 has no numerical values and no statistical test. Report the underlying numbers and the number of evaluation examples used.
Circularity Check
No circularity: the paper is an empirical training-and-evaluation report whose headline numbers are measured after held-out evaluation, not derived from fitted quantities or self-citations.
full rationale
Marco-LLM's central claims rest on multi-benchmark evaluations after continual pretraining, SFT, and DPO. No equation in the paper defines a predicted benchmark score from a fitted parameter, and no claimed result is equivalent to a training input by construction. The data mixture and learning rate are tuned on Marco-1.5B using evaluation families such as XStoryCloze, Belebele, and Flores, but this is hyperparameter selection, not a statistical reduction: the final 7B/72B scores are measured outcomes, not quantities forced by the tuning procedure. The paper does not invoke a uniqueness theorem, does not smuggle in an ansatz via citation, and contains no load-bearing self-citation; its baselines are external open models evaluated on public benchmarks. Concerns about possible benchmark contamination or inconsistent evaluation protocols are validity risks, not circularity, and the paper's Section 3.1.3 claim that common benchmark training sets are excluded from high-quality knowledge data is a data-handling statement rather than a derivation. Therefore the derivation chain is self-contained as an empirical system report, and no circular step can be quoted or exhibited.
Assumptions & free parameters
free parameters (6)
- Data mixture proportions =
Stage-I: en 32%, zh 17%, other HR 30%, LR 9%, parallel 6%, HQ 5%, synthetic 1%; Stage-II: 28/15/26/15/8/6/2
- Peak learning rates =
1e-5 (Stage-I), 6e-6 (Stage-II)
- Per-language token caps =
10.6B per high-resource language, 1.9B per low-resource language
- Filtering thresholds =
Not specified numerically
- Stage token budgets =
160B (Stage-I), 140B (Stage-II)
- SFT learning rate range =
6e-6 maximum, 6e-7 minimum
assumptions (4)
- domain assumption Qwen2's tokenizer and pretrained representations transfer to low-resource languages with continued training alone.
- domain assumption Translated versions of benchmarks and preference data are valid for measuring multilingual ability.
- ad hoc to paper Small-scale (1.5B) choices of learning rate and data mixture transfer to 7B and 72B.
- domain assumption Machine translation BLEU on Flores is an appropriate proxy for cross-lingual ability.
Cite this review
Pith. "Pith review of Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement." pith.science (2026). https://pith.science/paper/B3RNZNUE
@misc{pith2026241204003,
author = {Pith},
title = {Pith review of: Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement},
year = {2026},
howpublished = {\url{https://pith.science/paper/B3RNZNUE}},
note = {Machine review of arXiv:2412.04003}
}
read the original abstract
Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challenges with multilingual tasks, especially when it comes to low-resource languages. To address this issue, we introduced Marco-LLM: Massive multilingual training for cross-lingual enhancement LLM. We have collected a substantial amount of multilingual data for several low-resource languages and conducted extensive continual pre-training using the Qwen2 models. This effort has resulted in a multilingual LLM named Marco-LLM. Through comprehensive evaluations on various multilingual benchmarks, including MMMLU, AGIEval, Belebele, Flores-200, XCOPA and many others, Marco-LLM has demonstrated substantial improvements over state-of-the-art LLMs. Furthermore, Marco-LLM achieved substantial enhancements in any-to-any machine translation tasks, showing the effectiveness of our multilingual LLM. Marco-LLM is a pioneering multilingual LLM designed to not only perform exceptionally well in multilingual tasks, including low-resource languages, but also maintain strong performance in English and other major languages, closing the performance gap between high- and low-resource language capabilities. By bridging languages, this effort demonstrates our dedication to ensuring LLMs work accurately across various languages.
Forward citations
Cited by 2 Pith papers
-
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
A multilingual LLM training method that groups similar languages, converts high-deviation layers into mixture-of-experts layers, and assigns one expert per language group improves perplexity across 18 to 128 languages.
-
Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation
The WMT 2024 literary translation shared task finds that domain-enhanced systems lead in d-BLEU for Chinese-English, but human evaluators rank NLP2CT-UM and SJTU-LoveFiction at the top.
Reference graph
Works this paper leans on
-
[1]
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier - Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. \' A brego, J. Ahn, J. Austin, P. Barham, J. A. Botha, J. Bradbury, S. Brahma, K. ...
arXiv 2023
-
[2]
M. Artetxe, G. Labaka, E. Agirre, and K. Cho. Unsupervised neural machine translation. In International Conference on Learning Representations, 2018
work page 2018
-
[3]
V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, J. A. Campos, Y. C. Tan, K. Marchisio, M. Bartolo, S. Ruder, A. Locatelli, J. Kreutzer, N. Frosst, A. Gomez, P. Blunsom, M. Fadaee, A. Üstün, and S. Hooker. Aya 23: Open weight releases to further multilingual progress, 2024. URL https://arxiv.org/abs/2405.15032
arXiv 2024
-
[4]
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023
arXiv 2023
-
[5]
L. Bandarkar, D. Liang, B. Muller, M. Artetxe, S. N. Shukla, D. Husa, N. Goyal, A. Krishnan, L. Zettlemoyer, and M. Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
-
[6]
Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023
arXiv 2023
-
[7]
L. Barrault, O. Bojar, F. Bougares, R. Chatterjee, M. R. Costa-juss \`a , C. Federmann, M. Fishel, A. Fraser, Y. Graham, P. Guzman, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, A. Martins, M. Morishita, C. Monz, M. Nagata, T. Nakazawa, and M. Negri, editors. Proceedings of the Fifth Conference on Machine Translation, Online, Nov. 2020. Association for Compu...
work page 2020
-
[8]
O. Bojar, C. Buck, R. Chatterjee, C. Federmann, L. Guillou, B. Haddow, M. Huck, A. J. Yepes, A. N \'e v \'e ol, M. Neves, P. Pecina, M. Popel, P. Koehn, C. Monz, M. Negri, M. Post, L. Specia, K. Verspoor, J. Tiedemann, and M. Turchi, editors. Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, Berlin, Germany, Aug. 20...
Show all 76 references
-
[9]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Che...
1901
-
[10]
Chaudhary, Y
V. Chaudhary, Y. Tang, F. Guzmán, H. Schwenk, and P. Koehn. Low-resource corpus filtering using multilingual sentence embeddings. In Proceedings of the Fourth Conference on Machine Translation (Volume 3: Shared Task Papers, Day 2), pages 263--268, Florence, Italy, August 2019...
2019
-
[11]
Chiang, L
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024
2024
-
[12]
Chowdhery, S
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Au...
2024
-
[13]
J. H. Clark, E. Choi, M. Collins, D. Garrette, T. Kwiatkowski, V. Nikolaev, and J. Palomaki. T y D i QA : A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8: 0 454--470, 20...
2020 doi
-
[14]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. URL https://arxiv.org/abs/1803.05457
2018 arXiv
-
[15]
c4ai-command-r-plus-08-2024, 2024
Cohere For AI . c4ai-command-r-plus-08-2024, 2024. URL https://huggingface.co/CohereForAI/c4ai-command-r-plus-08-2024
2024
-
[16]
Conneau and G
A. Conneau and G. Lample. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems, 32: 0 7059--7069, 2019
2019
-
[17]
Conneau, R
A. Conneau, R. Rinott, G. Lample, A. Williams, S. Bowman, H. Schwenk, and V. Stoyanov. XNLI : Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475--2485, Brussels, Belgium, Oct....
2018 doi
-
[18]
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun. Ultrafeedback: Boosting language models with high-quality feedback, 2023
2023
-
[19]
Dac Lai, C
V. Dac Lai, C. Van Nguyen, N. T. Ngo, T. Nguyen, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. arXiv e-prints, pages arXiv--2307, 2023
2023
-
[20]
DeepSeek - AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Yang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, ...
2024 arXiv
-
[21]
Dubey and et al
A. Dubey and et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783
2024 arXiv
-
[22]
El-Kishky, V
A. El-Kishky, V. Chaudhary, F. Guzm \'a n, and P. Koehn. CCAligned : A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 5960--5969, Online, November 2020. Assoc...
2020 doi
-
[23]
A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, E. Grave, M. Auli, and A. Joulin. Beyond english-centric multilingual machine translation, 2020. URL https://arxiv.org/a...
2020 arXiv
-
[24]
Goyal, C
N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzm\' a n, and A. Fan. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. 2021
2021
-
[25]
Gunasekar, Y
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li. Textbooks are all you need. CoRR, abs/2306.11644, 2023
2023 arXiv
-
[26]
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang. Deepseek-coder: When the large language model meets programming -- the rise of code intelligence, 2024. URL https://arxiv.org/abs/2401.14196
2024 arXiv
-
[27]
Gurnee, N
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. Trans. Mach. Learn. Res., 2023, 2023
2023
-
[28]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In ICLR . OpenReview.net, 2021
2021
-
[29]
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun. Minicpm: Unveiling the potential of small language model...
2024 arXiv
-
[30]
Huang, Y
Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural Information Processing Systems, 2023
2023
-
[31]
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, A. Yang, R. Men, F. Huang, X. Ren, X. Ren, J. Zhou, and J. Lin. Qwen2.5-coder technical report, 2024. URL https://arxiv.org/abs/2409.12186
2024 arXiv
-
[32]
Hurst, A
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[33]
Ibrahim, B
A. Ibrahim, B. Th \' e rien, K. Gupta, M. L. Richter, Q. G. Anthony, E. Belilovsky, T. Lesort, and I. Rish. Simple and scalable strategies to continually pre-train large language models. Trans. Mach. Learn. Res., 2024, 2024
2024
-
[34]
W. Jiao, W. Wang, J.-t. Huang, X. Wang, and Z. Tu. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745, 1 0 (10), 2023
2023 arXiv
-
[35]
Joulin, E
A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jégou, and T. Mikolov. Fasttext.zip: Compressing text classification models. arXiv: Computation and Language,arXiv: Computation and Language, Nov 2016
2016
-
[36]
Z. Ke, Y. Shao, H. Lin, T. Konishi, G. Kim, and B. Liu. Continual pre-training of language models, 2023. URL https://arxiv.org/abs/2302.03241
2023 arXiv
-
[37]
V. D. Lai, N. T. Ngo, A. P. B. Veyseh, H. Man, F. Dernoncourt, T. Bui, and T. H. Nguyen. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning. arXiv preprint arXiv:2304.05613, 2023
2023 arXiv
-
[38]
Y. Li, S. Bubeck, R. Eldan, A. D. Giorno, S. Gunasekar, and Y. T. Lee. Textbooks are all you need II: phi-1.5 technical report. CoRR, abs/2309.05463, 2023
2023 arXiv
-
[39]
X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, R. Pasunuru, S. Shleifer, P. S. Koura, V. Chaudhary, B. O'Horo, J. Wang, L. Zettlemoyer, Z. Kozareva, M. T. Diab, V. Stoyanov, and X. Li. Few-shot learning with multilingual la...
2021 arXiv
-
[40]
Lovenia, R
H. Lovenia, R. Mahendra, S. M. Akbar, L. J. V. Miranda, J. Santoso, E. Aco, A. Fadhilah, J. Mansurov, J. M. Imperial, O. P. Kampman, et al. Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages. arXiv preprint arXiv:2406.10118, 2024
2024 arXiv
-
[41]
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. CoRR, abs/2308.09583, 2023
2023 arXiv
-
[42]
Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang. Wizardcoder: Empowering code large language models with evol-instruct. In ICLR . OpenReview.net, 2024
2024
-
[43]
Nguyen, C
T. Nguyen, C. V. Nguyen, V. D. Lai, H. Man, N. T. Ngo, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen. Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In LREC/COLING , pages 4226--4237. ELRA and ICCL , 2024
2024
-
[44]
GPT-4 Technical Report
OpenAI. GPT-4 Technical Report . arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[45]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Gray, et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, 2022
2022
-
[46]
Penedo, Q
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, H. Alobeidli, A. Cappelli, B. Pannier, E. Almazrouei, and J. Launay. The refinedweb dataset for falcon LLM: outperforming curated corpora with web data only. In NeurIPS, 2023
2023
-
[47]
Pires, E
T. Pires, E. Schlinger, and D. Garrette. How multilingual is multilingual bert? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996--5001, 2019
2019
-
[48]
E. M. Ponti, G. Glava s , O. Majewska, Q. Liu, I. Vuli \'c , and A. Korhonen. XCOPA : A multilingual dataset for causal commonsense reasoning. In B. Webber, T. Cohn, Y. He, and Y. Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process...
2020 doi
-
[49]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=HPuSIXJaa9
2023
-
[50]
Schwenk, V
H. Schwenk, V. Chaudhary, S. Sun, H. Gong, and F. Guzm \'a n. W iki M atrix: Mining 135 M parallel sentences in 1620 language pairs from W ikipedia. In P. Merlo, J. Tiedemann, and R. Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Associati...
2021 doi
-
[51]
S. She, W. Zou, S. Huang, W. Zhu, X. Liu, X. Geng, and J. Chen. MAPO : Advancing multilingual reasoning through multilingual-alignment-as-preference optimization. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for C...
2024 doi
-
[52]
F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations, 2023. URL https://ope...
2023
-
[53]
Singh, N
H. Singh, N. Gupta, S. Bharadwaj, D. Tewari, and P. Talukdar. Indicgenbench: A multilingual benchmark to evaluate generation capabilities of llms on indic languages, 2024 a . URL https://arxiv.org/abs/2404.16816
2024 arXiv
-
[54]
Singh, F
S. Singh, F. Vargus, D. Dsouza, B. F. Karlsson, A. Mahendiran, W.-Y. Ko, H. Shandilya, J. Patel, D. Mataciunas, L. OMahony, M. Zhang, R. Hettiarachchi, J. Wilson, M. Machado, L. S. Moura, D. Krzemiński, H. Fadaei, I. Ergün, I. Okoh, A. Alaagib, O. Mudannayake, Z. Alyafeai, V. ...
2024 arXiv
-
[55]
N. Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, ...
2022 arXiv
-
[56]
Tiedemann
J. Tiedemann. Parallel data, tools and interfaces in OPUS . In N. Calzolari, K. Choukri, T. Declerck, M. U. Do g an, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, and S. Piperidis, editors, Proceedings of the Eighth International Conference on Language Resources and Evaluation...
2012
-
[58]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. Llama: Open and efficient foundation language models, 2023 b . URL https://arxiv.org/abs/2302.13971
2023 arXiv
-
[59]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hoss...
2023 arXiv
-
[60]
\" U st \" u n, V
A. \" U st \" u n, V. Aryabumi, Z. X. Yong, W. Ko, D. D'souza, G. Onilude, N. Bhandari, S. Singh, H. Ooi, A. Kayid, F. Vargus, P. Blunsom, S. Longpre, N. Muennighoff, M. Fadaee, J. Kreutzer, and S. Hooker. Aya model: An instruction finetuned open-access multilingual language m...
2024
-
[61]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. CoRR, abs/1706.03762, 2017. URL http://arxiv.org/abs/1706.03762
2017 arXiv
-
[62]
L. Wang, C. Lyu, T. Ji, Z. Zhang, D. Yu, S. Shi, and Z. Tu. Document-level machine translation with large language models. arXiv preprint arXiv:2304.02210, 2023
2023 arXiv
-
[63]
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2022
2022
-
[64]
T. Wei, L. Zhao, L. Zhang, B. Zhu, L. Wang, H. Yang, B. Li, C. Cheng, W. L \" u , R. Hu, C. Li, L. Yang, X. Luo, X. Wu, L. Liu, W. Cheng, P. Cheng, J. Zhang, X. Zhang, L. Lin, X. Wang, Y. Ma, C. Dong, Y. Sun, Y. Chen, Y. Peng, X. Liang, S. Yan, H. Fang, and Y. Zhou. Skywork: A...
-
[65]
X. Wei, H. Wei, H. Lin, T. Li, P. Zhang, X. Ren, M. Li, Y. Wan, Z. Cao, B. Xie, T. Hu, S. Li, B. Hui, B. Yu, D. Liu, B. Yang, F. Huang, and J. Xie. Polylm: An open source polyglot large language model. CoRR, abs/2307.06018, 2023 b
2023 arXiv
-
[66]
Wenzek, M
G. Wenzek, M. Lachaux, A. Conneau, V. Chaudhary, F. Guzm \' a n, A. Joulin, and E. Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. In LREC , pages 4003--4012. European Language Resources Association, 2020
2020
-
[67]
Whitehouse, M
C. Whitehouse, M. Choudhury, and A. F. Aji. Llm-powered data augmentation for enhanced cross-lingual performance. In EMNLP , pages 671--686. Association for Computational Linguistics, 2023
2023
-
[68]
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, and D. Jiang. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In ICLR . OpenReview.net, 2024
2024
-
[69]
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. m T 5: A massively multilingual pre-trained text-to-text transformer. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty...
2021
-
[71]
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. ...
2024 arXiv
-
[72]
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[73]
Y. Yu, Y. Zhuang, J. Zhang, Y. Meng, A. J. Ratner, R. Krishna, J. Shen, and C. Zhang. Large language model as attributed training data generator: A tale of diversity and bias. In NeurIPS, 2023
2023
-
[74]
Zellers, A
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019
2019
-
[75]
Zhang, P
B. Zhang, P. Williams, I. Titov, and R. Sennrich. Improving massively multilingual neural machine translation and zero-shot translation. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational...
2020 doi
-
[76]
Zhong, R
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023. URL https://arxiv.org/abs/2304.06364
2023 arXiv
-
[77]
Çağatay Yıldız, N. K. Ravichandran, P. Punia, M. Bethge, and B. Ermis. Investigating continual pretraining in large language models: Insights and implications, 2024. URL https://arxiv.org/abs/2402.17400
2024 arXiv
-
[78]
Üstün, V
A. Üstün, V. Aryabumi, Z.-X. Yong, W.-Y. Ko, D. D'souza, G. Onilude, N. Bhandari, S. Singh, H.-L. Ooi, A. Kayid, F. Vargus, P. Blunsom, S. Longpre, N. Muennighoff, M. Fadaee, J. Kreutzer, and S. Hooker. Aya model: An instruction finetuned open-access multilingual language mode...
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.