REVIEW 4 major objections 5 minor 57 references
WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Battles among open-source code LLMs train a 6.7B model to 80.5% on HumanEval without proprietary data.
desk verdict WarriorCoder is a solid code-data flywheel paper with strong empirical results, undermined by a judge-based selection signal that is never validated against execution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the battle arena itself: one attacker expert mines an instruction by completing its own chat-template prefix, a defender expert answers it, and the remaining experts vote as judges. The selection rule combines the local vote share $x^i_{A>B}$ with an Elo rating expectation $X^{Elo}_{A>B}$ through Equation (5), with $\alpha=0.7$, so that global consistency tempers noisy local votes. Instruction quality is controlled by four-level difficulty filtering and KCenterGreedy embedding-based compression, where KCenterGreedy is a coreset selection algorithm that chooses a diverse subset from embeddings. The mechanism's work is to convert unlabeled expert knowledge into paired responses with a winner label, which then becomes supervised fine-tuning data.
What would settle it
Retrain the same 6.7B base on the same mined instructions with responses chosen randomly instead of by judge-plus-Elo selection; if the random-response model scores within a few points of WarriorCoder on HumanEval, then the battle-selection step is not what produces the gain.
Extended reading notes
Core claim
The central claim is that a code LLM can be improved to state-of-the-art same-size performance by learning from pairwise battles among open-source expert code LLMs, rather than from data expanded by proprietary models. The authors construct an arena in which each expert alternately attacks and defends, generating instructions by completing the prefix of its own chat template, and the other experts act as judges. The winning response for each instruction is selected by combining the local judge-vote share with an Elo-rating-based global expectation, then used as supervised fine-tuning data for a 6.7B DeepSeekCoder base. The paper reports 80.5% pass@1 on HumanEval, 75.6% on HumanEval+, 76.2% on MBPP, and 64.8% on MBPP+, all without relying on proprietary LLMs, and interprets this as direct evidence that competitive data generation can absorb the strengths of multiple experts.
Load-bearing premise
Everything rests on the open-source judge LLMs being able to tell which of two code answers is more correct and helpful; if their votes are noisy or systematically biased, the winner labels that form the training data may point the model at the wrong answers.
Editorial extensions
If this is right
- If the central claim holds, a 6.7B code model trained on battle-selected data reaches 80.5% pass@1 on HumanEval and 76.2% on MBPP, surpassing all same-size fine-tuned baselines in Table 1.
- The same model reaches 42.9% and 45.4% pass@1 on CRUXEval input/output and 38.1% overall on DS-1000, indicating gains extend beyond simple generation into code reasoning and library usage.
- Learning from more experts monotonically improves all four main benchmarks (Table 5), so the data flywheel should keep benefiting as the competitor pool grows.
- Because the pipeline needs no seed dataset, no human prompts, and no proprietary LLM annotations, it lowers the cost and widens access to building code instruction data.
- The paper also argues the mined instructions are largely novel, with ROUGE scores below 0.6 against existing datasets, so the approach adds independent training examples.
Reading between the lines
- An implication the paper leaves implicit: the same battle framework could generate fine-tuning data for other domains, but only in domains where a panel of open-source judges can reliably rank answers; judge quality would need to be validated first.
- The Elo-plus-vote blend suggests a general recipe for aggregating pairwise preferences under noisy judges; a natural next test is whether simpler aggregators, such as majority vote alone, lose the gains on harder problems.
- Because the instruction pool is mined from the chat templates of the five chosen experts, the diversity ceiling is set by those models; adding more or more diverse open-source experts should push the benchmarks higher, a trend the paper's Table 5 already hints at.
- A cautious reader would want a contamination check: the ROUGE comparisons were done against CodeAlpaca and CodeUltraFeedback, not against HumanEval or MBPP; testing overlap with the evaluation sets would clarify how much of the gain is genuine generalization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. WarriorCoder proposes a data flywheel for code LLMs in which open-source expert models compete in pairwise battles, LLM judges vote on the responses, and the target model is fine-tuned on the responses that win under a combination of local vote share and global Elo ratings. The pipeline mines instructions by prompting chat models with only the prefix of their chat templates, filters by judged difficulty, compresses via KCenterGreedy, and uses the battle winners as SFT targets. The paper reports pass@1 results on HumanEval, HumanEval+, MBPP, MBPP+, CRUXEval, and DS-1000, claiming state-of-the-art performance for a 6.7B model without proprietary LLM annotation.
Significance. If the central claims hold, the paper makes a useful contribution: it provides a fully open-source data construction pipeline for code instruction tuning, removes dependence on proprietary annotators, and reports large gains over same-scale baselines. The strengths are the concrete pipeline description, the analysis of instruction diversity and difficulty, and the multi-benchmark evaluation. The main limitation is that the selection signal (judge votes and Elo) is never validated against executable correctness, so the causal mechanism behind the gains remains unproven.
major comments (4)
- [Section 3.3, Eq. (5)] The summation \sum_{B\in Com\A} in Eq. (5) is not operationalized by the described arena. In each battle round only one attacker and one defender compete (Section 3.1), and Eq. (2) defines x_i^{A>B} for a single opponent B. If the final score is instead computed against all other competitors, the paper does not specify how responses from every model are obtained for the same instruction or how the Elo ratings used in the sum are updated. Since e_i^A is the selection criterion for the training response, the reader cannot reproduce or interpret Eq. (5) without clarification.
- [Section 3.3, Eqs. (3)-(5)] The Elo term is not an independent source of global consistency: R_A and R_B are updated in Eq. (4) from the same judge votes t_A and t_B that define the local score in Eq. (2). The combined score therefore reweights the same judge signal rather than adding a separate measurement of model strength. The paper should either justify analytically why this reweighting corrects judge noise or validate it empirically, for example by comparing selections made with and without the Elo term against a held-out execution-based correctness signal.
- [Section 2.3 and Appendix C] Judge reliability is the core assumption of the method, but it is never validated against executable ground truth. The paper itself notes that judge models struggle with complex problems and exhibit position, verbosity, and self-enhancement biases, and the only reported mitigation is order shuffling. Without reporting inter-judge agreement, judge agreement with test-case execution, or a human-annotated sample, the claim that WarriorCoder learns from the winner (Section 3.4) is not established; the selected responses could merely be the more verbose or stylistically preferred ones.
- [Table 5] The ablation varies the number of experts but does not isolate the contribution of the battle-based selection mechanism. A reader cannot tell whether the gains come from selecting the judge-preferred response, from the diversity of multiple expert responses, or simply from fine-tuning on additional open-source generated data. Adding comparisons against random selection among experts, local-vote-only selection, Elo-only selection, or a pooled dataset without any selection would be needed to support the causal claim implicit in the title and in Section 3.4.
minor comments (5)
- [Section 3.2] The paper states that 'we deduplicate the data and adopt judges to assess their difficulty' but does not specify which models serve as difficulty judges or how the 1-10 score is elicited; please provide the prompt and the judge model.
- [Section 4.1] The explanation for setting alpha to 0.7 ('because we need the Elo Rating only when judges' opinions are divided') is unclear, since alpha=0.7 makes the Elo term dominate rather than apply only in tied cases; please clarify the design.
- [Table 1] The 'Rely on proprietary LLMs?' column uses the symbols '%' and '!' without a legend in the table or caption, so the reader cannot interpret the last column.
- [Equations (2)-(4)] The paper should define what counts as a win and a draw when judges vote; currently only the vote counts t_A and t_B are given, and the mapping from votes to the actual scores s_i^{A>B} and s_i^{B>A} in Eq. (4) is not fully specified.
- [Section 4.4.1 and Figure 3] The ROUGE overlap analysis covers CodeAlpaca and CodeUltraFeedback but not the evaluation benchmarks used in Section 4; an overlap check against HumanEval, MBPP, CRUXEval, and DS-1000 would strengthen the contamination discussion.
Circularity Check
No significant circularity: benchmark evaluation is external; Elo and local votes are two views of the same judge signal by design, not a hidden reduction.
full rationale
WarriorCoder's claimed output—a 6.7B model scoring 80.5/75.6/76.2/64.8 on HumanEval/HumanEval+/MBPP/MBPP+—is verified against external execution-based benchmarks, not against the judge votes used to construct the training set. Within the data-generation pipeline, the local vote fraction (Eq. 2) and the Elo rating (Eqs. 3–4) are indeed derived from the same LLM-judge outcomes, and Eq. 5 combines them; however, the paper explicitly frames the Elo term as a temporally aggregated reweighting of the same preference signal for 'global consistency,' not as an independent correctness measurement. The combination is therefore a disclosed design choice rather than a hidden equivalence between a prediction and its input. The completion-based instruction mining, difficulty filtering, and KCenterGreedy compression operate on instructions sampled from the expert models before any external evaluation and do not import the benchmark outcomes into the selection. Self-citations (e.g., Luo et al. 2024a Arena Learning, Luo et al. 2024b WizardCoder, Xu et al. 2024a WizardLM) appear in related-work or baseline contexts and do not carry the load-bearing derivation; no uniqueness theorem or ansatz is imported from them. The lack of judge-vs-execution agreement analysis and the absence of a random-selection ablation are genuine validity and correctness limitations, but under the stated criteria they are not circularity, because the final state-of-the-art claim does not reduce by construction to the judge-vote input.
Assumptions & free parameters
free parameters (7)
- alpha =
0.7
- K =
40
- difficulty_threshold =
6
- temperature_settings =
1.0, 1.1, 1.2
- top_p_settings =
0.99, 0.995, 1.0
- battle_rounds =
70000
- initial_elo =
not specified
assumptions (5)
- domain assumption Judge votes from open-source LLMs accurately reflect the relative quality of code responses.
- domain assumption Completion-based mining from chat template prefixes produces useful, diverse instructions drawn from the expert's mastered distribution.
- domain assumption KCenterGreedy on RoBERTa embeddings preserves diversity and representativeness of instructions.
- domain assumption Public benchmark pass@1 and pass@5 scores are reliable indicators of code ability.
- domain assumption Elo ratings provide a valid global measure of model skill in this arena.
Cite this review
Pith. "Pith review of WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models." pith.science (2026). https://pith.science/paper/SM2WWDC5
@misc{pith2026241217395,
author = {Pith},
title = {Pith review of: WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/SM2WWDC5}},
note = {Machine review of arXiv:2412.17395}
}
read the original abstract
Despite recent progress achieved by code large language models (LLMs), their remarkable abilities are largely dependent on fine-tuning on the high-quality data, posing challenges for data collection and annotation. To address this, current methods often design various data flywheels to collect complex code instructions, enabling models to handle more intricate tasks. However, these approaches typically rely on off-the-shelf datasets and data augmentation from a limited set of proprietary LLMs (e.g., Claude, GPT4, and so on), which restricts the diversity of the constructed data and makes it prone to systemic biases. In this paper, we propose WarriorCoder, a novel paradigm learns from expert battles to address these limitations. Specifically, we create an arena where leading expert code LLMs challenge each other, with evaluations conducted by impartial judges. This competitive framework generates novel training data from scratch, leveraging the strengths of all participants. Experimental results show that WarriorCoder achieves state-of-the-art performance compared to previous models of the same size, even without relying on proprietary LLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . CoRR, abs/2108.07732
arXiv 2021
-
[2]
Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield - Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson...
-
[3]
Egor Bogomolov, Aleksandra Eliseeva, Timur Galimzyanov, Evgeniy Glukhov, Anton Shapkin, Maria Tigina, Yaroslav Golubev, Alexander Kovrigin, Arie van Deursen, Maliheh Izadi, and Timofey Bryksin. 2024. https://doi.org/10.48550/ARXIV.2406.11612 Long code arena: a set of benchmarks for long-context code models . CoRR, abs/2406.11612
-
[4]
Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca
2023
-
[5]
Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024 a . https://aclanthology.org/2024.emnlp-main.474 Humans or llms as the judge? A study on judgement bias . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024 , pages 8301--8327. Association for Co...
work page 2024
-
[6]
Jie Chen, Yupeng Zhang, Bingning Wang, Xin Zhao, Ji - Rong Wen, and Weipeng Chen. 2024 b . https://aclanthology.org/2024.findings-emnlp.873 Unveiling the flaws: Exploring imperfections in synthetic data and mitigation strategies for large language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, Novem...
work page 2024
-
[7]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bava...
arXiv 2021
-
[8]
David Cheng - Han Chiang and Hung - yi Lee. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.870 Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 15607--15631. Association fo...
Show all 57 references
-
[9]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\
2023
-
[10]
Jordan, Joseph E
Wei - Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. https://openreview.net/forum?id=3MW8GKNyzI Chatbot arena: An open platform for evaluating l...
2024
- [11]
-
[12]
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.183 Enhancing chat language models by scaling high-quality instructional conversations . In Proceedings of the 2023 Conference ...
2023 doi
- [13]
-
[14]
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. https://openreview.net/forum?id=hQwb-lbM6EL Incoder: A generative model for code infilling and synthesis . In The Eleventh Internation...
2023
-
[15]
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang. 2024. Cruxeval: A benchmark for code reasoning, understanding and execution. arXiv preprint arXiv:2401.03065
2024 arXiv
- [16]
-
[17]
Grundy, and Haoyu Wang
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John C. Grundy, and Haoyu Wang. 2023. https://doi.org/10.48550/ARXIV.2308.10620 Large language models for software engineering: A systematic literature review . CoRR, abs/2308.10620
- [18]
- [19]
-
[20]
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Wen tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2022. Ds-1000: A natural and reliable benchmark for data science code generation. ArXiv, abs/2211.11501
2022 arXiv
-
[21]
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy - Poirier, Jo \ a o Mont...
2023
-
[22]
Gonzalez, and Ion Stoica
Tianle Li, Wei - Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. https://doi.org/10.48550/ARXIV.2406.11939 From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline . CoRR, abs/2406.11939
- [23]
-
[24]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. https://openreview.net/forum?id=1qvx610Cu7 Is your code generated by chat GPT really correct? rigorous evaluation of large language models for code generation . In Thirty-seventh Conference on Neural Informa...
2023
-
[25]
Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. 2024. https://openreview.net/forum?id=IBCBMeAhmC Evaluating language models for efficient code generation . In First Conference on Language Modeling
2024
-
[26]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1907.11692 Roberta: A robustly optimized BERT pretraining approach . CoRR, abs/1907.11692
2019 arXiv
- [27]
-
[28]
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024 b . https://openreview.net/forum?id=UnUwSIgK5W Wizardcoder: Empowering code large language models with evol-instruct . In The Twelfth International Co...
2024
-
[29]
Lyu, Baishakhi Ray, Abhik Roychoudhury, Shin Hwei Tan, and Patanamon Thongtanunam
Michael R. Lyu, Baishakhi Ray, Abhik Roychoudhury, Shin Hwei Tan, and Patanamon Thongtanunam. 2024. https://doi.org/10.48550/ARXIV.2405.02213 Automatic programming: Large language models and beyond . CoRR, abs/2405.02213
-
[30]
Niklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. 2024. https://openreview.net/forum?id=mw1PWNSWZP Octopack: Instruction tuning code large language models . In The Twe...
2024
- [31]
-
[32]
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. https://openreview.net/forum?id=iaYcJKpY2B\_ Codegen: An open large language model for code with multi-turn program synthesis . In The Eleventh International Conf...
2023
-
[33]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[34]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...
2022
- [35]
-
[36]
Ozan Sener and Silvio Savarese. 2018. https://openreview.net/forum?id=H1aIuk-RW Active learning for convolutional neural networks: A core-set approach . In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Confe...
2018
-
[37]
DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng. 2025. https://arxiv.org/abs/2502.03275 Token assorted: Mixing latent and text tokens for improved language model reasoning . Preprint, arXiv:2502.03275
2025 arXiv
-
[38]
Qwen Team. 2024. https://qwenlm.github.io/blog/qwq-32b-preview/ Qwq: Reflect deeply on the boundaries of the unknown
2024
- [39]
-
[40]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...
2023 doi
-
[41]
Joty, and Steven C
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. https://doi.org/10.18653/V1/2021.EMNLP-MAIN.685 Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation . In Proceedings of the 2021 Conference on Empirical Met...
2021 doi
-
[42]
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2024. https://openreview.net/forum?id=XUeoOBid3x Magicoder: Empowering code generation with oss-instruct . In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2...
2024
- [43]
- [44]
- [45]
-
[46]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 a . https://openreview.net/forum?id=CfXh93NDgH Wizardlm: Empowering large pre-trained language models to follow complex instructions . In The Twelfth Internati...
2024
- [47]
-
[48]
Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J. Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/ae9500c4f5607caf2eff033c67daa9d7-Abstract-Datasets\_and\_Benchmarks.html Large language model as attributed ...
2023
-
[49]
Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, and Qiufeng Yin. 2024. https://doi.org/10.18653/V1/2024.ACL-LONG.280 Wavecoder: Widespread and versatile enhancement for code large language models by instruction tuning . In Proceedings of t...
2024 doi
-
[50]
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang, and Jian - Guang Lou. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.411 Large language models meet nl2code: A survey . In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023 doi
- [51]
-
[52]
Xing, Joseph E
Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2024. https://openreview.net/forum?id=BOfDKxfwt0 Lmsys-chat-1m: A large-scale real-world LLM con...
2024
-
[53]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 a . http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Ab...
2023
- [54]
- [55]
-
[56]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[57]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.