{"id":"b5b9755c-e4ea-46ac-bfac-91b37c2d2ac1","arxiv_id":"2412.08291","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"This paper presents a taxonomy-based review of code-focused large language models, grouping tasks, corpora, models, benchmarks, and challenges, and compiles code-generation benchmark scores.","lead":"This paper organizes recent AI coding models into a taxonomy: tasks, training data, model types, benchmarks, and challenges. It is a survey, not a new experiment, so readers should treat it as a map of the field rather than a source of new results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table IV presents pass@1 values labeled as the authors' own experiments with no protocol, harness, temperature, or code; those entries are unverifiable, and the benchmark comparison that anchors the survey's analysis cannot be trusted.","rationale":"The reader's load-bearing concern is correct. A survey's contribution is reliability, and Table IV is the only place where the paper offers first-party empirical content; that content is unverifiable by design. This is not a disagreement with consensus; it is an internal reproducibility failure: the paper itself labels the numbers as coming from the authors' own experiments while omitting the experimental conditions. The concrete test would settle it by rerunning standard EvalPlus evaluations. If the values reconstruct, the concern is resolved; if not, the survey's central comparison and rankings are unsupported. I also note that the paper contains several objective defects, including a truncated sentence and citation mismatches, and that the 'first taxonomy-based survey' claim is in tension with the cited taxonomy survey of She et al. [94]. However, the Table IV verification is the single most decisive check because it addresses the paper's own empirical claims rather than its novelty framing. The reader's REJECT verdict therefore remains appropriate, and no change to the verdict is needed.","tokens_in":16447,"tokens_out":5244,"duration_ms":52880,"concrete_test":"Run the official EvalPlus harness (Liu et al., 2024) on the released checkpoints for each model with an asterisked entry in Table IV, namely phi-1, phi-2, CodeLLaMA 7B, SantaCoder, and StarCoder 7B, using the standard EvalPlus prompt template and temperature 0.2 with the standard number of samples, and recompute HumanEval+ and MBPP+ pass@1. Compare each result to the corresponding Table IV entry. If any asterisked value cannot be reproduced, or differs by more than the usual run-to-run tolerance, the table's claim of comparable original measurements fails and the survey's benchmark analysis is unsupported. The paper should also publish the full evaluation protocol, but the numerical reproduction is the decisive check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a unified taxonomy that organizes Code LLMs and supports a comprehensive comparative analysis. That analysis depends critically on the benchmark comparison in Table IV. The table caption explicitly states: '*' denotes the results that are not reported in the original papers and obtained from our own experiments.' At least eight entries (phi-1 MBPP/MBPP+, phi-2 MBPP/MBPP+, CodeLLaMA 7B MBPP/MBPP+, SantaCoder MBPP/MBPP+, StarCoder 7B HumanEval+) are thus claimed as original measurements. The paper provides no evaluation harness, no prompt template, no sampling temperature, no number of samples, and no code or data release. Pass@1 on EvalPlus-style benchmarks is highly sensitive to these choices; without them, the values cannot be checked and their comparability with the non-asterisked numbers is asserted, not established. The table is load-bearing because the survey uses it to rank models and to support statements about performance trends. Additional reliability problems, including a truncated sentence at Section III.C.2 ('CodeParrot (1TB) [75], which contains about 80') and mismatched references (HumanEval cited to [48], a GPT-NeoX paper, and HumanEval+/MBPP+ cited to [49], a multilingual programmer paper), reinforce but do not replace this concern. For a survey whose central value is accuracy and organization, unsupported original benchmark numbers are a decisive flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a taxonomy-based survey of Code LLMs, organizing the field into tasks, corpora, models, benchmarks, and challenges. It reviews encoder-only, encoder-decoder, and decoder-only architectures; summarizes common pretraining corpora and code-generation benchmarks; and presents a comparative table of pass@1 scores for a wide range of models. The stated contribution is a unified classification framework that can serve as a structured reference for researchers and practitioners entering the Code LLM area, together with a discussion of open problems in benchmarks, multilingual coverage, synthetic data, scaling laws, and resource efficiency.","tokens_in":16734,"tokens_out":4244,"duration_ms":51809,"significance":"If the taxonomy and the associated comparisons were accurate and verifiable, the survey would be a useful entry point for newcomers to the field. The paper's breadth is a genuine strength: it covers encoder-only models, encoder-decoder models, decoder-only base and fine-tuned models, pretraining corpora, benchmarks, and open challenges in one place. The proposed NL-NL/PL-NL/PL-PL/NL-PL task classification is coherent and sensible. However, the survey's central comparative analysis depends on Table IV, and a substantial part of that table consists of unreported experimental numbers from the authors' own runs with no protocol, harness, code, or data release. In a survey, where the value proposition is accuracy and organization, such unverifiable numbers are a load-bearing reliability problem, as are the multiple reference mismatches.","major_comments":[{"comment":"The caption of Table IV states that asterisked entries are \"results that are not reported in the original papers and obtained from our own experiments.\" At least eight entries (phi-1 MBPP and MBPP+, phi-2 MBPP and MBPP+, CodeLLaMA 7B MBPP and MBPP+, StarCoder 7B HumanEval+, SantaCoder MBPP and MBPP+) are presented as original measurements. No evaluation harness, prompt template, sampling temperature, number of samples, or code release is provided. Pass@1 values on EvalPlus-style benchmarks are highly sensitive to these choices, so the numbers cannot be verified and their comparability with the non-asterisked values is asserted rather than established. Because the table is used to rank models and support claims about performance trends, this issue undermines the central comparative analysis. The authors should either remove all own-experiment entries or supply a complete, reproducible evaluation protocol and release the code and data.","section":"Table IV and Section III.C.3"},{"comment":"Several citations do not match the works they claim to support. HumanEval is cited as [48], which is the GPT-NeoX paper, rather than the OpenAI HumanEval paper; HumanEval+ and MBPP+ are cited as [49], a multilingual programmer paper, rather than the EvalPlus work; GPT3.5 Turbo is cited as [48] in Table IV; Command R+ is cited as [74], which is a paper about reasoning breakdowns rather than the Command R+ model; CommitPack is cited as [70], which is the Chinchilla scaling paper; DBRX is cited as [58], which is the Granite Code Models paper; and Codestral is cited as [53], the Mistral 7B paper. These mismatches make it difficult for readers to trace the survey's factual claims and must be corrected throughout.","section":"References and citations throughout Section III"},{"comment":"The passage describing code-pretrained models ends mid-sentence with \"such as CodeParrot (1TB) [75], which contains about 80\". This is not merely a typo; it breaks a substantive claim about a training corpus. The sentence should be completed or removed. The same section also contains a factual inconsistency in Table IV: Claude 3 [Opus] is marked \"Only Weights\" even though Claude 3 is proprietary and not open-weight, and Codestral is marked \"No\" despite having a research-license release. These details matter for a survey whose goal is to organize the model landscape accurately.","section":"Section III.C.3.a.2"}],"minor_comments":[{"comment":"There is a typo \"JA V A\" that should read \"JAVA\".","section":"III.C.3.b"},{"comment":"The benchmark name is written inconsistently as both \"Multiple-E\" and \"Multipl-E\"; the correct name is \"Multipl-E\".","section":"IV.A.b"},{"comment":"The model name is written inconsistently as \"MagiCoder\" in Table IV and \"Magicoder\" in the text; please unify.","section":"III.C.3.b"},{"comment":"The \"Open Source?\" column mixes categories such as \"Only Weights\", \"Weights & Data\", and \"No\" without a legend; please define what each term means and use it consistently.","section":"Table IV"},{"comment":"The conclusion states that the survey covers works \"rigorously evaluated\" on recognized benchmarks, but Table IV includes unreported self-measured values that do not meet this standard; the statement should be reconciled with the table.","section":"V"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful taxonomy and broad coverage, but the unverifiable own-experiment entries in Table IV and the widespread reference mismatches are serious quality problems for a survey. I recommend major revision rather than reject because these issues are fixable in principle: the authors can remove or fully document the asterisked numbers, correct the references, and repair the truncated sentence. However, the revision must be thorough; the current version does not meet the standard of reliability expected of a survey."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou can skip the 'first taxonomy-based survey' claim: the paper itself cites [94], which is a taxonomy and survey of pitfalls in code LLMs, so that claim is simply wrong. What the paper does well is organize the area into tasks, corpora, models, benchmarks, and challenges. That five-way split is a clean, sensible framework for a newcomer, and the model coverage is broad, including recent open and closed models. The open problems section is reasonable, if not deep.\n\nThe soft spot is the benchmark table. Table IV reports pass@1 values, and several entries are marked as 'obtained from our own experiments' with no evaluation harness, no sampling temperature, no number of samples, no prompts, and no released code. Those numbers are not verifiable, and pass@1 on EvalPlus-style benchmarks is known to be sensitive to those choices. The table is load-bearing because the survey uses it to rank models and to support performance claims. That alone is enough to make the survey unreliable as a reference.\n\nOther issues are smaller but real. The citation for HumanEval points to a GPT-NeoX paper, and HumanEval+/MBPP+ point to a multilingual-programmer paper; both are mismatches. Section III.C.2 has a truncated sentence about CodeParrot. The self-citations to mHumanEval, CSEPrompts, and MojoBench are fine in principle, but they could have been flagged as the authors' own benchmarks.\n\nIf the authors removed the asterisked entries or provided full evaluation details, corrected the citations, and dropped the 'first' claim, the survey would be a useful orientation for researchers entering the field. As it stands, I would not send it to peer review, because the central comparative analysis rests on unverifiable numbers. I'd desk-reject with an invitation to resubmit after those fixes.","headline":"A serviceable taxonomy that is let down by unverifiable in-house benchmark numbers and an overclaimed 'first survey' title.","tokens_in":17223,"tokens_out":3292,"would_cite":false,"duration_ms":31892,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes a five-part taxonomy to organize Code LLM research into tasks, corpora, models, benchmarks, and challenges.","keywords":["Code LLMs","taxonomy","code generation","large language models","benchmarks","corpora","encoder-decoder models","open problems"],"falsifier":"Apply the taxonomy to a random sample of 50 recent Code LLM publications: if a notable share (say, more than one in ten) cannot be assigned to a single leaf of the five-branch scheme without ad-hoc exceptions, the survey's claim of a unified classification system collapses.","tokens_in":16253,"feed_emoji":"🤖","tokens_out":6624,"duration_ms":64920,"temperature":0.7,"pith_summary":"The paper argues that Code LLM research has grown too scattered to navigate, and responds with a taxonomy-based survey that classifies the field into five sub-areas: tasks, corpora, models, benchmarks, and challenges. It further sorts coding tasks by input/output type (NL-NL, NL-PL, PL-PL, PL-NL) and sorts models by architecture (encoder-only, encoder-decoder, decoder-only), with decoder-only models split into foundational and fine-tuned categories. The paper compiles performance numbers on code-generation benchmarks to compare decoder and encoder-decoder models, and it closes by listing open problems such as benchmark leakage, multilingual coverage, scaling laws, and synthetic data quality. The intended payoff is that a newcomer can enter the field with a map rather than a pile of papers.","feed_headline":"A five-branch taxonomy maps the Code LLM field","feed_subtitle":"Newcomers get one unified classification for tasks, corpora, models, benchmarks, and the field's open challenges.","key_machinery":"The central object is the five-branch taxonomy (tasks, corpora, models, benchmarks, challenges), with the task branch as the organizing spine. Tasks are partitioned by the modality of input and output into NL-NL, NL-PL, PL-PL, and PL-NL; Code LLMs are then linked to each branch, so that every model, corpus, benchmark, or challenge in the survey is positioned relative to this scheme. The taxonomy does not itself predict model performance; it carries the argument by making coverage and gaps visible, such as showing that most benchmarks are Python-only and most models are trained on English instructions.","core_discovery":"The central claim is that the entire Code LLM research area can be captured by a single taxonomy, and that organizing the literature this way reveals both the structure of progress and the gaps. The taxonomy's backbone is a task classification based on whether inputs and outputs are natural language or programming language; around that backbone, the survey groups corpora, model architectures, training strategies, benchmarks, and challenges. On the model side, the paper distinguishes encoder-only, encoder-decoder, and decoder-only architectures, and further separates decoder-only models into foundational (base, code, base+code) and fine-tuned variants. On the evaluation side, it treats pass@k on HumanEval and MBPP as the standard measure for generative code models and tabulates reported pass@1 values, including some the authors ran themselves. The open-problems section then extends the taxonomy's logic into a research agenda.","pith_inferences":["If the taxonomy is adopted by the community, it could become a living index: each new Code LLM paper would carry its taxonomy coordinates, making the survey easy to update rather than a point-in-time snapshot.","The taxonomy's modality split suggests a natural extension the paper does not develop: multimodal code tasks, such as generating code from screenshots or diagrams, would need a new branch or a redefinition of natural-language input.","The authors' own asterisked benchmark numbers could be tested by re-running them on public harnesses; doing so would show how much of the ranking depends on evaluation conditions.","The survey's emphasis on synthetic data quality leads to a testable hypothesis: for a fixed compute budget, a small curated programming corpus may outperform a large noisy one, which is a direct experimental follow-up."],"forward_implications":["Researchers entering the field can locate any given model, dataset, or benchmark within the taxonomy and see which task category it serves.","The taxonomy's NL-PL emphasis marks code generation as the central application of decoder-only models, and the paper's pass@1 tables provide a snapshot ranking of the main models as of August 2024.","The open-problems section converts the taxonomy into a research agenda, pointing to multilingual benchmarks, blind test sets, quality-focused synthetic data, and a re-examination of scaling laws for code.","Because the survey classifies benchmarks by language coverage and test-case counts, it gives readers criteria for choosing an evaluation benchmark for a new code model.","The distinction between foundational models (base, code, base+code) and fine-tuned models gives a vocabulary for describing how any new Code LLM was built."],"supporting_citations":[{"why":"The transformer architecture that grounds the survey's three-way model split into encoder-only, encoder-decoder, and decoder-only.","marker":"[16]"},{"why":"Defines the encoder-only branch that code-oriented models such as CodeBERT extend.","marker":"[4]"},{"why":"The concrete encoder-only Code LLM used as the reference for understanding-oriented coding tasks.","marker":"[2]"},{"why":"The encoder-decoder instance that anchors that branch of the model taxonomy.","marker":"[26]"},{"why":"A decoder-only code foundation model used for the survey's discussion of code-pretrained and base-plus-code decoder families.","marker":"[27]"},{"why":"The marker the paper uses for the HumanEval benchmark, the primary dataset in its code-generation comparison table.","marker":"[48]"},{"why":"The MBPP benchmark, the second dataset in the pass@1 comparison table.","marker":"[50]"},{"why":"The pass@k metric that makes the benchmark table's scores comparable across models.","marker":"[87]"},{"why":"The phi-1 study that grounds the survey's open problem about synthetic corpora and quality over quantity.","marker":"[62]"},{"why":"The scaling hypothesis the survey identifies as needing re-evaluation for Code LLMs.","marker":"[70]"}],"fun_headline_variants":["Code LLMs mapped by one unified taxonomy","Taxonomy drives new code LLM survey","One taxonomy to classify code LLMs","A taxonomy-based guide to code LLMs","How a taxonomy reveals code LLM gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pass@1 scores in the main comparison table, including the values the authors say they obtained from their own experiments, were all measured under the same benchmark conditions; the paper provides no experimental setup, prompts, or sampling details to verify that.","fun_headline_variants_meta":{"raw":{"variants":["Code LLMs mapped by one unified taxonomy","Taxonomy drives new code LLM survey","One taxonomy to classify code LLMs","A taxonomy-based guide to code LLMs","How a taxonomy reveals code LLM gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000435,"raw_usage":{"total_tokens":2150,"prompt_tokens":816,"completion_tokens":1334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":1269}},"tokens_in":432,"tokens_out":1334,"duration_ms":10536,"temperature":1.0,"reasoning_tokens":1269,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:59:00.693049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the taxonomy to a random sample of 50 recent Code LLM publications: if a notable share (say, more than one in ten) cannot be assigned to a single leaf of the five-branch scheme without ad-hoc exceptions, the survey's claim of a unified classification system collapses.","supporting_citations":[{"cited_title":"Gpt-neox-20b: An open- source autoregressive language model,","cited_arxiv_id":null,"evidence_quote":"The marker the paper uses for the HumanEval benchmark, the primary dataset in its code-generation comparison table."}],"review_version":1}