{"id":"b793b56c-c40a-4d7c-8ae1-ec6def4ab958","arxiv_id":"2505.07425","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A systematic literature review that organizes 57 neural code translation papers into seven research themes and identifies current trends and open problems.","lead":"This paper reviews 57 studies published between 2020 and 2025 on using neural networks and large language models to translate code between programming languages. It maps the field into seven dimensions and highlights open challenges such as low-resource languages, evaluation data leakage, and repository-level translation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Search string omits common alternative terminology (e.g., 'transpilation', 'source-to-source translation'), and snowballing cannot recover studies that neither cite nor are cited by the initial set; the 'comprehensive SLR' claim is therefore not yet established.","rationale":"The paper is a useful SLR; the methodological skeleton is standard and the seven-RQ framework is reasonable. The single load-bearing premise is corpus completeness, because every quantitative trend and research-gap conclusion is a statement about the field, not merely about the 57 selected papers. The authors' own Section 6 threat analysis flags incomplete search; the mitigation (broader keywords and snowballing) is plausible but incomplete, since snowballing cannot discover disconnected literature using alternate vocabulary. The expanded-search test is cheap and decisive. The 'no prior SLR' novelty claim is also asserted without systematic comparison to existing surveys; I regard the search-coverage issue as prior to that claim, because if the corpus is incomplete the map itself is biased. No issue with author conduct; the concern is about the evidential support for 'comprehensive'. The Reader's weakest_assumption identified the same search-string coverage issue; I agree and recommend no change to the CONDITIONAL verdict.","tokens_in":30763,"tokens_out":3187,"duration_ms":29177,"concrete_test":"Re-run the Section 3.2 search on the same six databases with an expanded string: ('code translation' OR 'code-to-code translation' OR 'program translation' OR 'code migration' OR 'transpilation' OR 'transpiler' OR 'source-to-source translation' OR 'cross-language migration' OR 'program porting' OR 'code conversion') AND ('deep learning' OR 'large language model'). Apply the same inclusion/exclusion criteria and the QA >= 8 threshold from Table 1, with backward and forward snowballing on the expanded initial set. Count how many additional primary studies qualify and recompute the RQ1-RQ7 percentages. If the additional set is non-empty and shifts any headline percentage by more than 5 percentage points, or adds studies whose absence is material to the open-issues discussion, the comprehensiveness claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that this is the first comprehensive SLR depends on the search in Section 3.2 retrieving the relevant literature. The string ('code translation' OR 'code-to-code translation' OR 'program translation' OR 'code migration') AND ('deep learning' OR 'large language model') omits widely used terms such as 'transpilation', 'transpiler', 'source-to-source translation', 'cross-language migration', 'program porting', and 'code conversion'. The authors state in Section 3.2 that snowballing mitigates missing studies that adopt alternative terminology, but snowballing operates on the reference lists and citations of the initially retrieved studies; a relevant paper that uses none of the included terms and is not linked to the initial set remains invisible. Section 6 acknowledges the 'Incomplete literature search' threat but asserts rather than demonstrates that the chosen keywords plus snowballing are sufficient. Because all RQ statistics (e.g., 71.9% bidirectional, 73.7% function-level, 77.2% plain-text modeling) and the open-challenge list are computed over the 57 selected studies, a systematically missed segment could change trend claims and the gap analysis. The 8/10 quality threshold is also arbitrary, but the terminology gap is the more load-bearing issue for the comprehensiveness claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature review (SLR) of neural code translation. It claims to be the first comprehensive SLR on this topic, selects 57 primary studies published between 2020 and 2025 from six databases, and analyzes them along seven research questions covering task characteristics, data preprocessing, code modeling, model construction, post-processing, evaluation subjects, and evaluation metrics. The paper reports quantitative trends (e.g., 71.9% bidirectional translation, 73.7% function-level granularity, 77.2% plain-text modeling, 42.1% BLEU usage), identifies open challenges (low-resource language translation, repository-level translation, data leakage, security, code repair and refinement), and proposes future directions.","tokens_in":31029,"tokens_out":7239,"duration_ms":64193,"significance":"If the 57-study set is representative, the review provides a useful structured map of an active area: it compiles and classifies recent work, quantifies methodological choices, and catalogs datasets and metrics. The manuscript follows a named SLR guideline (Kitchenham), defines research questions, uses two independent extractors, applies explicit inclusion/exclusion and quality criteria, and includes a threats-to-validity section. The main quantitative findings are novel in their granularity and are falsifiable. However, the scientific value is conditional on the completeness and auditability of the study selection; the paper's own Section 6 acknowledges the 'Incomplete literature search' threat but does not provide the evidence needed to retire it. The contribution is therefore potentially strong but currently under-supported.","major_comments":[{"comment":"The search string is too narrow to support the 'comprehensive' claim. It contains the terms 'code translation', 'code-to-code translation', 'program translation', and 'code migration' combined with 'deep learning' or 'large language model', but omits common alternative terminology such as 'transpilation', 'transpiler', 'source-to-source translation', 'cross-language migration', 'program porting', and 'code conversion'. The paper argues in Section 3.2 that snowballing can recover studies using alternative terminology, but backward and forward snowballing can only find papers that are connected by references or citations to the initially retrieved set; a relevant paper that uses none of the included terms and is not linked to the initial set is systematically missed. The manuscript also does not report how many studies were found by the automated search versus by snowballing, or the per-database hit counts, so the completeness of the 57-study set cannot be audited. Since the trend statistics in Section 4 (e.g., 71.9% bidirectional, 73.7% function-level) and the open-challenge list in Section 5 are computed over this set, a systematically missed segment could change the main conclusions. I request a rerun with an expanded keyword set, a PRISMA-style flow diagram with search and snowballing counts, and a sensitivity analysis of the key RQ statistics.","section":"Section 3.2 (Search Strategy)"},{"comment":"The quality assessment is not auditable. The ten criteria QA1-QA10 in Table 1 are introduced as adopted from prior SLRs, but the paper provides no external validation of their construct validity and no calibration against known high- or low-quality studies. The scoring of binary items as 1/-1 and ternary items as 1/0/-1 makes the 8/10 threshold extremely strict (a single 'No' and one 'Partially' yields 7), and QA10 for arXiv papers ('does it satisfy the quality criteria?') is circular because it refers to the same unvalidated instrument. In addition, the paper does not report per-study QA scores, the number of studies that passed each stage, inter-rater agreement on QA, or how disagreements were resolved. Without this information, the statement that the final set is 'representative, rigorous, and of high reference value' cannot be verified, and the threshold directly shapes every result in Section 4.","section":"Section 3.3 (Study Selection and Quality Assessment)"},{"comment":"The RQ1 statistics are not reproducible from the presented table. Table 2 counts the same study in multiple language-pair cells (for example, a study on Java↔Python and Python↔C++ appears in multiple rows), and the text states that Tang et al. [74] and Lano et al. [43] are excluded from the table, yet the percentages in the RQ1 summary divide by 57. The paper should report the per-study coding (one row per included study with all seven RQ dimensions) in an appendix or online artifact and state clearly whether percentages are computed over studies or over study-category occurrences. This will also allow readers to verify the subsequent claims in Section 5.","section":"Section 4.1 and Table 2 (RQ1)"},{"comment":"The claim of being the first comprehensive SLR is not contextualized against the closest prior work. The paper includes [57] ('On ML-based program translation: perils and promises') as a primary study and cites code-generation surveys [32, 89] for methodology, but it does not discuss how this SLR differs from, or improves on, those earlier reviews in coverage, research questions, or conclusions. Because the novelty claim is a central contribution, a dedicated comparison with prior surveys and a statement of the incremental contribution are needed.","section":"Sections 1 and 7 (Novelty Claim)"}],"minor_comments":[{"comment":"The exact search string is given only in prose; providing the per-database queries with field restrictions and year filters (since the paper claims coverage of 2020-2025) would help reproducibility.","section":"Section 3.2"},{"comment":"There is a typo in the exclusion criteria list: 'Non-English papers.:' has a double colon.","section":"Section 3.3"},{"comment":"The text mentions Tang et al. [74] and Lano et al. [43] as bidirectional studies excluded from Table 2; please clarify why Lano et al. [43], which is described as model-driven engineering, satisfies the inclusion criterion 'Focus on neural code translation'.","section":"Section 4.1 and Table 2"},{"comment":"Several table entries have missing spaces, for example 'C++↔CUDA[76]' and 'Solidity→Move[39]'.","section":"Table 2"},{"comment":"The word cloud of venues is difficult to read; a frequency table of venues would be more informative and precise.","section":"Figure 3b"}],"recommendation":"major_revision","confidential_remarks":"The topic is timely and the review has clear potential, but the comprehensiveness claim is stronger than the evidence supports. The key requests are an expanded and auditable search, a transparent quality-assessment and selection protocol, and a per-study coding artifact. I would not reject the paper on current evidence, but the 'first comprehensive SLR' claim should not be accepted until these issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a well-structured systematic review, and it probably is the first dedicated SLR on neural code translation. The seven-RQ framework (task characteristics, preprocessing, modeling, model construction, post-processing, evaluation subjects, metrics) gives a genuinely useful map of 57 studies from 2020 to 2025. The methodology is largely sound: explicit inclusion/exclusion criteria, two independent extractors with consensus, quality assessment, and clear reporting of trend statistics. Newcomers will find this valuable, and the open-challenge list is sensible.\n\nThe main soft spot is search coverage. The string in Section 3.2—('code translation' OR 'code-to-code translation' OR 'program translation' OR 'code migration') AND ('deep learning' OR 'large language model')—omits common alternatives like 'transpilation', 'source-to-source translation', and 'program porting'. Snowballing helps, but only for papers linked to the initial set; an isolated paper using different vocabulary stays invisible. The paper acknowledges the 'Incomplete literature search' threat in Section 6 but asserts rather than demonstrates sufficiency. So the claim to be 'the first comprehensive SLR' is plausible but not fully proven. This is fixable with a wider string, a supplementary table of all included studies, and per-database search logs.\n\nThe 8/10 quality threshold is arbitrary, as the authors admit by citing precedent; that is common in SLRs and not a fatal flaw. The lack of reported inter-rater reliability (Cohen's kappa or similar) is a minor omission. The self-citation in [88] is a primary study under review, not a red flag.\n\nOverall, the central contribution stands: a useful, reproducible-in-principle map of the field. It deserves peer review. I would want a reviewer to push for the widened search and the appendix before publication, but this is conditional acceptance material, not a reject.","headline":"A competent, genuinely useful first map of neural code translation, with a search-coverage gap that is real but fixable before it can be called comprehensive.","tokens_in":31534,"tokens_out":2058,"would_cite":true,"duration_ms":18840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to be the first comprehensive systematic literature review of neural code translation, synthesizing 57 studies from 2020 to 2025 across seven dimensions to map techniques and unresolved challenges.","keywords":["neural code translation","systematic literature review","large language models","deep learning","code migration","code modeling","evaluation metrics","code translation benchmarks"],"falsifier":"Re-run the review's search string plus candidate synonyms such as 'transpilation', 'cross-language migration', and 'program conversion' in the same six academic databases, applying the same inclusion criteria and 8-out-of-10 quality threshold; if this yields a non-trivial set of qualifying studies outside the 57, the review's comprehensiveness claim fails. Independently, one could re-score the 57 studies against the ten quality criteria; a study scoring below 8 would weaken the reported representativeness.","tokens_in":30554,"feed_emoji":"🔁","tokens_out":4727,"duration_ms":41085,"temperature":0.7,"pith_summary":"This paper tries to establish that neural code translation had no comprehensive systematic literature review and that this review now provides the missing map, covering 57 primary studies published from 2020 to 2025. It synthesizes the field along seven axes: task characteristics, data preprocessing, code modeling, model construction, post-processing, evaluation subjects, and evaluation metrics. If the review is right, researchers and practitioners now have a baseline of what methods dominate, which benchmarks are standard, and which problems remain open.","feed_headline":"Neural code translation gets its first systematic map","feed_subtitle":"Review of 57 studies shows fine-tuning and Java-Python/C++ domination, plus open challenges.","key_machinery":"The machinery is the systematic review protocol itself: an automated search over six academic databases with a query combining code-translation terms with deep-learning and large-language-model terms, backward and forward snowballing, inclusion and exclusion criteria requiring empirical evaluation, and a ten-item quality-assessment checklist with an 8-out-of-10 inclusion threshold. The protocol yields 57 primary studies, and a fixed seven-question data-extraction scheme organises the synthesis. This machinery is what turns scattered papers into comparable statistics and research-gap conclusions.","core_discovery":"The central claim is that prior work on automatically translating source code from one programming language to another using neural models had not been summarised in a comprehensive, systematic way, and that this paper fills that gap by collecting 57 studies and analysing them across seven dimensions. The review reports that the field centres on function-level translation among statically typed languages such as Java, C++, and Python; that fine-tuning pre-trained models is the most common construction strategy; and that BLEU and CodeBLEU are the dominant evaluation metrics despite known limitations. It also identifies open challenges: scarce parallel corpora, low-resource languages, data leakage, repository-level context, and security of translation models.","pith_inferences":["My inference: because the review shows static metrics BLEU and CodeBLEU are the most common while execution-based metrics appear less frequently, reported quality improvements may partly measure textual resemblance rather than runnable correctness; a neighbouring test would correlate BLEU improvements with pass@k changes on the same benchmarks.","My inference: the sharp rise in 2023 and 2024 publications and the heavy reliance on arXiv submissions suggest the field's empirical base is consolidating around a few benchmarks; if so, performance saturation on TransCoder and CodeXGLUE could push the next wave toward repository-level evaluation.","My inference: the review's emphasis on low-resource languages and repository-level context points to a testable target: building a large synthetic parallel corpus for a pair such as Julia and Python and measuring whether transfer learning plus synthesized tests closes the gap to high-resource pairs."],"forward_implications":["Researchers entering neural code translation can use the seven-perspective taxonomy as a structured baseline of existing methods and benchmarks.","Fine-tuning pre-trained models is the current default construction strategy, with prompt engineering as a close, cost-effective alternative, while multi-agent and retrieval-augmented methods remain early-stage.","TransCoder and CodeXGLUE are the de facto evaluation standards, so new work should expect to compare against them.","Class-level, file-level, and repository-level translation, plus low-resource languages, are the review's identified open areas; progress there would directly address the field's unserved migration needs."],"supporting_citations":[{"why":"Supplies the systematic-literature-review methodology the paper follows.","marker":"[41]"},{"why":"Provides the template for the ten-item quality-assessment criteria and scoring scheme.","marker":"[27]"},{"why":"Cited as precedent for applying backward and forward snowballing in the search process.","marker":"[53, 81, 93]"},{"why":"Describes TransCoder, the most-used evaluation subject and a representative unsupervised baseline.","marker":"[68]"},{"why":"Introduces CodeXGLUE, the second most-used evaluation benchmark for code translation.","marker":"[55]"},{"why":"Introduces AVATAR, a frequently used Java-Python parallel corpus and evaluation subject.","marker":"[6]"}],"fun_headline_variants":["First systematic review of neural code translation spans 57 studies","Neural code translation: 57 studies mapped across seven dimensions","Systematic review reveals gaps in neural code translation research","How neural code translation stacks up: 57-study review","Neural code translation's blind spots exposed in new review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the search string, snowballing, and 8-out-of-10 quality threshold together capture the relevant literature on neural code translation; if a substantial body of work using different terminology such as 'transpilation' or 'cross-language migration' is missed, the trend statistics and gap conclusions do not describe the whole field.","fun_headline_variants_meta":{"raw":{"variants":["First systematic review of neural code translation spans 57 studies","Neural code translation: 57 studies mapped across seven dimensions","Systematic review reveals gaps in neural code translation research","How neural code translation stacks up: 57-study review","Neural code translation's blind spots exposed in new review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000683,"raw_usage":{"total_tokens":3040,"prompt_tokens":827,"completion_tokens":2213,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":2132}},"tokens_in":443,"tokens_out":2213,"duration_ms":13307,"temperature":1.0,"reasoning_tokens":2132,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:16:13.277621+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the review's search string plus candidate synonyms such as 'transpilation', 'cross-language migration', and 'program conversion' in the same six academic databases, applying the same inclusion criteria and 8-out-of-10 quality threshold; if this yields a non-trivial set of qualifying studies outside the 57, the review's comprehensiveness claim fails. Independently, one could re-score the 57 studies against the ten quality criteria; a study scoring below 8 would weaken the reported representativeness.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes TransCoder, the most-used evaluation subject and a representative unsupervised baseline."}],"review_version":1}