Pith. sign in

REVIEW 1 cited by

Measuring The Impact Of Programming Language Distribution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.01973 v3 pith:L5UTBG66 submitted 2023-02-03 cs.LG cs.CLcs.PL

classification cs.LGcs.CLcs.PL
keywords languageslanguagepassprogrammingbabelcodebenchmarklow-resourcepython
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this issue, we present the BabelCode framework for execution-based evaluation of any benchmark in any language. BabelCode enables new investigations into the qualitative performance of models' memory, runtime, and individual test case results. Additionally, we present a new code translation dataset called Translating Python Programming Puzzles (TP3) from the Python Programming Puzzles (Schuster et al. 2021) benchmark that involves translating expert-level python functions to any language. With both BabelCode and the TP3 benchmark, we investigate if balancing the distributions of 14 languages in a training dataset improves a large language model's performance on low-resource languages. Training a model on a balanced corpus results in, on average, 12.34% higher $pass@k$ across all tasks and languages compared to the baseline. We find that this strategy achieves 66.48% better $pass@k$ on low-resource languages at the cost of only a 12.94% decrease to high-resource languages. In our three translation tasks, this strategy yields, on average, 30.77% better low-resource $pass@k$ while having 19.58% worse high-resource $pass@k$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ransomware 3.0: Self-Composing and LLM-Orchestrated

    cs.CR 2025-08 conditional novelty 8.0 of 10

    A prototype LLM-orchestrated ransomware successfully executes reconnaissance, payload selection, encryption/exfiltration/destruction, and personalized extortion across three environments, with open-source models.

Pith tools