ResearchCodeBench, a 212-task benchmark built from 20 recent ML papers, finds that even the best LLM (Gemini-2.5-Pro-Preview) implements only 37.3% of the weighted code correctly.
calculate area
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
ResearchCodeBench, a 212-task benchmark built from 20 recent ML papers, finds that even the best LLM (Gemini-2.5-Pro-Preview) implements only 37.3% of the weighted code correctly.