Otter generates fail-to-pass tests from issue descriptions alone, reaching 37% success with an ensemble, and the tests can lift SWE agent patch precision from about 61% to 92% at reduced recall.
Write one file name in each line
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Otter: Generating Tests from Issues to Validate SWE Patches
Otter generates fail-to-pass tests from issue descriptions alone, reaching 37% success with an ensemble, and the tests can lift SWE agent patch precision from about 61% to 92% at reduced recall.