An error taxonomy and GPT-4o auto-evaluator show that LLMs answering Civil Procedure MCQs have high accuracy but much lower soundness and correctness, with misinterpretation as the dominant step-level error.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
An error taxonomy and GPT-4o auto-evaluator show that LLMs answering Civil Procedure MCQs have high accuracy but much lower soundness and correctness, with misinterpretation as the dominant step-level error.