Pith. sign in

REVIEW 1 cited by

Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07075 v1 pith:MZQTV7Z7 submitted 2023-06-12 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords legalllmscapabilitiesexampleslanguagelargelevelsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Better understanding of Large Language Models' (LLMs) legal analysis abilities can contribute to improving the efficiency of legal services, governing artificial intelligence, and leveraging LLMs to identify inconsistencies in law. This paper explores LLM capabilities in applying tax law. We choose this area of law because it has a structure that allows us to set up automated validation pipelines across thousands of examples, requires logical reasoning and maths skills, and enables us to test LLM capabilities in a manner relevant to real-world economic lives of citizens and companies. Our experiments demonstrate emerging legal understanding capabilities, with improved performance in each subsequent OpenAI model release. We experiment with retrieving and utilising the relevant legal authority to assess the impact of providing additional legal context to LLMs. Few-shot prompting, presenting examples of question-answer pairs, is also found to significantly enhance the performance of the most advanced model, GPT-4. The findings indicate that LLMs, particularly when combined with prompting enhancements and the correct legal texts, can perform at high levels of accuracy but not yet at expert tax lawyer levels. As LLMs continue to advance, their ability to reason about law autonomously could have significant implications for the legal profession and AI governance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Using Large Language Models for Legal Decision-Making in Austrian Value-Added Tax Law: An Experimental Study

    cs.CL 2025-07 conditional novelty 6.0 of 10

    RAG-enhanced LLMs slightly outperform fine-tuned LLMs on Austrian/EU VAT questions, but not significantly, and neither approach is ready for full automation.

Pith tools