Pith. sign in

REVIEW 2 cited by

Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.13906 v1 pith:TN5JKKWV submitted 2023-06-24 cs.CL

classification cs.CL
keywords gpt-4domainexpertisehighlyperformancepredictionsspecializedtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We evaluated the capability of generative pre-trained transformers~(GPT-4) in analysis of textual data in tasks that require highly specialized domain expertise. Specifically, we focused on the task of analyzing court opinions to interpret legal concepts. We found that GPT-4, prompted with annotation guidelines, performs on par with well-trained law student annotators. We observed that, with a relatively minor decrease in performance, GPT-4 can perform batch predictions leading to significant cost reductions. However, employing chain-of-thought prompting did not lead to noticeably improved performance on this task. Further, we demonstrated how to analyze GPT-4's predictions to identify and mitigate deficiencies in annotation guidelines, and subsequently improve the performance of the model. Finally, we observed that the model is quite brittle, as small formatting related changes in the prompt had a high impact on the predictions. These findings can be leveraged by researchers and practitioners who engage in semantic/pragmatic annotations of texts in the context of the tasks requiring highly specialized domain expertise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are manual annotations necessary for statutory interpretations retrieval?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    With a large DeBERTa model, annotating 500 to 1000 sentences per legal concept matches full annotation, and LLM-based annotation (Qwen 2.5) achieves NDCG scores close to or better than human-annotation-trained models ...

  2. Structuring Radiology Reports: Challenging LLMs with Lightweight Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Fully finetuned T5 and BERT2BERT models match or beat prompt-adapted LLMs up to 70B parameters on radiology report structuring, at less than 1% of the inference cost.

Pith tools