Pith. sign in

REVIEW 1 cited by

ChatGPT may Pass the Bar Exam soon, but has a Long Way to Go for the LexGLUE benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.12202 v1 pith:AEKBW34O submitted 2023-03-09 cs.CL

classification cs.CL
keywords chatgptlexgluemodelavailablebenchmarkdatasetsgpt-3micro-f1
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Following the hype around OpenAI's ChatGPT conversational agent, the last straw in the recent development of Large Language Models (LLMs) that demonstrate emergent unprecedented zero-shot capabilities, we audit the latest OpenAI's GPT-3.5 model, `gpt-3.5-turbo', the first available ChatGPT model, in the LexGLUE benchmark in a zero-shot fashion providing examples in a templated instruction-following format. The results indicate that ChatGPT achieves an average micro-F1 score of 47.6% across LexGLUE tasks, surpassing the baseline guessing rates. Notably, the model performs exceptionally well in some datasets, achieving micro-F1 scores of 62.8% and 70.2% in the ECtHR B and LEDGAR datasets, respectively. The code base and model predictions are available for review on https://github.com/coastalcph/zeroshot_lexglue.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A temporal mixture-of-experts model with cluster-based shift signals improves cross-time multilingual document classification and reveals language-specific temporal performance drops.

Pith tools