Pith. sign in

REVIEW 1 cited by

JEC-QA: A Legal-Domain Question Answering Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.12011 v1 pith:56JEBOZF submitted 2019-11-27 cs.CL

classification cs.CL
keywords answeringdatasetjec-qaexaminationhumanslegalquestionreasoning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present JEC-QA, the largest question answering dataset in the legal domain, collected from the National Judicial Examination of China. The examination is a comprehensive evaluation of professional skills for legal practitioners. College students are required to pass the examination to be certified as a lawyer or a judge. The dataset is challenging for existing question answering methods, because both retrieving relevant materials and answering questions require the ability of logic reasoning. Due to the high demand of multiple reasoning abilities to answer legal questions, the state-of-the-art models can only achieve about 28% accuracy on JEC-QA, while skilled humans and unskilled humans can reach 81% and 64% accuracy respectively, which indicates a huge gap between humans and machines on this task. We will release JEC-QA and our baselines to help improve the reasoning ability of machine comprehension models. You can access the dataset from http://jecqa.thunlp.org/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding

    cs.CL 2025-06 conditional novelty 5.0 of 10

    EpiCoDe builds an extrapolated checkpoint from early and late finetuned models, then subtracts the late model's logits from the extrapolated model's logits during decoding to boost accuracy.

Pith tools