REVIEW 4 major objections 6 minor 2 cited by
Legal Evalutions and Challenges of Large Language Models
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that OpenAI's O1-preview model produces the most human-aligned legal judgments among ten tested LLMs, and that automated text-overlap scores are a poor proxy for that quality.
desk verdict An unfinished bilingual legal-LLM eval whose headline ranking is not measurable from the reported evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluation protocol pairs two families of automated text-overlap metrics, ROUGE (n-gram recall between the generated judgment and the reference judgment) and BLEU (modified n-gram precision), with a 1-5 human evaluation in which law students trained in legal analysis score how well each model's decision output aligns with the legal reasoning and outcomes of real cases. The gap between these two measurement families is what carries the paper's argument that human judgment cannot be replaced by similarity scores.
What would settle it
A re-evaluation of the same 26 cases with a pre-registered scoring rubric, at least three independent legal-expert raters per output, and reported inter-annotator agreement would settle whether O1-preview truly outranks Qwen2-7B-Instruct; if the mean difference shrinks below the inter-rater standard deviation, the paper's headline ranking is not established.
Extended reading notes
Core claim
Across both languages, the O1-preview model achieved the highest overall human evaluation score of 3.96, demonstrating strong alignment with human judgment across diverse legal cases. Phi-3.5-mini-instruct posted the best overall ROUGE-1 and ROUGE-L scores (0.41) but a human score of only 2.62; lawyer-llama-13b-v2 had the best ROUGE-2 (0.28) yet scored 2.58. The paper takes this divergence as evidence that lexical-overlap metrics measure surface similarity, not the interpretative accuracy that matters in legal reasoning, and that general-purpose frontier models can outperform models fine-tuned on legal corpora in human-judged quality.
Load-bearing premise
The ranking rests on the assumption that the human evaluation protocol yields reliable measurements of legal judgment quality, but the paper reports no rubric, number of raters, per-case scores, or inter-annotator agreement, so the 0.11-point gap between O1-preview (3.96) and Qwen2-7B (3.85) could be rating noise.
Editorial extensions
If this is right
- If O1-preview's top human rating holds, it suggests that general-purpose reasoning-oriented models can outperform models fine-tuned on legal corpora for human-judged legal alignment.
- High ROUGE and BLEU scores do not imply legally sound output; evaluations of legal LLMs should include human or outcome-based assessment.
- The divergence between automated and human scores implies that current lexical-overlap benchmarks are not fit to rank legal reasoning quality.
- Legal-specific models such as LawGPT_zh and Lawyer-LLaMA, despite domain training, trail general models on human judgment, indicating that current legal fine-tuning methods have not yet delivered an advantage.
- Human evaluation of legal outputs needs greater methodological standardization before cross-model claims can be made with confidence.
Reading between the lines
- A plausible extension is to replace 1-5 holistic scores with a rubric distinguishing legal rule identification, factual application, and outcome prediction; that would show whether O1-preview's lead comes from one component or all.
- If the ranking is replicated by practicing lawyers, it would strengthen the case for using frontier reasoning models as drafting assistants in legal practice, but not for autonomous adjudication, since the paper itself emphasizes unexplained hallucinations and liability gaps.
- The correlation between human scores and ROUGE could be tested quantitatively from the reported tables: for example, Spearman rank correlation across the ten models would quantify how weak the relationship is, and the paper's own numbers likely show a negative or near-zero association.
- Because the Chinese and US cases differ in language, legal system, and document style, a cross-lingual factor analysis could reveal whether O1-preview's margin is driven mainly by its English performance (4.08) or holds equally in Chinese (3.85).
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a review of legal-domain LLM evaluation plus an original comparative study of ten models (closed-source, open-source, and legal-specific) on 13 Chinese and 13 English legal cases. For each case, the authors compute ROUGE-1/2/L and BLEU scores and obtain human evaluation ratings from law students on a 1--5 scale. The paper's central empirical claim is that O1-preview achieves the highest overall human evaluation score of 3.96 across both languages, followed by Qwen2-7B-Instruct at 3.85, and that automated lexical metrics do not predict these human preferences. The paper also surveys legal LLMs and discusses challenges such as data privacy, liability, ethics, and technical limitations. However, the human evaluation protocol is described only at a high level, no uncertainty or significance measures accompany the reported scores, and the presented results contain an internal inconsistency regarding lawyer-llama-13b-v2's Chinese versus English scores.
Significance. If the central ranking were well supported, the paper would provide a useful cross-lingual comparison of general and legal-specific LLMs on a realistically structured legal judgment task, with the interesting finding that ROUGE/BLEU are poor proxies for human-judged legal quality. The authors are also to be credited for including both Chinese and English civil, criminal, and administrative cases, for evaluating a heterogeneous set of open-source, closed-source, and domain-tuned models, and for using human evaluators as an external reference rather than relying solely on automated metrics. That said, the study's empirical value is currently limited because the evaluation data are not released, the scoring rubric and rater pool are unspecified, and no statistical measures support the pairwise differences that drive the conclusions. The paper contains no fitted parameters or circular derivations, so the circularity burden is low.
major comments (4)
- [Section IV-B, Tables I-III] The central claim that O1-preview outperforms all other models on human evaluation (overall 3.96 vs. Qwen2-7B-Instruct 3.85 in Table III) is not supported by the reported protocol. The manuscript states only that law students scored outputs on a 1--5 scale (Section IV-B, Human Evaluation Score); it does not report the number of raters, the rubric or scoring dimensions, whether raters were blind to model identity, per-case scores, or inter-annotator agreement. With 26 cases total and a 0.11-point gap, the difference could easily be rating noise. Without error bars, confidence intervals, or significance tests, the ranking in Tables I-III should be presented as descriptive only, and the wording ``demonstrating strong alignment with human judgment'' is not justified.
- [Section IV-B2, Table II] There is an internal contradiction in the English legal texts subsection. The text says lawyer-llama-13b-v2 received ``a noticeably lower score of 2.92 on Chinese texts compared to its English score of 2.23'', but Table I reports a Chinese score of 2.92 and Table II reports an English score of 2.23; since 2.92 > 2.23, the sentence inverts the direction of the comparison. This is a concrete reporting error in a passage that directly supports the cross-language analysis, and it undermines confidence in the accuracy of the tables and surrounding prose.
- [Section IV-A, IV-B] The experimental setup is underspecified in ways that affect the interpretation of the automated metrics and the human evaluation. The paper does not state the prompting strategy, decoding parameters, output length constraints, or whether the models were run zero-shot; it also does not describe how the reference texts for ROUGE/BLEU were constructed or whether the same reference judgments were used for all models. Because the paper's secondary claim is that ROUGE/BLEU scores do not predict human preference, these details are needed to rule out artifacts such as length bias or reference mismatch. Additionally, the dataset of 26 cases is not released, so the results are not reproducible.
- [Section I and Section IV] The paper frames itself as a review but introduces a new experiment without a clear statement of how the 26 cases were selected and whether the human raters were given any calibration or anchor examples. If the rater pool consisted of a small number of law students, the reported scores, which are averaged to two decimal places, imply a precision that the protocol cannot support. The authors should either provide the full rating data, inter-rater reliability measures, and a significance analysis, or explicitly limit the conclusions to qualitative observations about model behavior.
minor comments (6)
- [Title] The title contains a typo: ``Evalutions'' should be ``Evaluations''.
- [Section I] The sentence beginning ``As shown in Fig 1 Based on this background...'' is grammatically incomplete; it should be split or rephrased.
- [Section VII] The Acknowledgements section still contains the LaTeX placeholder text ``This should be a simple paragraph before the bibliography to thank those individuals and institutions who have supported your work on this article.'' This placeholder must be replaced before submission.
- [References] Several references are not in a consistent format; for example, some entries lack a publisher or venue (e.g., [8], [10]), and some online references rely on tinyurl redirects without a stable archive link. Please standardize the bibliography.
- [Section IV-B1] In the Chinese human evaluation results, the text states that GPT-4o, Qwen2-7B-Instruct, and O1-preview each scored 3.85, but Table I shows all three at 3.85; this is consistent, yet the passage immediately calls them ``the highest human evaluation scores'' without noting that several other models are statistically indistinguishable from these values given the absence of error bars.
- [Section III-C] The description of LexNLP as a ``legal language model'' is imprecise; LexNLP is an NLP toolkit for legal text processing, not an LLM. Please correct this characterization.
Circularity Check
No circularity: the paper's claims are empirical evaluations against external human ratings and reference texts, with no fitted inputs, derivation chain, or load-bearing self-citation.
full rationale
This paper is an empirical evaluation, not a derivation. The central claim—o1-preview achieving the highest overall human evaluation score of 3.96—rests on human ratings by law students comparing model outputs to real case judgments, which is an external reference. No parameter is fitted from the reported data, no equation is derived, and no prediction is produced from a fitted input. The ROUGE/BLEU scores are standard text-overlap metrics computed against external reference judgments, and the human evaluation scores are independent subjective judgments, not constructed from the models' outputs. Self-citations appear only in background sections (e.g., refs. [1], [2], [18]–[23]) and are not load-bearing for the legal evaluation results. The lack of rubric details, rater counts, and inter-annotator agreement is a measurement-validity concern, not a circularity concern. There is an internal contradiction in Section IV-B2: the text says lawyer-llama-13b-v2 received a 'noticeably lower score of 2.92 on Chinese texts compared to its English score of 2.23,' although 2.92 is higher than 2.23; this is a data-reporting error that undermines confidence in the human-evaluation tables, but it is not a circular step. No step in the paper's chain reduces to its own inputs by construction, and no load-bearing premise is justified solely by self-citation. Thus the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Law student ratings on an unvalidated 1-5 scale accurately measure alignment of model judgments with real legal reasoning.
- domain assumption The 26 selected cases (13 Chinese, 13 US) are representative of civil, criminal, and administrative practice in both jurisdictions.
- domain assumption ROUGE/BLEU computed against the true judgment is a meaningful similarity metric for legal judgment generation.
Cite this review
Pith. "Pith review of Legal Evalutions and Challenges of Large Language Models." pith.science (2026). https://pith.science/paper/PGVJJUKW
@misc{pith2026241110137,
author = {Pith},
title = {Pith review of: Legal Evalutions and Challenges of Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/PGVJJUKW}},
note = {Machine review of arXiv:2411.10137}
}
read the original abstract
In this paper, we review legal testing methods based on Large Language Models (LLMs), using the OPENAI o1 model as a case study to evaluate the performance of large models in applying legal provisions. We compare current state-of-the-art LLMs, including open-source, closed-source, and legal-specific models trained specifically for the legal domain. Systematic tests are conducted on English and Chinese legal cases, and the results are analyzed in depth. Through systematic testing of legal cases from common law systems and China, this paper explores the strengths and weaknesses of LLMs in understanding and applying legal texts, reasoning through legal issues, and predicting judgments. The experimental results highlight both the potential and limitations of LLMs in legal applications, particularly in terms of challenges related to the interpretation of legal language and the accuracy of legal reasoning. Finally, the paper provides a comprehensive analysis of the advantages and disadvantages of various types of models, offering valuable insights and references for the future application of AI in the legal field.
Figures
Forward citations
Cited by 2 Pith papers
-
LaQual: An Automated Framework for LLM App Quality Evaluation
LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.
-
QueEn: A Large Language Model for Quechua-English Translation
QueEn reports BLEU 17.6 for Quechua-English translation, but its internal table shows BLEU 0.235, the method is not reproducible, and the translation direction is inconsistent.
Reference graph
Works this paper leans on
-
[1]
Review of large vision models and visual prompt engineering,
J. Wang, Z. Liu, L. Zhao, Z. Wu, C. Ma, S. Yu, H. Dai, Q. Yang, Y . Liu, S. Zhang et al. , “Review of large vision models and visual prompt engineering,” Meta-Radiology, p. 100047, 2023
2023
-
[2]
J. Wang, H. Jiang, Y . Liu, C. Ma, X. Zhang, Y . Pan, M. Liu, P. Gu, S. Xia, W. Li et al. , “A comprehensive review of multimodal large language models: Performance and challenges across different tasks,” arXiv preprint arXiv:2408.01319 , 2024. JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, SEPTEMBER 2020 9
arXiv 2024
-
[3]
Language models are few-shot learners,
T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020
arXiv 2005
-
[4]
Attention is all you need,
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, 2017
2017
-
[5]
Improving language understanding by generative pre- training,
A. Radford, “Improving language understanding by generative pre- training,” 2018
2018
-
[6]
Large language models in law: A survey,
J. Lai, W. Gan, J. Wu, Z. Qi, and S. Y . Philip, “Large language models in law: A survey,” AI Open, 2024
work page 2024
-
[7]
The judge, the ai, and the crown: a collusive network,
B. Chaudhary, P. Covarrubia, and G. Y . Ng, “The judge, the ai, and the crown: a collusive network,”Information & Communications Technology Law, vol. 33, no. 3, pp. 330–367, 2024
work page 2024
-
[8]
Lawgpt: A chinese legal knowledge-enhanced large language model,
Z. Zhou, J.-X. Shi, P.-X. Song, X.-W. Yang, Y .-X. Jin, L.-Z. Guo, and Y .-F. Li, “Lawgpt: A chinese legal knowledge-enhanced large language model,” arXiv preprint arXiv:2406.04614 , 2024
arXiv 2024
Show all 59 references
-
[9]
Chatgpt may pass the bar exam soon, but has a long way to go for the lexglue benchmark,
I. Chalkidis, “Chatgpt may pass the bar exam soon, but has a long way to go for the lexglue benchmark,” arXiv preprint arXiv:2304.12202, 2023
2023 arXiv
-
[10]
Xiezhi: Chinese law large language model,
L. Hongcheng, L. Yusheng, M. Yutong, and Y . Wang, “Xiezhi: Chinese law large language model,” https://github.com/LiuHC0428/LAW GPT, 2023
2023
-
[11]
Language in the legal process,
B. Danet, “Language in the legal process,” Law & Society Review , vol. 14, no. 3, pp. 445–564, 1980
1980
-
[12]
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models,
N. Guha, J. Nyarko, D. Ho, C. R ´e, A. Chilton, A. Chohlas-Wood, A. Peters, B. Waldon, D. Rockmore, D. Zambrano et al., “Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models,” Advances in Neural Information Processing Systems , v...
2024
-
[13]
Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects,
M. U. Hadi, Q. Al Tashi, A. Shah, R. Qureshi, A. Muneer, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wuet al., “Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects,” Authorea Preprints, 2024
2024
-
[14]
Chatgpt: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope,
P. P. Ray, “Chatgpt: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope,” Internet of Things and Cyber-Physical Systems , vol. 3, pp. 121–154, 2023
2023
-
[15]
Safeguarding human values: rethinking us law for generative ai’s societal impacts,
I. Cheong, A. Caliskan, and T. Kohno, “Safeguarding human values: rethinking us law for generative ai’s societal impacts,” AI and Ethics , pp. 1–27, 2024
2024
-
[16]
Reviewing the ethical implications of ai in decision making processes,
F. Osasona, O. O. Amoo, A. Atadoga, T. O. Abrahams, O. A. Faray- ola, and B. S. Ayinla, “Reviewing the ethical implications of ai in decision making processes,” International Journal of Management & Entrepreneurship Research, vol. 6, no. 2, pp. 322–335, 2024
2024
-
[17]
Legal challenges of artificial intelligence and robotics: a comprehensive review,
C. U. Akpuokwe, A. O. Adeniyi, and S. S. Bakare, “Legal challenges of artificial intelligence and robotics: a comprehensive review,” Computer Science & IT Research Journal , vol. 5, no. 3, pp. 544–561, 2024
2024
-
[18]
Eg-spikeformer: Eye-gaze guided transformer on spiking neural networks for medical image analysis,
Y . Pan, H. Jiang, J. Chen, Y . Li, H. Zhao, Y . Zhou, P. Shu, Z. Wu, Z. Liu, D. Zhu et al. , “Eg-spikeformer: Eye-gaze guided transformer on spiking neural networks for medical image analysis,” arXiv preprint arXiv:2410.09674, 2024
-
[19]
Echopulse: Ecg controlled echocardio-grams video generation,
Y . Li, S. Kim, Z. Wu, H. Jiang, Y . Pan, P. Jin, S. Song, Y . Shi, T. Yang, T. Liu et al. , “Echopulse: Ecg controlled echocardio-grams video generation,” arXiv preprint arXiv:2410.03143 , 2024
2024 arXiv
-
[20]
Mgh radiology llama: A llama 3 70b model for radiology,
Y . Shi, P. Shu, Z. Liu, Z. Wu, Q. Li, and X. Li, “Mgh radiology llama: A llama 3 70b model for radiology,” arXiv preprint arXiv:2408.11848 , 2024
2024 arXiv
-
[21]
Llms for coding and robotics education,
P. Shu, H. Zhao, H. Jiang, Y . Li, S. Xu, Y . Pan, Z. Wu, Z. Liu, G. Lu, L. Guan et al., “Llms for coding and robotics education,” arXiv preprint arXiv:2402.06116, 2024
2024 arXiv
-
[22]
Assessing large language models in mechanical engineering education: A study on mechanics-focused conceptual un- derstanding,
J. Tian, J. Hou, Z. Wu, P. Shu, Z. Liu, Y . Xiang, B. Gu, N. Filla, Y . Li, N. Liu et al. , “Assessing large language models in mechanical engineering education: A study on mechanics-focused conceptual un- derstanding,” arXiv preprint arXiv:2401.12983 , 2024
2024 arXiv
-
[23]
Multimodality of ai for education: To- wards artificial general intelligence,
G.-G. Lee, L. Shi, E. Latif, Y . Gao, A. Bewersdorff, M. Nyaaba, S. Guo, Z. Wu, Z. Liu, H. Wang et al., “Multimodality of ai for education: To- wards artificial general intelligence,” arXiv preprint arXiv:2312.06037 , 2023
2023 arXiv
-
[24]
The use of large language mod- els in legaltech,
N. Shaver, “The use of large language mod- els in legaltech,” 2023, accessed: 2024-10-27. [On- line]. Available: https://www.legaltechnologyhub.com/contents/ the-use-of-large-language-models-in-legaltech/
2023
-
[25]
Summize uses openai to supercharge contract summaries with gpt-3.5,
T. Dunlop, “Summize uses openai to supercharge contract summaries with gpt-3.5,” 2023, accessed: 2024-10-27. [Online]. Available: https://www.summize.com/resources/summize-and-chatgpt
2023
-
[26]
Docket alarm incorporates gpt-3.5 to auto-summarize pdf litigation filings in complex dockets,
S. Wilkins, “Docket alarm incorporates gpt-3.5 to auto-summarize pdf litigation filings in complex dockets,” 2023, accessed: 2024-10-27. [Online]. Available: https://tinyurl.com/5n7yvtzd
2023
-
[27]
Docket alarm incorporates gpt-3.5 to auto-summarize pdf litigation filings in complex dockets,
L. M. D. Droit, “Docket alarm incorporates gpt-3.5 to auto-summarize pdf litigation filings in complex dockets,” 2023, accessed: 2024- 10-27. [Online]. Available: https://www.lemondedudroit.fr/professions/ 337-legaltech/85951-chatgpt-predictice-integremoteur-recherche.html
2023
-
[28]
Navigating ai for legal documents: Tips & tricks,
T. Hassonjee, “Navigating ai for legal documents: Tips & tricks,” 2024, accessed: 2024-10-27. [Online]. Available: https://www.docdraft. ai/blogs/navigating-ai-for-legal-documents-tips-tricks
2024
-
[29]
Lexisnexis to acquire henchman: A new chapter in legal drafting,
Henchman, “Lexisnexis to acquire henchman: A new chapter in legal drafting,” 2024, accessed: 2024-10-27. [Online]. Available: https://henchman.io/blog/lexisnexis-announcement
2024
-
[30]
Colin lachance on jurisage’s myjr and how he’s looking at ai to assist in the synthesis and reading of legal cases,
G. Lambert, “Colin lachance on jurisage’s myjr and how he’s looking at ai to assist in the synthesis and reading of legal cases,” 2023, accessed: 2024-10-27. [Online]. Available: https://tinyurl.com/4tnd29uv
2023
-
[31]
Ask blue j enhances user experience through gpt-4 and conversational,
B. J, “Ask blue j enhances user experience through gpt-4 and conversational,” 2023, accessed: 2024-10-27. [Online]. Available: https://www.bluej.com/blog/ask-blue-j-enhances-user-experience
2023
-
[32]
Lexglue: A benchmark dataset for legal language understanding in english,
I. Chalkidis, A. Jana, D. Hartung, M. Bommarito, I. Androutsopoulos, D. M. Katz, and N. Aletras, “Lexglue: A benchmark dataset for legal language understanding in english,” arXiv preprint arXiv:2110.00976 , 2021
2021 arXiv
-
[33]
How ready are pre-trained abstrac- tive models and llms for legal case judgement summarization?
A. Deroy, K. Ghosh, and S. Ghosh, “How ready are pre-trained abstrac- tive models and llms for legal case judgement summarization?” arXiv preprint arXiv:2306.01248, 2023
2023 arXiv
-
[34]
Explaining legal concepts with augmented large language models (gpt- 4),
J. Savelka, K. D. Ashley, M. A. Gray, H. Westermann, and H. Xu, “Explaining legal concepts with augmented large language models (gpt- 4),” arXiv preprint arXiv:2306.09525 , 2023
2023 arXiv
-
[35]
Garbage in, garbage out: Zero-shot detection of crime using large language models,
A. Simmons and R. Vasa, “Garbage in, garbage out: Zero-shot detection of crime using large language models,” arXiv preprint arXiv:2307.06844, 2023
2023 arXiv
-
[36]
Lawyer llama technical report,
Q. Huang, M. Tao, C. Zhang, Z. An, C. Jiang, Z. Chen, Z. Wu, and Y . Feng, “Lawyer llama technical report,” 2023
2023
-
[37]
Lawyer llama,
——, “Lawyer llama,” https://github.com/AndrewZhe/lawyer-llama, 2023
2023
-
[38]
Lexilaw: a chinese legal large language model,
LexiLaw, “Lexilaw: a chinese legal large language model,” 2023, accessed: 2024-10-27. [Online]. Available: https://github.com/CSHaitao/ LexiLaw?tab=readme-ov-file#readme
2023
-
[39]
Lexgpt 0.1: pre-trained gpt-j models with pile of law,
J.-S. Lee, “Lexgpt 0.1: pre-trained gpt-j models with pile of law,” arXiv preprint arXiv:2306.05431, 2023
2023 arXiv
-
[40]
Chatlaw: Open-source legal large language model with integrated external knowledge bases,
J. Cui, Z. Li, Y . Yan, B. Chen, and L. Yuan, “Chatlaw: Open-source legal large language model with integrated external knowledge bases,” arXiv preprint arXiv:2306.16092 , 2023
2023 arXiv
-
[41]
Disc-lawllm: Fine-tuning large language models for intelligent legal services,
S. Yue, W. Chen, S. Wang, B. Li, C. Shen, S. Liu, Y . Zhou, Y . Xiao, S. Yun, X. Huang et al. , “Disc-lawllm: Fine-tuning large language models for intelligent legal services,” arXiv preprint arXiv:2309.11325, 2023
2023 arXiv
-
[42]
Meet kl3m: the first legal large language model,
KL3M, “Meet kl3m: the first legal large language model,” 2024, accessed: 2024-10-27. [Online]. Available: https://273ventures.com/ kl3m-the-first-legal-large-language-model/#note1
2024
-
[43]
Internlm-law: An open source chinese legal large language model,
Z. Fei, S. Zhang, X. Shen, D. Zhu, X. Wang, M. Cao, F. Zhou, Y . Li, W. Zhang, D. Lin et al. , “Internlm-law: An open source chinese legal large language model,” arXiv preprint arXiv:2406.14887 , 2024
2024 arXiv
-
[44]
Saullm-54b & saullm-141b: Scaling up domain adaptation for the legal domain,
P. Colombo, T. Pires, M. Boudiaf, R. Melo, D. Culver, S. Morgado, E. Malaboeuf, G. Hautreux, J. Charpentier, and M. Desa, “Saullm-54b & saullm-141b: Scaling up domain adaptation for the legal domain,” arXiv preprint arXiv:2407.19584 , 2024
2024 arXiv
-
[45]
wisdominterrogatory,
Y . Wu, Y . Liu, Y . Liu, A. Li, S. Zhou, and K. Kuang, “wisdominterrogatory,” 2024, accessed: 2024-10-27. [Online]. Available: https://github.com/zhihaiLLM/wisdomInterrogatory
2024
-
[46]
Syllogistic reasoning for legal judgment analysis,
W. Deng, J. Pei, K. Kong, Z. Chen, F. Wei, Y . Li, Z. Ren, Z. Chen, and P. Ren, “Syllogistic reasoning for legal judgment analysis,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: As...
2023
-
[47]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023 arXiv
-
[48]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican et al., “Gemini: a family of highly capable multimodal models,”arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[49]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,
G. Team, P. Georgiev, V . I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang et al., “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530, 2024
2024 arXiv
-
[50]
Claude 3.5 sonnet model card addendum,
A. Anthropic, “Claude 3.5 sonnet model card addendum,” Claude-3.5 Model Card, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, SEPTEMBER 2020 10
2024
-
[51]
Yi: Open foundation models by 01. ai,
A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang et al., “Yi: Open foundation models by 01. ai,” arXiv preprint arXiv:2403.04652, 2024
2024 arXiv
-
[52]
The llama 3 herd of models,
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan et al. , “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783 , 2024
2024 arXiv
-
[53]
Mistral 7b,
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier et al., “Mistral 7b,” arXiv preprint arXiv:2310.06825 , 2023
2023 arXiv
-
[54]
Gemma: Open models based on gemini research and technology,
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Love et al. , “Gemma: Open models based on gemini research and technology,” arXiv preprint arXiv:2403.08295, 2024
2024 arXiv
-
[55]
Gemma 2: Improving open language models at a practical size,
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ram ´e et al. , “Gemma 2: Improving open language models at a practical size,” arXiv preprint arXiv:2408.00118, 2024
2024 arXiv
-
[56]
Phi-3 technical report: A highly capable language model locally on your phone,
M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl et al., “Phi-3 technical report: A highly capable language model locally on your phone,” arXiv preprint arXiv:2404.14219, 2024
2024 arXiv
-
[57]
Qwen2 technical report,
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang et al. , “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024
2024 arXiv
-
[58]
Chatglm: A family of large language models from glm-130b to glm-4 all tools,
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao et al. , “Chatglm: A family of large language models from glm-130b to glm-4 all tools,” arXiv preprint arXiv:2406.12793, 2024
2024 arXiv
-
[59]
Lexnlp: Natural language processing and information extraction for legal and regulatory texts,
M. J. Bommarito II, D. M. Katz, and E. M. Detterman, “Lexnlp: Natural language processing and information extraction for legal and regulatory texts,” in Research handbook on big data law. Edward Elgar Publishing, 2021, pp. 216–227
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.