Pith. sign in

REVIEW 1 cited by

BLESS: Benchmarking Large Language Models on Sentence Simplification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15773 v1 pith:CTIBV2EH submitted 2023-10-24 cs.CL

classification cs.CL
keywords llmsmodelsanalysisbenchmarkblessdifferenteditevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS). We examine how well off-the-shelf LLMs can solve this challenging task, assessing a total of 44 models, differing in size, architecture, pre-training methods, and accessibility, on three test sets from different domains (Wikipedia, news, and medical) under a few-shot setting. Our analysis considers a suite of automatic metrics as well as a large-scale quantitative investigation into the types of common edit operations performed by the different models. Furthermore, we perform a manual qualitative analysis on a subset of model outputs to better gauge the quality of the generated simplifications. Our evaluation indicates that the best LLMs, despite not being trained on TS, perform comparably with state-of-the-art TS baselines. Additionally, we find that certain LLMs demonstrate a greater range and diversity of edit operations. Our performance benchmark will be available as a resource for the development of future TS methods and evaluation metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Redefining Simplicity: Benchmarking Large Language Models from Lexical to Document Simplification

    cs.CL 2025-02 conditional novelty 4.0 of 10

    In a four-task benchmark, GPT-4o, Llama3.1-70B, and Gemma2-2B outperform traditional text simplification systems on most automatic metrics, and GPT-4o is preferred over human-written references in a small human study.

Pith tools