Pith. sign in

REVIEW 1 cited by

Overview of the BioLaySumm 2024 Shared Task on the Lay Summarization of Biomedical Research Articles

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08566 v1 pith:FNAF7UUL submitted 2024-08-16 cs.CL

classification cs.CL
keywords taskeditionresearchapproachesarticlesbiolaysummbiomedicalinterest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents the setup and results of the second edition of the BioLaySumm shared task on the Lay Summarisation of Biomedical Research Articles, hosted at the BioNLP Workshop at ACL 2024. In this task edition, we aim to build on the first edition's success by further increasing research interest in this important task and encouraging participants to explore novel approaches that will help advance the state-of-the-art. Encouragingly, we found research interest in the task to be high, with this edition of the task attracting a total of 53 participating teams, a significant increase in engagement from the previous edition. Overall, our results show that a broad range of innovative approaches were adopted by task participants, with a predictable shift towards the use of Large Language Models (LLMs).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Human-written plain language summaries led to significantly better reader comprehension than LLM-generated ones, despite similar subjective ratings, and most automated metrics did not predict comprehension.

Pith tools