Pith. sign in

REVIEW 1 cited by

SWiPE: A Dataset for Document-Level Simplification of Wikipedia Pages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19204 v1 pith:DPEPAPMS submitted 2023-05-30 cs.CL

classification cs.CL
keywords editssimplificationwikipediadocument-levelmodelsswipeworkarticles
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text simplification research has mostly focused on sentence-level simplification, even though many desirable edits - such as adding relevant background information or reordering content - may require document-level context. Prior work has also predominantly framed simplification as a single-step, input-to-output task, only implicitly modeling the fine-grained, span-level edits that elucidate the simplification process. To address both gaps, we introduce the SWiPE dataset, which reconstructs the document-level editing process from English Wikipedia (EW) articles to paired Simple Wikipedia (SEW) articles. In contrast to prior work, SWiPE leverages the entire revision history when pairing pages in order to better identify simplification edits. We work with Wikipedia editors to annotate 5,000 EW-SEW document pairs, labeling more than 40,000 edits with proposed 19 categories. To scale our efforts, we propose several models to automatically label edits, achieving an F-1 score of up to 70.6, indicating that this is a tractable but challenging NLU task. Finally, we categorize the edits produced by several simplification models and find that SWiPE-trained models generate more complex edits while reducing unwanted edits.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Progressive Document-level Text Simplification via Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A three-stage hierarchical LLM pipeline for document simplification outperforms direct ChatGPT prompts and earlier methods on Wiki-auto and Newsela, with caveats about self-evaluation.

Pith tools