Pith. sign in

REVIEW 2 cited by

Neural CRF Model for Sentence Alignment in Text Simplification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.02324 v4 pith:2CBBZWII submitted 2020-05-05 cs.CL

classification cs.CL
keywords sentencesimplificationtextalignmentdatasetsmodelneuralquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of a text simplification system heavily depends on the quality and quantity of complex-simple sentence pairs in the training corpus, which are extracted by aligning sentences between parallel articles. To evaluate and improve sentence alignment quality, we create two manually annotated sentence-aligned datasets from two commonly used text simplification corpora, Newsela and Wikipedia. We propose a novel neural CRF alignment model which not only leverages the sequential nature of sentences in parallel documents but also utilizes a neural sentence pair model to capture semantic similarity. Experiments demonstrate that our proposed approach outperforms all the previous work on monolingual sentence alignment task by more than 5 points in F1. We apply our CRF aligner to construct two new text simplification datasets, Newsela-Auto and Wiki-Auto, which are much larger and of better quality compared to the existing datasets. A Transformer-based seq2seq model trained on our datasets establishes a new state-of-the-art for text simplification in both automatic and human evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The paper introduces Ace-CEFR, a dataset of short conversational English texts with expert CEFR labels, and shows that a fine-tuned BERT model predicts these labels more accurately than a single human expert.

  2. DLM-One: Diffusion Language Models for One-Step Sequence Generation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    DLM-One distills a continuous diffusion language model into a one-step student, achieving roughly 500x inference speedup while staying within a few percent of the teacher on BLEU, ROUGE, and BERTScore, with substantia...

Pith tools