Pith. sign in

REVIEW 1 cited by

Encode, Tag, Realize: High-Precision Text Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.01187 v1 pith:XU4CEQCK submitted 2019-09-03 cs.CL

classification cs.CL
keywords textapproachediteditingexampleslasertaggernumberoperations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose LaserTagger - a sequence tagging approach that casts text generation as a text editing task. Target texts are reconstructed from the inputs using three main edit operations: keeping a token, deleting it, and adding a phrase before the token. To predict the edit operations, we propose a novel model, which combines a BERT encoder with an autoregressive Transformer decoder. This approach is evaluated on English text on four tasks: sentence fusion, sentence splitting, abstractive summarization, and grammar correction. LaserTagger achieves new state-of-the-art results on three of these tasks, performs comparably to a set of strong seq2seq baselines with a large number of training examples, and outperforms them when the number of examples is limited. Furthermore, we show that at inference time tagging can be more than two orders of magnitude faster than comparable seq2seq models, making it more attractive for running in a live environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Algorithm for Automatic Legislative Text Consolidation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A LoRA-fine-tuned 13B language model can automatically consolidate French legislative texts, outperforming a span-extraction baseline and approaching GPT-4 on a subset of a real finance bill.

Pith tools