Pith. sign in

REVIEW 1 cited by

Disco-Bench: A Discourse-Aware Evaluation Benchmark for Language Modelling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.08074 v2 pith:67IKDQ3G submitted 2023-07-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords discoursedisco-benchevaluationmodelsbenchmarklanguagephenomenadocument-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modeling discourse -- the linguistic phenomena that go beyond individual sentences, is a fundamental yet challenging aspect of natural language processing (NLP). However, existing evaluation benchmarks primarily focus on the evaluation of inter-sentence properties and overlook critical discourse phenomena that cross sentences. To bridge the gap, we propose Disco-Bench, a benchmark that can evaluate intra-sentence discourse properties across a diverse set of NLP tasks, covering understanding, translation, and generation. Disco-Bench consists of 9 document-level testsets in the literature domain, which contain rich discourse phenomena (e.g. cohesion and coherence) in Chinese and/or English. For linguistic analysis, we also design a diagnostic test suite that can examine whether the target models learn discourse knowledge. We totally evaluate 20 general-, in-domain and commercial models based on Transformer, advanced pretraining architectures and large language models (LLMs). Our results show (1) the challenge and necessity of our evaluation benchmark; (2) fine-grained pretraining based on literary document-level training data consistently improves the modeling of discourse information. We will release the datasets, pretrained models, and leaderboard, which we hope can significantly facilitate research in this field: https://github.com/longyuewangdcu/Disco-Bench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    The WMT 2024 literary translation shared task finds that domain-enhanced systems lead in d-BLEU for Chinese-English, but human evaluators rank NLP2CT-UM and SJTU-LoveFiction at the top.

Pith tools