REVIEW 2 major objections 6 minor 85 references
Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper proposes DiscoEval, a seven-task benchmark for whether sentence representations encode discourse context, and shows that Wikipedia-derived training objectives and deep BERT/ELMo layers capture different aspects of it.
desk verdict DiscoEval is a genuinely useful discourse-focused probing suite plus Wikipedia-derived training objectives, honestly evaluated; its main weakness is that several tasks may reward lexical/register cues rather than broader context, and the paper provides no surface baselines to rule that out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is DiscoEval itself: a collection of seven probing task groups in which a frozen sentence encoder produces vectors and a logistic-regression classifier (one task uses a 2000-unit hidden layer) must predict discourse-related targets—RST and PDTB relations, a shuffled sentence's original position among five, the correct order of two sentences, whether six sentences form a coherent passage, and whether a sentence comes from an abstract. The companion machinery is a multi-task training setup with one Bi-GRU encoder and decoders predicting neighboring sentences, nesting level, sentence and paragraph position, and section and document titles. DiscoEval supplies the measurement; the losses supply the learning signal.
What would settle it
On the same frozen embeddings, replace the logistic-regression probe with a single-hidden-layer probe (as the paper itself already does for the Discourse Coherence task) on PDTB-I and RST-DT. If accuracy rises substantially above the Table 2 numbers, DiscoEval's ranking reflects linear accessibility, not whether discourse information is present in the representation.
Extended reading notes
Core claim
DiscoEval is presented as a valid probe of whether sentence representations include discourse information. Across seven task groups built from Wikipedia, stories, dialogues, and scientific papers, the paper benchmarks fixed sentence encoders and finds that BERT-Large has the highest average DiscoEval accuracy, that ELMo is competitive, and that models trained with neighboring-sentence prediction plus two structure-aware losses—nesting level combined with sentence and paragraph position—improve over the neighboring-sentence baseline, while the section-and-document-title loss tends to hurt position-related tasks. Per-layer analysis shows that for BERT and ELMo, deeper layers perform best on DiscoEval, in contrast to sentence-level evaluation where shallower layers lead.
Load-bearing premise
The load-bearing premise is that a logistic-regression classifier trained on frozen sentence vectors tells us what a representation 'includes': if discourse information is present but not linearly separable, DiscoEval will score the encoder low even though the information is encoded.
Editorial extensions
If this is right
- Sentence-embedding evaluation that stops at standalone semantics misses a whole axis: the same representation that looks weak on semantic similarity can be strong on discourse tasks, so future benchmarks should include both.
- Discourse-aware training does not require expensive annotation: Wikipedia's headings, positions, and nesting levels give free labels, and combining position and nesting losses is better than either alone.
- For BERT and ELMo, the best layer for discourse probing is consistently deep, so applications that need discourse structure should select layers accordingly rather than defaulting to the pooled output.
- The section-and-document-title objective's negative effect on position and ordering tasks suggests a real trade-off between topical invariance and positional sensitivity in learned sentence vectors.
- DiscoEval gives a shared measurement that future discourse-aware encoders can be compared against, alongside the released preprocessing and evaluation scripts.
Reading between the lines
- If the linear-probe assumption is relaxed, some encoders now ranked low on DiscoEval might turn out to encode discourse information nonlinearly; reporting both linear and nonlinear probe scores would separate representation content from probe accessibility.
- The layer-depth result suggests discourse structure is a higher-level abstraction in contextual encoders; one testable consequence is that freezing deeper BERT layers should help downstream discourse-heavy tasks such as summarization and coherence ranking.
- The same Wikipedia-derived objectives could transfer to other structured corpora with section headers or positions, such as legal or scientific documents, and the DiscoEval tasks could be re-run there to test generality.
- A simple extension would be to anneal or gate the section-title loss during training so the model gets topical signal early without erasing the positional distinctions that sentence-position tasks require.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DiscoEval, a suite of seven probing tasks for evaluating whether sentence representations encode discourse-related information: discourse relations (PDTB-E, PDTB-I, RST-DT), sentence position (SP), binary sentence ordering (BSO), discourse coherence (DC), and sentence section prediction (SSP). It also proposes auxiliary training objectives for sentence encoders trained on Wikipedia: NSP, nesting level (NL), sentence/paragraph position (SPP), and section/document title prediction (SDT). The authors benchmark their models against Skip-thought, InferSent, DisSent, ELMo, and BERT on DiscoEval and SentEval. The main empirical findings are that the NL+SPP combination gives a small average improvement over the NSP baseline, SDT tends to hurt discourse-task performance, and BERT/ELMo perform strongly, with deeper layers better on DiscoEval than on SentEval.
Significance. If the benchmark's validity were established, DiscoEval would fill a real gap: existing sentence-embedding benchmarks mostly test stand-alone sentence meaning, and a discourse-focused suite with public code and data would be widely used. The paper is commendably transparent about negative results, makes data and scripts available, and provides a per-layer analysis that is a useful reference. The main risk is that the tasks may partly measure sentence-internal lexical or register cues rather than broader-context information, and the training-objective comparisons lack variance estimates. The contribution is promising as a resource and an empirical study, but the central claims need additional support.
major comments (2)
- [Section 3.2, Section 3.5, Table 6] The benchmark's central claim is that DiscoEval evaluates whether sentence representations include broader context information (Abstract). This is not yet established. In SSP (Section 3.5), the classifier receives only the sentence embedding, and the authors themselves note that a within-sentence cue such as 'Empirically' is predictive of the Abstract section; the task can therefore be solved as sentence-internal register classification. In SP (Table 6), the 'w/o context' baseline reaches 43.2% versus 20% random, and adding the four surrounding sentences yields only 47.3%, so most of the signal is sentence-internal. I ask for non-neural surface baselines (bag-of-words, TF-IDF, n-grams, and randomized embeddings) on every DiscoEval task so the context-dependent component is quantified; without them, the high scores of BERT/ELMo and the training-objective effects are underdetermined.
- [Section 5.2, Table 2] The conclusions about the proposed training objectives rest on small differences with no error bars or significance tests. The NL+SPP model beats the NSP baseline by 0.7% on the DiscoEval average, and the claim that SDT hurts performance rests on similarly small margins. Moreover, the clearest gains for NL+SPP are on SP and BSO derived from Wikipedia, the same corpus used for the proposed training objectives, while the external human-annotated tasks show mixed or negligible effects (e.g., PDTB-I 39.1 to 40.5; RST-DT 56.7 to 56.4 for the NSP baseline vs. NL+SPP). Please report standard deviations across classifier seeds and a breakdown separating Wikipedia-derived tasks from external tasks, and temper the generalization claims accordingly.
minor comments (6)
- [Appendix A] The citations for GloVe and Adam appear as unresolved placeholders ('?'); these need to be filled in, and the typo 'intialize' should be corrected.
- [Human Evaluation, Table 5] The human evaluation is based on a single native-English-speaking annotator with 50 examples per domain and no inter-annotator reliability information; the human accuracies should be presented as indicative rather than definitive.
- [Section 3.4, Table 2] The DC task uses a 2000-unit hidden layer while all other tasks use logistic regression; this difference in classifier capacity should be flagged in the main results table so readers do not compare the DC column directly with the other columns.
- [Section 5.1] For the main BERT and ELMo results the paper averages across layers, but the per-layer analysis uses [CLS] for BERT and layer-specific averages for ELMo; please clarify whether the pooling is otherwise identical so the layer comparison is clean.
- [Figure 8] Because each column is standardized separately, cross-column differences in raw accuracy are not visually comparable; adding raw values or a shared color scale would help.
- [Section 3.2, Table 6] The 'w/o context' condition is not defined precisely (input x1 only, or the full input with the x1 - xi terms removed?); please state the exact input representation.
Circularity Check
No circularity: DiscoEval is an empirical benchmark with held-out test sets; no prediction reduces to its inputs.
full rationale
The paper's central claims are empirical: DiscoEval evaluates discourse-related knowledge in sentence representations, and the proposed Wikipedia-derived training objectives (NL, SPP, SDT) help encoders capture different discourse aspects. These claims are supported by classifier evaluations on held-out test splits of multiple datasets, including external human-annotated resources (PDTB, RST-DT) and the independent SentEval suite. No equation in the paper defines a predicted quantity in terms of the fitted inputs; the training objectives and the evaluation tasks are distinct constructions, and the reported numbers are measured accuracies rather than identities. Some DiscoEval tasks (SP-Wiki, BSO-Wiki, DC-Wiki) are drawn from Wikipedia, the same corpus used to train the proposed objectives, but the splits are held out and the benchmark also includes ROC, arXiv, Ubuntu, and human-annotated discourse datasets, so this is a domain-overlap limitation, not a reduction by construction. The only self-citation (Chen et al., 2019) appears in a list of prior probing-task work and is not load-bearing for any conclusion. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. The skeptical concern about surface lexical cues is a construct-validity question, not a circularity of the derivation chain, and the paper even reports a 'w/o context' ablation on Sentence Position. Therefore the analysis is self-contained and non-circular.
Assumptions & free parameters
free parameters (3)
- PDTB category frequency threshold =
10 instances
- DC classifier hidden layer size =
2000
- Context window sizes =
5 sentences for SP/BSO, 6 for DC
assumptions (5)
- domain assumption The distributional semantics assumption: text statistics in large corpora carry semantic and discourse information.
- domain assumption Probing with simple linear classifiers is a valid way to measure what sentence representations encode.
- domain assumption Wikipedia document structure (section titles, nesting levels, sentence/paragraph positions) provides a useful discourse signal for general-purpose sentence representations.
- domain assumption RST-DT can be evaluated by averaging EDU representations and using right-branching binarization.
- ad hoc to paper The PDTB category filtering threshold (fewer than 10 instances removed) does not bias the relation distribution.
Cite this review
Pith. "Pith review of Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations." pith.science (2026). https://pith.science/paper/5YVWUVCU
@misc{pith2026190900142,
author = {Pith},
title = {Pith review of: Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/5YVWUVCU}},
note = {Machine review of arXiv:1909.00142}
}
read the original abstract
Prior work on pretrained sentence embeddings and benchmarks focus on the capabilities of stand-alone sentences. We propose DiscoEval, a test suite of tasks to evaluate whether sentence representations include broader context information. We also propose a variety of training objectives that makes use of natural annotations from Wikipedia to build sentence encoders capable of modeling discourse. We benchmark sentence encoders pretrained with our proposed training objectives, as well as other popular pretrained sentence encoders on DiscoEval and other sentence evaluation tasks. Empirically, we show that these training objectives help to encode different aspects of information in document structures. Moreover, BERT and ELMo demonstrate strong performances over DiscoEval with individual hidden layers showing different characteristics.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017. Fine-grained analysis of sentence embeddings using auxiliary prediction tasks. In Proceedings of ICLR
2017
-
[2]
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe. 2015. https://doi.org/10.18653/v1/S15-2045 Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability . In Proceeding...
-
[3]
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014. https://doi.org/10.3115/v1/S14-2010 Semeval-2014 task 10: Multilingual semantic textual similarity . In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pages 81--91. ...
-
[4]
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016. https://doi.org/10.18653/v1/S16-1081 Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation . In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), pages 497--...
-
[5]
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012. http://aclweb.org/anthology/S12-1051 Semeval-2012 task 6: A pilot on semantic textual similarity . In *SEM 2012: The First Joint Conference on Lexical and Computational Semantics -- Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth ...
2012
-
[6]
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. http://aclweb.org/anthology/S13-1004 *sem 2013 shared task: Semantic textual similarity . In Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity, pages 32--43. As...
2013
-
[7]
Regina Barzilay and Mirella Lapata. 2008. Modeling local coherence: An entity-based approach. Computational Linguistics, 34(1):1--34
2008
-
[8]
Yonatan Belinkov, Llu \' s M \`a rquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017. https://www.aclweb.org/anthology/I17-1001 Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks . In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volu...
2017
Show all 85 references
-
[9]
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, and Yejin Choi. 2018. https://doi.org/10.18653/v1/N18-1016 Discourse-aware neural rewards for coherent text generation . In Proceedings of the 2018 Conference of the North A merican Chapter of the Ass...
2018 doi
-
[10]
Lynn Carlson, Daniel Marcu, and Mary Ellen Okurovsky. 2001. https://www.aclweb.org/anthology/W01-1605 Building a discourse-tagged corpus in the framework of rhetorical structure theory . In Proceedings of the Second SIG dial Workshop on Discourse and Dialogue
2001
-
[11]
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017. https://doi.org/10.18653/v1/S17-2001 Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation . In Proceedings of the 11th International Workshop on Semant...
2017 doi
-
[12]
Mingda Chen, Zewei Chu, Karl Stratos, and Kevin Gimpel. 2019. Enteval: A holistic evaluation benchmark for entity representations. In Proc. of EMNLP
2019
-
[13]
Xinchi Chen, Xipeng Qiu, and Xuanjing Huang. 2016. Neural sentence ordering. arXiv preprint arXiv:1607.06952
2016 arXiv
-
[14]
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555
2014 arXiv
-
[15]
Alexis Conneau and Douwe Kiela. 2018. https://www.aclweb.org/anthology/L18-1269 S ent E val: An evaluation toolkit for universal sentence representations . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC -2018) , Miyazaki, Ja...
2018
-
[16]
Alexis Conneau, Douwe Kiela, Holger Schwenk, Lo \"i c Barrault, and Antoine Bordes. 2017. https://doi.org/10.18653/v1/D17-1070 Supervised learning of universal sentence representations from natural language inference data . In Proceedings of the 2017 Conference on Empirical Me...
2017 doi
-
[17]
Alexis Conneau, Germ \'a n Kruszewski, Guillaume Lample, Lo \" c Barrault, and Marco Baroni. 2018. https://www.aclweb.org/anthology/P18-1198 What you can cram into a single \ & ! \# * vector: Probing sentence embeddings for linguistic properties . In Proceedings of the 56th An...
2018
-
[18]
Baiyun Cui, Yingming Li, Ming Chen, and Zhongfei Zhang. 2018. http://aclweb.org/anthology/D18-1465 Deep attentive sentence ordering network . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4340--4349. Association for Computatio...
2018
-
[19]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[20]
Bill Dolan, Chris Quirk, and Chris Brockett. 2004. Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources. In Proceedings of the 20th International Conference on Computational Linguistics (COLING), pages 350--356, Geneva, Switzerland
2004
-
[21]
Micha Elsner and Eugene Charniak. 2008. https://www.aclweb.org/anthology/P08-1095 You talking to me? a corpus and algorithm for conversation disentanglement . In Proceedings of ACL-08: HLT, pages 834--842, Columbus, Ohio. Association for Computational Linguistics
2008
-
[22]
Micha Elsner and Eugene Charniak. 2010. https://doi.org/10.1162/coli_a_00003 Disentangling chat . Computational Linguistics, 36(3):389--409
2010 doi
-
[23]
Allyson Ettinger. 2019. What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models. arXiv preprint arXiv:1907.13528
2019 arXiv
-
[24]
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, and Katrin Erk. 2019. https://www.aclweb.org/anthology/P19-1062 Evaluating discourse in structured text representations . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 646--653, ...
2019
-
[25]
Zhe Gan, Yunchen Pu, Ricardo Henao, Chunyuan Li, Xiaodong He, and Lawrence Carin. 2017. https://doi.org/10.18653/v1/D17-1254 Learning generic sentence representations using convolutional neural networks . In Proceedings of the 2017 Conference on Empirical Methods in Natural La...
2017 doi
-
[26]
Ng, and Bita Nejat
Shima Gerani, Yashar Mehdad, Giuseppe Carenini, Raymond T. Ng, and Bita Nejat. 2014. https://doi.org/10.3115/v1/D14-1168 Abstractive summarization of product reviews using discourse structure . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Proc...
2014 doi
-
[27]
Grosz, Scott Weinstein, and Aravind K
Barbara J. Grosz, Scott Weinstein, and Aravind K. Joshi. 1995. http://aclweb.org/anthology/J95-2003 Centering: A framework for modeling the local coherence of discourse . Computational Linguistics, 21(2)
1995
-
[28]
Francisco Guzm \'a n, Shafiq Joty, Llu \' s M \`a rquez, and Preslav Nakov. 2014. https://doi.org/10.3115/v1/P14-1065 Using discourse structure improves machine translation evaluation . In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics ...
2014 doi
-
[29]
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016. https://doi.org/10.18653/v1/N16-1162 Learning distributed representations of sentences from unlabelled data . In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistic...
2016 doi
-
[30]
Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 168--177. ACM
2004
-
[31]
Yacine Jernite, Samuel R Bowman, and David Sontag. 2017. Discourse-based objectives for fast unsupervised sentence representation learning. arXiv preprint arXiv:1705.00557
2017 arXiv
-
[32]
Yangfeng Ji and Jacob Eisenstein. 2014. https://doi.org/10.3115/v1/P14-1002 Representation learning for text-level discourse parsing . In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13--24, Baltimore, M...
2014 doi
-
[33]
Yangfeng Ji and Jacob Eisenstein. 2015. http://aclweb.org/anthology/Q15-1024 One vector is not enough: Entity-augmented distributed semantics for discourse relations . Transactions of the Association for Computational Linguistics, 3:329--344
2015
-
[34]
Yangfeng Ji and Noah A. Smith. 2017. https://doi.org/10.18653/v1/P17-1092 Neural discourse structure for text categorization . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 996--1005, Vancouver, Canada...
2017 doi
-
[35]
Jyun-Yu Jiang, Francine Chen, Yan-Ying Chen, and Wei Wang. 2018. https://doi.org/10.18653/v1/N18-1164 Learning to disentangle interleaved conversational threads with a siamese hierarchical network and similarity ranking . In Proceedings of the 2018 Conference of the North Amer...
2018 doi
-
[36]
Daniel Jurafsky and James H. Martin. 2009. Speech and Language Processing (2Nd Edition). Prentice-Hall, Inc., Upper Saddle River, NJ, USA
2009
-
[37]
Nal Kalchbrenner and Phil Blunsom. 2013. http://aclweb.org/anthology/W13-3214 Recurrent convolutional neural networks for discourse compositionality . In Proceedings of the Workshop on Continuous Vector Space Models and their Compositionality, pages 119--126. Association for C...
2013
-
[38]
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. 2018. https://doi.org/10.18653/v1/N18-1149 A dataset of peer reviews (peerread): Collection, insights and nlp applications . In Proceedings of the 2018 Confere...
2018 doi
-
[39]
Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2015. http://dl.acm.org/citation.cfm?id=2969442.2969607 Skip-thought vectors . In Proceedings of the 28th International Conference on Neural Information Processin...
2015
-
[40]
Kummerfeld, Sai R
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, and Walter Lasecki. 2019. https://www.aclweb.org/anthology/P19-1374 A large-scale corpus for conversation disentangleme...
2019
-
[41]
Quoc Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In International conference on machine learning, pages 1188--1196
2014
-
[42]
Jiwei Li and Dan Jurafsky. 2017. https://doi.org/10.18653/v1/D17-1019 Neural net models of open-domain discourse coherence . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 198--209. Association for Computational Linguistics
2017 doi
-
[43]
Xin Li and Dan Roth. 2002. Learning question classifiers. In Proceedings of the 19th international conference on Computational linguistics-Volume 1, pages 1--7. Association for Computational Linguistics
2002
-
[44]
Xiang Lin, Shafiq Joty, Prathyusha Jwalapuram, and M Saiful Bari. 2019. https://www.aclweb.org/anthology/P19-1410 A unified linear-time framework for sentence-level discourse parsing . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, ...
2019
-
[45]
Ziheng Lin, Min-Yen Kan, and Hwee Tou Ng. 2009. http://aclweb.org/anthology/D09-1036 Recognizing implicit discourse relations in the penn discourse treebank . In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 343--351. Association...
2009
-
[46]
Cohen, and Mirella Lapata
Jiangming Liu, Shay B. Cohen, and Mirella Lapata. 2018. https://doi.org/10.18653/v1/P18-1040 Discourse representation structure parsing . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 429--439, Melbour...
2018 doi
-
[47]
Liu, Matt Gardner, Yonatan Belinkov, Matthew E
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 a . https://doi.org/10.18653/v1/N19-1112 Linguistic knowledge and transferability of contextual representations . In Proceedings of the 2019 Conference of the North A merican Chapter of t...
2019 doi
-
[48]
Yang Liu and Mirella Lapata. 2018. https://doi.org/10.1162/tacl_a_00005 Learning structured text representations . Transactions of the Association for Computational Linguistics, 6:63--75
2018 doi
-
[49]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 b . Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692
2019 arXiv
-
[50]
Lajanugen Logeswaran and Honglak Lee. 2018. https://openreview.net/forum?id=rJvJXZb0W An efficient framework for learning sentence representations . In International Conference on Learning Representations
2018
-
[51]
Lajanugen Logeswaran, Honglak Lee, and Dragomir Radev. 2016. http://arxiv.org/abs/1611.02654 Sentence ordering and coherence modeling using recurrent neural networks
2016 arXiv
-
[52]
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015. https://doi.org/10.18653/v1/W15-4640 The U buntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems . In Proceedings of the 16th Annual Meeting of the Special Interest Group o...
2015 doi
-
[53]
William C Mann and Sandra A Thompson. 1988. http://scholar.google.com/scholar.bib?q=info:BEw8CIWbucoJ:scholar.google.com/&output=citation&scisig=AAGBfm0AAAAAU3X_1Dq4ULnWfFzMeRsqGJcha1fReMSl&scisf=4&hl=en Rhetorical structure theory: Toward a functional theory of text organizat...
1988
-
[54]
Daniel Marcu. 2000. The Theory and Practice of Discourse Parsing and Summarization. MIT Press, Cambridge, MA, USA
2000
-
[55]
Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. 2014. SemEval -2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment. In Proceedin...
2014
-
[56]
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017. http://papers.nips.cc/paper/7209-learned-in-translation-contextualized-word-vectors.pdf Learned in translation: Contextualized word vectors . In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S....
2017
-
[57]
Shikib Mehri and Giuseppe Carenini. 2017. https://www.aclweb.org/anthology/I17-1062 Chat disentanglement: Identifying semantic reply relationships with random forests and recurrent neural networks . In Proceedings of the Eighth International Joint Conference on Natural Languag...
2017
-
[58]
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. http://www.aclweb.org/anthology/N16-1098 A corpus and cloze evaluation for deeper understanding of commonsense stories . In Proceedings of t...
2016
-
[59]
Karthik Narasimhan and Regina Barzilay. 2015. https://doi.org/10.3115/v1/P15-1121 Machine comprehension with discourse relations . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural L...
2015 doi
-
[60]
Allen Nie, Erin Bennett, and Noah Goodman. 2019. https://www.aclweb.org/anthology/P19-1442 D is S ent: Learning sentence representations from explicit discourse relations . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4497--...
2019
-
[61]
Hamid Palangi, Li Deng, Yelong Shen, Jianfeng Gao, Xiaodong He, Jianshu Chen, Xinying Song, and Rabab Ward. 2016. https://doi.org/10.1109/TASLP.2016.2520371 Deep sentence embedding using long short-term memory networks: Analysis and application to information retrieval . IEEE/...
2016
-
[62]
Boyuan Pan, Yazheng Yang, Zhou Zhao, Yueting Zhuang, Deng Cai, and Xiaofei He. 2018. https://doi.org/10.18653/v1/P18-1091 Discourse marker augmented network with reinforcement learning for natural language inference . In Proceedings of the 56th Annual Meeting of the Associatio...
2018 doi
-
[63]
Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of the 42nd annual meeting on Association for Computational Linguistics, page 271. Association for Computational Linguistics
2004
-
[64]
Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the 43rd annual meeting on association for computational linguistics, pages 115--124. Association for Computational Linguistics
2005
-
[65]
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 a . https://doi.org/10.18653/v1/N18-1202 Deep contextualized word representations . In Proceedings of the 2018 Conference of the North American Chapter of the Ass...
2018 doi
-
[66]
Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018 b . https://doi.org/10.18653/v1/D18-1179 Dissecting contextual word embeddings: Architecture and representation . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pa...
2018 doi
-
[67]
Karl Pichotta and Raymond J. Mooney. 2016. https://doi.org/10.18653/v1/P16-1027 Using sentence-level LSTM language models for script inference . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 279--289, ...
2016 doi
-
[68]
Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018. https://doi.org/10.18653/v1/D18-1007 Collecting diverse natural language inference problems for sentence representation evaluation . In Proceedings of...
2018 doi
-
[69]
Rashmi Prasad, Nikhil Dinesh, Alan Lee, Eleni Miltsakaki, Livio Robaldo, Aravind Joshi, and Bonnie Webber. 2008. The penn discourse treebank 2.0. In In Proceedings of LREC
2008
-
[70]
Xing Shi, Inkit Padhi, and Kevin Knight. 2016. https://doi.org/10.18653/v1/D16-1159 Does string-based neural MT learn source syntax? In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1526--1534, Austin, Texas. Association for Comp...
2016 doi
-
[71]
Damien Sileo, Tim Van De Cruys, Camille Pradel, and Philippe Muller. 2019. https://doi.org/10.18653/v1/N19-1351 Mining discourse markers for unsupervised sentence representation learning . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association ...
2019 doi
-
[72]
Manning, Andrew Ng, and Christopher Potts
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language ...
2013
-
[73]
Shuai Tang and Virginia R. de Sa. 2019. https://www.aclweb.org/anthology/P19-1397 Exploiting invertible decoders for unsupervised sentence representation learning . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4050--4060, Fl...
2019
-
[74]
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4631--4640
2016
-
[75]
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019. https://openreview.net/forum?id=SJzSgnRcKX What do you learn from context? probing for sentence structure in contex...
2019
-
[76]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537
2019 arXiv
-
[77]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 a . http://aclweb.org/anthology/W18-5446 Glue: A multi-task benchmark and analysis platform for natural language understanding . In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: An...
2018
-
[78]
Su Wang, Eric Holgate, Greg Durrett, and Katrin Erk. 2018 b . http://aclweb.org/anthology/D18-1175 Picking apart story salads . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1455--1465. Association for Computational Linguistics
2018
-
[79]
Yizhong Wang, Sujian Li, and Jingfeng Yang. 2018 c . https://doi.org/10.18653/v1/D18-1116 Toward fast and accurate neural discourse segmentation . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 962--967, Brussels, Belgium. Asso...
2018 doi
-
[80]
Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39(2):165--210
2005
-
[81]
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016. Towards universal paraphrastic sentence embeddings. In Proceedings of International Conference on Learning Representations
2016
-
[82]
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237
2019 arXiv
-
[83]
Zhi-Min Zhou, Yu Xu, Zheng-Yu Niu, Man Lan, Jian Su, and Chew Lim Tan. 2010. http://aclweb.org/anthology/C10-2172 Predicting discourse connectives for implicit discourse relation recognition . In Coling 2010: Posters, pages 1507--1514. Coling 2010 Organizing Committee
2010
-
[84]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[85]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.