REVIEW 4 major objections 6 minor 69 references
BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Concatenating a Bengali religious news article's category, aspect, and sentiment labels with its text yields better generated headlines than using the text alone, across four transformer models.
desk verdict Useful corpus, plausible but under-tested method: MultiGen's gains all come with gold labels at inference and no significance tests, so treat the headline numbers as upper bounds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the fused input sequence $I = [A_1, \ldots, A_n, \mathrm{[SEP]}, C_1, \ldots, C_k, \mathrm{[SEP]}, P_1, \ldots, P_m, \mathrm{[SEP]}, S_1, \ldots, S_p]$ in which $A$ is the tokenized article, $C$ the category tokens, $P$ the aspect tokens, and $S$ the sentiment tokens. This single text-to-text input replaces the article-only encoder input, so a T5-style encoder-decoder conditions the headline on the three categorical signals without any architectural change. The paper's experiments hold the model fixed and vary only this input composition, isolating the effect of the fused context.
What would settle it
Run the best BanglaT5 configuration twice: once with the corpus's gold labels and once with labels predicted by an independent classifier, then compare BLEU and ROUGE-L on the same test split; if the predicted-label run falls back to the content-only baseline, the fusion gain depends on oracle access. A within-corpus check is to shuffle the category, aspect, and sentiment tokens across articles before training; if scores do not drop, the gain comes from the extra tokens, not their semantics.
Extended reading notes
Core claim
The central claim is that giving a headline generator the article's category, aspect, and sentiment explicitly, rather than letting it infer them from the text, produces better headlines for Bengali religious news. The paper evaluates this by comparing a content-only input string with an input that appends the three labels, keeping the model architecture and hyperparameters fixed for each of the four pre-trained transformer models. Across nearly all model-metric combinations the fused input wins; the single reported exception is mBART's ROUGE-2, where the fused input drops from 7.90 to 7.78. The largest gains come from BanglaT5, the model pre-trained on Bengali; the paper reads this as evidence that contextual side information compensates for the limited guidance available to content-only systems in a low-resource setting.
Load-bearing premise
The reported improvement assumes the category, aspect, and sentiment labels are available and correct at generation time; the experiments feed the corpus's own manual labels, so nothing in the paper shows how the method behaves when those features must be predicted automatically.
Editorial extensions
If this is right
- The BeliN corpus (2,520 labeled articles) becomes a public benchmark for Bengali religious headline generation, summarization, and classification.
- MultiGen-style feature fusion is reported to improve BLEU and ROUGE-L over the content-only baseline for all four tested transformer backbones, suggesting the benefit is not specific to one architecture.
- The best configuration, BanglaT5 with fused context, sets the reported state-of-the-art numbers on BeliN (BLEU 18.61, ROUGE-L 24.19).
- Because the method only changes the input string, it can be dropped into any text-to-text model without architectural modification.
- Category, aspect, and sentiment labels, often discarded in summarization, are shown to carry signal worth exploiting in low-resource headline generation.
Reading between the lines
- If the labels were predicted by an upstream classifier instead of taken from the corpus, the reported margins would likely shrink; the paper only evaluates with oracle labels, so deployment would require an annotation or prediction pipeline.
- A plausible ablation is to shuffle the labels or replace them with dummy constants; if gains persist, the improvement may come from the extra input length rather than from the semantic content of the labels.
- The corpus is strongly skewed (about 79% Islam-related), so per-category evaluation would clarify whether the fusion helps all religious groups or mainly the majority class.
- The same concatenation recipe could be tested on other low-resource languages that have side information (e.g., document genre or user metadata) to see whether contextual fusion generalizes beyond Bengali religious news.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BeliN, a new Bengali religious news corpus of 2,520 article-headline pairs annotated with category, aspect, and sentiment, and proposes MultiGen, a headline generation approach that concatenates these labels with the article text as input to transformer-based sequence-to-sequence models (BanglaT5, mBART, mT5, mT0). Experiments compare MultiGen against a content-only baseline on a 500-article test split, reporting that the proposed approach improves most metrics across all four models, with BanglaT5 achieving BLEU 18.61 and ROUGE-L 24.19 versus 16.08 and 23.08 for the baseline. The authors argue that the additional contextual features improve headline quality and that the corpus addresses a gap in low-resource Bengali NLP.
Significance. If the reported gains are robust, the BeliN corpus is a useful resource for an underexplored domain, and the MultiGen input-fusion recipe is a simple, reproducible baseline for future work. The paper ships public data and code, and the experimental setup covers four diverse pretrained models, which strengthens the empirical description. However, the central claim depends on test-time availability of gold category, aspect, and sentiment labels, and the evaluation lacks significance testing; these issues currently leave the magnitude and even the existence of a real advantage over the baseline uncertain.
major comments (4)
- [Sections 4.2 and 5.5] The MultiGen input in Section 4.2 concatenates article text with category, aspect, and sentiment labels, and Table 10 reports gains using gold labels from BeliN as if they were available at inference time. A deployed system would need to predict these labels, and the labels are derived from the same article that gives rise to the reference headline, so the model can exploit a label-headline correlation that is not present in the content-only baseline. No experiment replaces gold labels with predicted labels, and no ablation with shuffled or random labels is reported. I ask for a predicted-label evaluation or a label-permutation control; without one, the reported BLEU and ROUGE improvements may reflect oracle leakage rather than a robust property of MultiGen.
- [Section 5.5, Table 10] The evaluation uses only 500 test examples and reports no confidence intervals, error bars, or paired significance tests. Several differences are small in absolute terms (for example, BanglaT5 BERTScore improves from 73.57 to 75.12, and mBART ROUGE-2 actually decreases from 7.90 to 7.78), and the mBART ROUGE-2 result moves opposite to the main claim. The consistent direction across most metrics and models is encouraging, but without a paired test such as bootstrap resampling or Wilcoxon signed-rank, the central claim that MultiGen outperforms the baseline is not statistically established.
- [Section 5.5 and Abstract] The paper refers to the improvements as 'SOTA' or state-of-the-art gains, but the only comparison is against the paper's own content-only baseline on BeliN. No prior headline-generation system is evaluated on BeliN, and no comparison is made to existing Bengali headline-generation systems such as Shironaam on a shared dataset. The label should be changed to 'relative improvement over the content-only baseline' in Table 10 and the Abstract, or the authors should provide an actual comparison against prior published methods.
- [Section 3.3, Table 4] The corpus is highly imbalanced: Islam has 2,001 samples while Christianity has 28 and Buddhism 29; the Report aspect accounts for 1,204 of 2,520 samples; Positive sentiment accounts for 1,717 of 2,520. These imbalances are not discussed in the experimental analysis. If MultiGen's gains are partly driven by the model learning label-specific surface patterns from the majority classes, then the reported aggregate improvement may not transfer to minority categories or aspects. I ask for per-category or class-stratified results, or at least a discussion of how the imbalance interacts with the label-conditioning mechanism.
minor comments (6)
- [Section 4.2 and 4.3] The input sequence I is defined with [SEP] tokens in Section 4.2, but the equivalent equation in Section 4.3 omits the [SEP] separators, making the two definitions inconsistent.
- [Table 10] The column heading 'Delta SOTA' is misleading because the comparison is against the paper's own baseline, not against published state-of-the-art results. Rename it to indicate relative change over the baseline.
- [Section 1 and Section 6.3] The limitations section mentions dataset scarcity and hardware constraints but does not mention the oracle-label limitation or the absence of significance testing; these should be acknowledged explicitly.
- [Table 2] Several source names and URLs contain typographical errors, including 'Alokito Bnagladesh' and 'Daily V orer Pata'; these should be corrected.
- [Section 5.2] The BLEU formula in Eq. (1) uses N=1 implicitly but the text describes multi-gram precision; clarify that N is the maximum n-gram order used in the computation.
- [Section 3.2] The paper states that labels were assigned manually but provides no inter-annotator agreement statistics or detailed annotation guidelines; at least a brief description of the annotation protocol would help readers assess label quality.
Circularity Check
No circularity: MultiGen's gain is an empirical result, not a derivation from its inputs.
full rationale
This paper contains no analytic derivation chain whose output could be equivalent to its inputs. The central claim is empirical: fine-tuned transformer models conditioned on I=[A;C;P;S] score higher BLEU/ROUGE than models conditioned on A alone (Table 10). Category, aspect, and sentiment are manually annotated properties of the source article, not functions of the reference headline, and no parameter is fitted to the test labels and then reported as a prediction. The comparisons are same-corpus baseline-vs-proposed with standard metrics; the extra features are additional conditioning variables, so a gain is a contingent empirical finding rather than a tautology. The only self-citations (refs. [24], [57]) are peripheral: [57] supports the general effectiveness of contextual fusion in sentiment analysis and [24] is mentioned as a future LLM direction; neither is load-bearing for the headline-generation result. The oracle-label evaluation at inference (gold C/P/S instead of predicted labels) is a legitimate deployment and information-availability weakness that could inflate the reported gain, but it is not a definitional equivalence or fitted-input circularity. No circular step can be exhibited, so the score is 0.
Assumptions & free parameters
free parameters (8)
- learning_rate_BanglaT5 =
1e-4
- learning_rate_mBART =
1e-3
- learning_rate_mT5 =
1e-4
- learning_rate_mT0 =
2e-5
- epochs =
5
- batch_size =
8
- input_token_length =
512
- target_token_length =
64
assumptions (4)
- domain assumption Manual labels (category, aspect, sentiment) are accurate and consistent.
- domain assumption Auxiliary features are available at inference time.
- domain assumption Evaluation metrics are valid for Bengali headline quality.
- domain assumption Pre-trained models are suitable backbones.
Cite this review
Pith. "Pith review of BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion." pith.science (2026). https://pith.science/paper/F3QWKO23
@misc{pith2026250101069,
author = {Pith},
title = {Pith review of: BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/F3QWKO23}},
note = {Machine review of arXiv:2501.01069}
}
read the original abstract
Automatic text summarization, particularly headline generation, remains a critical yet underexplored area for Bengali religious news. Existing approaches to headline generation typically rely solely on the article content, overlooking crucial contextual features such as sentiment, category, and aspect. This limitation significantly hinders their effectiveness and overall performance. This study addresses this limitation by introducing a novel corpus, BeliN (Bengali Religious News) - comprising religious news articles from prominent Bangladeshi online newspapers, and MultiGen - a contextual multi-input feature fusion headline generation approach. Leveraging transformer-based pre-trained language models such as BanglaT5, mBART, mT5, and mT0, MultiGen integrates additional contextual features - including category, aspect, and sentiment - with the news content. This fusion enables the model to capture critical contextual information often overlooked by traditional methods. Experimental results demonstrate the superiority of MultiGen over the baseline approach that uses only news content, achieving a BLEU score of 18.61 and ROUGE-L score of 24.19, compared to baseline approach scores of 16.08 and 23.08, respectively. These findings underscore the importance of incorporating contextual features in headline generation for low-resource languages. By bridging linguistic and cultural gaps, this research advances natural language processing for Bengali and other underrepresented languages. To promote reproducibility and further exploration, the dataset and implementation code are publicly accessible at https://github.com/akabircs/BeliN.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
P. Cai, K. Song, S. Cho, H. Wang, X. Wang, H. Yu, F. Liu, D. Yu, Generating user-engaging news headlines, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 2023, pp. 3265–3280. doi:10.18653/v1/2023.acl-long.183
-
[2]
G. De Francisci Morales, A. Gionis, C. Lucchese, From chatter to headlines: harnessing the real-time web for personalized news recommendation, in: Proceedings of the fifth ACM international conference on Web search and data mining, 2012, pp. 153–162. doi:10.1145/2124295.2124315
arXiv 2012
-
[3]
H. Y . Koh, J. Ju, M. Liu, S. Pan, An empirical survey on long document summarization: Datasets, models, and metrics, ACM computing surveys 55 (2022) 1–35. doi:10.1145/3545176
doi:10.1145/3545176 2022
-
[4]
A. Rao, S. Aithal, S. Singh, Single-document abstractive text summarization: A systematic literature review, ACM Comput. Surv. 57 (2024). doi:10.1145/3700639
-
[5]
S. Banerjee, S. Mukherjee, S. Bandyopadhyay, P. Pakray, An extract-then-abstract based method to generate disaster-news headlines using a dnn extractor followed by a transformer abstractor, Information Processing & Management 60 (2023) 103291. doi: 10.1016/j.ipm.2023. 103291
-
[6]
N. Giarelis, C. Mastrokostas, N. Karacapilidis, Abstractive vs. extractive summarization: An experimental review, Applied Sciences 13 (2023). doi:10.3390/app13137620
-
[7]
A. Alomari, N. Idris, A. Q. M. Sabri, I. Alsmadi, Deep reinforcement and transfer learning for abstractive text summarization: A review, Computer Speech & Language 71 (2022) 101276. doi:10.1016/j.csl.2021.101276
-
[8]
W. S. El-Kassas, C. R. Salama, A. A. Rafea, H. K. Mohamed, Automatic text summarization: A comprehensive survey, Expert systems with applications 165 (2021) 113679. doi:10.1016/j.eswa.2020.113679
arXiv 2021
Show all 69 references
-
[9]
Ahuir, J.-A
V . Ahuir, J.-A. Gonzalez, L.-F. Hurtado, E. Segarra, Abstractive summarizers become emotional on news summarization, Applied Sciences 14 (2024). doi:10.3390/app14020713
2024 doi
-
[10]
Shen, Y .-K
Ayana, S.-Q. Shen, Y .-K. Lin, C.-C. Tu, Y . Zhao, Z.-Y . Liu, M.-S. Sun, Recent advances on neural headline generation, Journal of Computer Science and Technology 32 (2017) 768–784. doi:10.1007/s11390-017-1758-3
2017 doi
-
[11]
Hagar, N
N. Hagar, N. Diakopoulos, Optimizing content with a /b headline testing: Changing newsroom practices, Media and Communication 7 (2019) 117–127
2019
-
[12]
Banerjee, O
A. Banerjee, O. Urminsky, The language that drives engagement: A systematic large-scale analysis of headline experiments, Marketing Science (2024)
2024
-
[13]
A. U. Akash, M. T. Nayeem, F. T. Shohan, T. Islam, Shironaam: Bengali news headline generation using auxiliary information, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Association for Computational Linguistic...
2023
-
[14]
A. M. Saad, U. N. Mahi, M. S. Salim, S. I. Hossain, Bangla news article dataset, Data in Brief 57 (2024) 110874. URL: 10.1016/j.dib. 2024.110874. xxiv
2024
-
[15]
D. M. Eberhard, G. F. Simons, C. D. Fennig, Ethnologue: Languages of the world. twenty-seventh edition., https://www.ethnologue. com/, 2024. [last accessed 30 December 2024]
2024
-
[16]
A. Y . Shaibani, A. M. Elnagar, A survey of text summarization and headline generation methods in arabic: A survey of text summarization and headline generation methods in arabic, in: Proceedings of the 2024 9th International Conference on Machine Learning Technologies, ICMLT ...
2024
-
[17]
A. M. A. Zeyad, A. Biradar, Advancements in the e fficacy of flan-t5 for abstractive text summarization: A multi-dataset evaluation using rouge and bertscore, 2024 International Conference on Advancements in Power, Communication and Intelligent Systems (APCI) (2024) 1–5. URL: ...
2024
-
[18]
A. K. Yadav, Ranvijay, R. S. Yadav, A. K. Maurya, State-of-the-art approach to extractive text summarization: a comprehensive review, Multimedia Tools and Applications 82 (2023) 29135–29197. doi:10.1007/s11042-023-14613-9
2023 doi
-
[19]
Bharathi Mohan, R
G. Bharathi Mohan, R. Prasanna Kumar, S. Parathasarathy, S. Aravind, K. B. Hanish, G. Pavithria, Text Summarization for Big Data Analyt- ics: A Comprehensive Review of GPT 2 and BERT Approaches, Springer Nature Switzerland, Cham, 2023, pp. 247–264. doi:10.1007/978- 3-031-33808-3_14
2023 doi
-
[20]
D. O. Cajueiro, A. G. Nery, I. Tavares, M. K. D. Melo, S. A. dos Reis, L. Weigang, V . R. R. Celestino, A comprehensive review of automatic text summarization techniques: method, data, evaluation and coding, 2023. arXiv:2301.03403
2023 arXiv
-
[21]
T. Liu, H. Li, J. Zhu, J. Zhang, C. Zong, Review headline generation with user embedding, in: M. Sun, T. Liu, X. Wang, Z. Liu, Y . Liu (Eds.), Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data, Springer International Publis...
2018 doi
-
[22]
Salehin, A
M. Salehin, A. Rafat, F. Khan, S. Abujar, Generating bengali news headlines: An attentive approach with sequence-to-sequence networks, in: Proceedings of the 8th International Conference System Modeling and Advancement in Research Trends (SMART), 2019, pp. 256–261. doi:10.1109...
2019
-
[23]
S. M. A. I. Hayat, A. Das, M. Hoque, Abstractive bengali text summarization using transformer-based learning, in: 6th International Conference on Electrical Information and Communication Technology (EICT), 2023, pp. 1–6. doi:10.1109/EICT61409.2023.10427906
2023 arXiv
-
[24]
Kabir, M
M. Kabir, M. S. Islam, M. T. R. Laskar, M. T. Nayeem, M. S. Bari, E. Hoque, BenLLM-eval: A comprehensive evaluation into the potentials and pitfalls of large language models on Bengali NLP, in: N. Calzolari, M.-Y . Kan, V . Hoste, A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings...
2024
-
[25]
Grusky, M
M. Grusky, M. Naaman, Y . Artzi, Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies, in: M. Walker, H. Ji, A. Stent (Eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...
2018 doi
-
[26]
Nallapati, B
R. Nallapati, B. Zhou, C. N. dos santos, C. Gulcehre, B. Xiang, Abstractive text summarization using sequence-to-sequence rnns and beyond, in: Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, 2016, pp. 280–290. doi:10.18653/v1/K16- 1028
2016 doi
-
[27]
R. D. Lins, H. Oliveira, L. Cabral, J. Batista, B. Tenorio, R. Ferreira, R. Lima, G. de França Pereira e Silva, S. J. Simske, The cnn-corpus: A large textual corpus for single-document extractive summarization, in: Proceedings of the ACM Symposium on Document Engineering 2019,...
2019
-
[28]
Jiang, M
X. Jiang, M. Dreyer, CCSum: A large-scale and high-quality dataset for abstractive news summarization, in: K. Duh, H. Gomez, S. Bethard (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...
2024 doi
-
[29]
Narayan, S
S. Narayan, S. B. Cohen, M. Lapata, Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization, in: E. Rilo ff, D. Chiang, J. Hockenmaier, J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in xxv Nat...
2018 doi
-
[30]
X. Gu, Y . Mao, J. Han, J. Liu, Y . Wu, C. Yu, D. Finnie, H. Yu, J. Zhai, N. Zukoski, Generating representative headlines for news stories, in: Proceedings of The Web Conference 2020, WWW ’20, Association for Computing Machinery, New York, NY , USA, 2020, p. 1773–1784. doi:10....
2020
-
[31]
X. Ao, X. Wang, L. Luo, Y . Qiao, Q. He, X. Xie, PENS: A dataset and generic framework for personalized news headline generation, in: C. Zong, F. Xia, W. Li, R. Navigli (Eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th...
2021 doi
-
[32]
X. Ao, L. Luo, X. Wang, Z. Yang, J.-H. Chen, Y . Qiao, Q. He, X. Xie, Put your voice on stage: Personalized headline generation for news articles, ACM Transactions on Knowledge Discovery from Data 18 (2023). doi:10.1145/3629168
2023 doi
-
[33]
D. Jin, Z. Jin, J. T. Zhou, L. Orii, P. Szolovits, Hooks in the headline: Learning to generate headlines with controlled styles, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics,...
2020 doi
-
[34]
Sandhaus, The New York Times Annotated Corpus, 2008
E. Sandhaus, The New York Times Annotated Corpus, 2008. URL: https://hdl.handle.net/11272.1/AB2/GZC6PL. doi:11272.1/ AB2/GZC6PL
2008
-
[35]
Takase, J
S. Takase, J. Suzuki, N. Okazaki, T. Hirao, M. Nagata, Neural headline generation on Abstract Meaning Representation, in: J. Su, K. Duh, X. Carreras (Eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computa- tional ...
2016 doi
-
[36]
Accessed: 30 December 2024
National Institute of Standard and Technology, Document understanding conferences, https://www-nlpir.nist.gov/projects/duc/ data.html, 2014. Accessed: 30 December 2024
2014
-
[37]
Gra ff, C
D. Gra ff, C. Cieri, English gigaword, https://catalog.ldc.upenn.edu/LDC2003T05, 2003. doi:10.35111/0z6y-q265
2003 doi
-
[38]
Napoles, M
C. Napoles, M. Gormley, B. Van Durme, Annotated gigaword, in: Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-Scale Knowledge Extraction, AKBC-WEKEX ’12, Association for Computational Linguistics, 2012, p. 95–100. URL: https://aclanthology.or...
2012
-
[39]
R. K. Singh, S. Khetarpaul, R. Gorantla, S. G. Allada, Sheg: summarization and headline generation of news articles using deep learning, Neural Computing and Applications 33 (2021) 3251–3265. doi:10.1007/s00521-020-05188-9
2021 doi
-
[40]
M. Sen, B. Yanikoglu, Document classification of suder turkish news corpora, in: Proceedings of the 2018 26th Signal Processing and Communications Applications Conference (SIU), 2018, pp. 1–4. doi:10.1109/SIU.2018.8404790
2018
-
[41]
Ogunremi, S
T. Ogunremi, S. sessi Akojenu, A. Soronnadi, O. Adekanmbi, D. I. Adelani, AfriHG: News headline generation for african languages, in: 5th Workshop on African Natural Language Processing, 2024, p. 4. URL: https://openreview.net/forum?id=fw7g7pNUDl
2024
-
[42]
Bukhtiyarov, I
A. Bukhtiyarov, I. Gusev, Advances of Transformer-Based Models for News Headline Generation, Springer, 2020, pp. 54–61. doi:10.1007/ 978-3-030-59082-6_4
2020
-
[43]
Gavrilov, P
D. Gavrilov, P. Kalaidin, V . Malykh, Self-attentive model for headline generation, 2019. doi:10.1007/978-3-030-15719-7_11
2019 doi
-
[44]
Yutkin, Lenta, https://github.com/yutkin/Lenta.Ru-News-Dataset , 2019
D. Yutkin, Lenta, https://github.com/yutkin/Lenta.Ru-News-Dataset , 2019. Accessed: 30 December 2024
2019
-
[45]
B. Hu, Q. Chen, F. Zhu, LCSTS: A large scale Chinese short text summarization dataset, in: L. Màrquez, C. Callison-Burch, J. Su (Eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Lisbon, Po...
2015 doi
-
[46]
Madasu, G
L. Madasu, G. Kanumolu, N. Surange, M. Shrivastava, Mukhyansh: A headline generation dataset for Indic languages, in: C.-R. Huang, Y . Harada, J.-B. Kim, S. Chen, Y .-Y . Hsu, E. Chersoni, P. A, W. H. Zeng, B. Peng, Y . Li, J. Li (Eds.), Proceedings of the 37th Pacific Asia Co...
2023
-
[47]
Aralikatte, Z
R. Aralikatte, Z. Cheng, S. Doddapaneni, J. C. K. Cheung, Varta: A large-scale headline-generation dataset for Indic languages, in: A. Rogers, xxvi J. Boyd-Graber, N. Okazaki (Eds.), Findings of the Association for Computational Linguistics: ACL 2023, Association for Computati...
2023 doi
-
[48]
P. Li, J. Yu, J. Chen, B. Guo, Hg-news: News headline generation based on a generative pre-training model, IEEE Access 9 (2021) 110039–110046. doi:10.1109/ACCESS.2021.3102741
2021
-
[49]
Hasan, A
T. Hasan, A. Bhattacharjee, M. S. Islam, K. Mubasshir, Y .-F. Li, Y .-B. Kang, M. S. Rahman, R. Shahriyar, XL-Sum: Large-scale multilingual abstractive summarization for 44 languages, in: The Joint Conference of the 59th Annual Meeting of the Association for Computational Ling...
2021
-
[50]
Ahmad, F
I. Ahmad, F. Alqurashi, R. Mehmood, Potrika: Raw and balanced newspaper datasets in the bangla language with eight topics and five attributes, 2022. arXiv:2210.09389
2022 arXiv
-
[51]
Karaca, O
A. Karaca, O. Aydın, Generating headlines for turkish news texts with transformer architecture based deep learning method, Gazi Üniversitesi Mühendislik-Mimarlık Fakültesi Dergisi 39 (2023) 485–495. doi:10.17341/gazimmfd.963240
2023 doi
-
[52]
Theledi, V
K. Theledi, V . M. Pule, President’s speech and terminology used during the covid-19 pandemic: The interpretation of linguistic meaning in context and situational context, in: Public Health Communication Challenges to Minority and Indigenous Communities, IGI Global, 2024, pp. 92–107
2024
-
[53]
D. Liu, Y . Gong, Y . Yan, J. Fu, B. Shao, D. Jiang, J. Lv, N. Duan, Diverse, controllable, and keyphrase-aware: A corpus and method for news multi-headline generation, in: B. Webber, T. Cohn, Y . He, Y . Liu (Eds.), Proceedings of the 2020 Conference on Empirical Methods in N...
2020 doi
-
[54]
Kiefer, Case: Explaining text classifications by fusion of local surrogate explanation models with contextual and semantic knowledge, Information Fusion 77 (2022) 184–195
S. Kiefer, Case: Explaining text classifications by fusion of local surrogate explanation models with contextual and semantic knowledge, Information Fusion 77 (2022) 184–195. doi:10.1016/j.inffus.2021.07.014
2022 doi
-
[55]
Chen, Ci-snf: Exploiting contextual information to improve snf based information retrieval, Information Fusion 52 (2019) 175–186
N. Chen, Ci-snf: Exploiting contextual information to improve snf based information retrieval, Information Fusion 52 (2019) 175–186. doi:j.inffus.2018.08.004
2019
-
[56]
L. Zhu, Z. Zhu, C. Zhang, Y . Xu, X. Kong, Multimodal sentiment analysis based on fusion methods: A survey, Information Fusion 95 (2023) 306–325. doi:j.inffus.2023.02.028
2023
-
[57]
A. Aziz, N. K. Chowdhury, M. A. Kabir, A. N. Chy, M. J. Siddique, Mmtf-des: A fusion of multimodal transformer models for desire, emotion, and sentiment analysis of social media data, 2023. arXiv:2310.14143
2023 arXiv
-
[58]
Hasan, A
T. Hasan, A. Bhattacharjee, K. Samin, M. Hasan, M. Basak, M. S. Rahman, R. Shahriyar, Not low-resource anymore: Aligner ensembling, batch filtering, and new datasets for Bengali-English machine translation, in: Proceedings of the 2020 Conference on Empirical Methods in Natural...
2020 doi
-
[59]
Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, pp
C.-Y . Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, pp. 74–81. URL: https://aclanthology.org/W04-1013
2004
-
[60]
Papineni, S
K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a method for automatic evaluation of machine translation, in: P. Isabelle, E. Charniak, D. Lin (Eds.), Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Lingu...
2002
-
[61]
Banerjee, A
S. Banerjee, A. Lavie, METEOR: An automatic metric for MT evaluation with improved correlation with human judgments, in: J. Goldstein, A. Lavie, C.-Y . Lin, C. V oss (Eds.), Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation ...
2005
-
[62]
Zhang, V
T. Zhang, V . Kishore, F. Wu, K. Q. Weinberger, Y . Artzi, BERTScore: Evaluating text generation with bert, in: International Conference on Learning Representations, 2020, p. 43. URL: https://openreview.net/forum?id=SkeHuCVFDr
2020
-
[63]
Ra ffel, N
C. Ra ffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, P. J. Liu, Exploring the limits of transfer learning with xxvii a unified text-to-text transformer, Journal of Machine Learning Research 21 (2020) 1–67. URL: http://jmlr.org/papers/v21/20- 074.html
2020
-
[64]
Lewis, Y
M. Lewis, Y . Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V . Stoyanov, L. Zettlemoyer, BART: Denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Pro...
2020 doi
-
[65]
Bhattacharjee, T
A. Bhattacharjee, T. Hasan, W. U. Ahmad, R. Shahriyar, BanglaNLG: Benchmarks and resources for evaluating low-resource natural language generation in bangla, 2022. arXiv:2205.11081
2022 arXiv
-
[66]
Muennigho ff, T
N. Muennigho ff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. Le Scao, M. S. Bari, S. Shen, Z. X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Ra ff, C. Raffel, Crosslingual generalization through multitask finetu...
2023
-
[67]
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, C. Ra ffel, mT5: A massively multilingual pre-trained text- to-text transformer, in: K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, Y...
2021
-
[68]
Y . Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V . Chaudhary, J. Gu, A. Fan, Multilingual translation from denoising pre-training, in: The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference ...
2021
-
[69]
Rahman, A
A. Rahman, A. Mamun, The rise of clickbait headlines: A study on media platforms from bangladesh, Athens Journal of Mass Media and Communications 10 (2024) 109–130. doi:10.30958/ajmmc.10-2-3 . xxviii
2024 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.