REVIEW 2 cited by
BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla pretraining data (dubbed `Bangla2B+') by crawling 110 popular Bangla sites. We introduce two downstream task datasets on natural language inference and question answering and benchmark on four diverse NLU tasks covering text classification, sequence labeling, and span prediction. In the process, we bring them under the first-ever Bangla Language Understanding Benchmark (BLUB). BanglaBERT achieves state-of-the-art results outperforming multilingual and monolingual models. We are making the models, datasets, and a leaderboard publicly available at https://github.com/csebuetnlp/banglabert to advance Bangla NLP.
Forward citations
Cited by 2 Pith papers
-
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
The authors release eight Bengali benchmarks translated from English and report that models with more fragmented Bengali tokenization tend to score lower.
-
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
Fine-tuning mT5, BanglaT5, mBART50 and a basic Transformer on a newly translated Bengali math word problem dataset yields up to 97.3% solution accuracy on elementary arithmetic problems.
Discussion (0). Continue with ORCID to comment.