Pith. sign in

REVIEW 2 cited by

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.00204 v4 pith:YLH57M47 submitted 2021-01-01 cs.CL

classification cs.CL
keywords banglalanguagebanglabertunderstandingbenchmarkdatasetsintroducelow-resource
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla pretraining data (dubbed `Bangla2B+') by crawling 110 popular Bangla sites. We introduce two downstream task datasets on natural language inference and question answering and benchmark on four diverse NLU tasks covering text classification, sequence labeling, and span prediction. In the process, we bring them under the first-ever Bangla Language Understanding Benchmark (BLUB). BanglaBERT achieves state-of-the-art results outperforming multilingual and monolingual models. We are making the models, datasets, and a leaderboard publicly available at https://github.com/csebuetnlp/banglabert to advance Bangla NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis

    cs.CL 2025-07 reject novelty 5.0 of 10

    The authors release eight Bengali benchmarks translated from English and report that models with more fragmented Bengali tokenization tend to score lower.

  2. Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Fine-tuning mT5, BanglaT5, mBART50 and a basic Transformer on a newly translated Bengali math word problem dataset yields up to 97.3% solution accuracy on elementary arithmetic problems.

Pith tools