Pith. sign in

REVIEW 1 cited by

Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.07884 v1 pith:LT5BE4SL submitted 2020-07-15 cs.CL cs.SI

classification cs.CLcs.SI
keywords corporalanguagesinhalafacebooklankanlargerstopwordstext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corpora spans 2010 to 2020 and contains 28,825,820 to 29,549,672 words of multilingual text posted by 533 Sri Lankan Facebook pages, including politics, media, celebrities, and other categories; the smaller corpus amounts to 5,402,76 words of only Sinhala text extracted from the larger. Both corpora have markers for their date of creation, page of origin, and content type.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new dataset of 63,471 Sinhala YouTube comments on music videos and 964 derived stop-words is presented and compared against general Sinhala corpora.

Pith tools