REVIEW 1 cited by
Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corpora spans 2010 to 2020 and contains 28,825,820 to 29,549,672 words of multilingual text posted by 533 Sri Lankan Facebook pages, including politics, media, celebrities, and other categories; the smaller corpus amounts to 5,402,76 words of only Sinhala text extracted from the larger. Both corpora have markers for their date of creation, page of origin, and content type.
Forward citations
Cited by 1 Pith paper
-
Linguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study
A new dataset of 63,471 Sinhala YouTube comments on music videos and 964 derived stop-words is presented and compared against general Sinhala corpora.
Discussion (0). Continue with ORCID to comment.