Pith. sign in

REVIEW 2 cited by

Transformer-based Model for Word Level Language Identification in Code-mixed Kannada-English Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.14459 v1 pith:EWNLD2NF submitted 2022-11-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords code-mixedlanguageidentificationmodelf1-scoremediasocialtexts
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Using code-mixed data in natural language processing (NLP) research currently gets a lot of attention. Language identification of social media code-mixed text has been an interesting problem of study in recent years due to the advancement and influences of social media in communication. This paper presents the Instituto Polit\'ecnico Nacional, Centro de Investigaci\'on en Computaci\'on (CIC) team's system description paper for the CoLI-Kanglish shared task at ICON2022. In this paper, we propose the use of a Transformer based model for word-level language identification in code-mixed Kannada English texts. The proposed model on the CoLI-Kenglish dataset achieves a weighted F1-score of 0.84 and a macro F1-score of 0.61.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bilingual Word Level Language Identification for Omotic Languages

    cs.CL 2025-09 conditional novelty 4.0 of 10

    On a new 144,000-word annotated dataset for Wolayta and Gofa, BERT-base-uncased embeddings with an LSTM classifier reach 0.72 F1, the best of seven compared approaches.

  2. Multiclass Sentiment Analysis for Identifying Political Viewpoints

    cs.CL 2026-08 conditional novelty 3.0 of 10

    On the DravidianLangTech 2025 Tamil political sentiment test set, XGBoost scored 0.2835 macro F1 and BERT scored 0.2806, showing the task is far from solved.

Pith tools