Pith. sign in

REVIEW 1 cited by

Subword-Level Language Identification for Intra-Word Code-Switching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.01989 v1 pith:ZJ77AZOK submitted 2019-04-03 cs.CL cs.LG

classification cs.CLcs.LG
keywords languageidentificationwordscode-switchingdatasethoweverintra-wordmixed
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Language identification for code-switching (CS), the phenomenon of alternating between two or more languages in conversations, has traditionally been approached under the assumption of a single language per token. However, if at least one language is morphologically rich, a large number of words can be composed of morphemes from more than one language (intra-word CS). In this paper, we extend the language identification task to the subword-level, such that it includes splitting mixed words while tagging each part with a language ID. We further propose a model for this task, which is based on a segmental recurrent neural network. In experiments on a new Spanish--Wixarika dataset and on an adapted German--Turkish dataset, our proposed model performs slightly better than or roughly on par with our best baseline, respectively. Considering only mixed words, however, it strongly outperforms all baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bilingual Word Level Language Identification for Omotic Languages

    cs.CL 2025-09 conditional novelty 4.0 of 10

    On a new 144,000-word annotated dataset for Wolayta and Gofa, BERT-base-uncased embeddings with an LSTM classifier reach 0.72 F1, the best of seven compared approaches.

Pith tools