Pith. sign in

REVIEW 3 cited by

What talking you?: Translating Code-Mixed Messaging Texts to English

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05253 v1 pith:5ZYGUSMB submitted 2024-11-08 cs.CL

classification cs.CL
keywords code-mixedenglishlanguagessinglishtranslationanalysistextsasian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Translation of code-mixed texts to formal English allow a wider audience to understand these code-mixed languages, and facilitate downstream analysis applications such as sentiment analysis. In this work, we look at translating Singlish, which is colloquial Singaporean English, to formal standard English. Singlish is formed through the code-mixing of multiple Asian languages and dialects. We analysed the presence of other Asian languages and variants which can facilitate translation. Our dataset is short message texts, written as informal communication between Singlish speakers. We use a multi-step prompting scheme on five Large Language Models (LLMs) for language detection and translation. Our analysis show that LLMs do not perform well in this task, and we describe the challenges involved in translation of code-mixed languages. We also release our dataset in this link https://github.com/luoqichan/singlish.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Code-switching hurts LLM comprehension when non-English tokens enter English text, but inserting English into other languages often improves accuracy; fine-tuning mitigates losses more reliably than prompting.

  2. Stylistic Evolution and LLM Neutrality in Singlish Language

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Singlish changed cumulatively over a decade, and LLM-generated Singlish remains tied to particular time periods: realistic outputs carry temporal bias, while neutral outputs lose authenticity.

  3. Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A two-stage pipeline combining human-curated Singlish examples with embedding-based LLM ranking selects GPT-4o mini for toxicity-preserving translation, reaching gold-level human scores for Chinese and Malay but not Tamil.

Pith tools