REVIEW 3 cited by
From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generating code-switched text is a problem of growing interest, especially given the scarcity of corpora containing large volumes of real code-switched text. In this work, we adapt a state-of-the-art neural machine translation model to generate Hindi-English code-switched sentences starting from monolingual Hindi sentences. We outline a carefully designed curriculum of pretraining steps, including the use of synthetic code-switched text, that enable the model to generate high-quality code-switched text. Using text generated from our model as data augmentation, we show significant reductions in perplexity on a language modeling task, compared to using text from other generative models of CS text. We also show improvements using our text for a downstream code-switched natural language inference task. Our generated text is further subjected to a rigorous evaluation using a human evaluation study and a range of objective metrics, where we show performance comparable (and sometimes even superior) to code-switched text obtained via crowd workers who are native Hindi speakers.
Forward citations
Cited by 3 Pith papers
-
Spin-down of solar-mass protostars in magnetospheric accretion paradigm
Three-dimensional effects in magnetospheric accretion, especially failed winds and a conical disk wind, spin down solar-mass protostars within about a million years.
-
CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback
CHAI trains a reward model on GPT-4o preference labels and uses PPO to align Llama-3.1-8B for English-to-Hinglish translation, claiming a 25.66% human win-rate improvement over baselines.
-
Pre-training a Transformer-Based Generative Model Using a Small Sepedi Dataset
On a roughly 11-million-token Sepedi corpus, standard autoregressive pre-training gives lower validation loss and perplexity, while occlusion-based pre-training gives a slightly higher BLEU score on generated text.
Discussion (0). Continue with ORCID to comment.