Pith. sign in

REVIEW 6 cited by

Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17790 v1 pith:BCWLLW7T submitted 2024-04-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords japanesepre-trainingcontinualenglishperformancecross-linguallanguagevocabulary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-lingual continual pre-training of large language models (LLMs) initially trained on English corpus allows us to leverage the vast amount of English language resources and reduce the pre-training cost. In this study, we constructed Swallow, an LLM with enhanced Japanese capability, by extending the vocabulary of Llama 2 to include Japanese characters and conducting continual pre-training on a large Japanese web corpus. Experimental results confirmed that the performance on Japanese tasks drastically improved through continual pre-training, and the performance monotonically increased with the amount of training data up to 100B tokens. Consequently, Swallow achieved superior performance compared to other LLMs that were trained from scratch in English and Japanese. An analysis of the effects of continual pre-training revealed that it was particularly effective for Japanese question answering tasks. Furthermore, to elucidate effective methodologies for cross-lingual continual pre-training from English to Japanese, we investigated the impact of vocabulary expansion and the effectiveness of incorporating parallel corpora. The results showed that the efficiency gained through vocabulary expansion had no negative impact on performance, except for the summarization task, and that the combined use of parallel corpora enhanced translation ability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

    cs.CL 2026-08 conditional novelty 6.0 of 10

    The measured native-versus-translate reasoning gap on MGSM depends strongly on the output-token budget, nearly vanishing at saturation and reversing direction under tight caps.

  2. Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

    cs.CL 2026-08 conditional novelty 6.0 of 10

    For Hindi vocabulary extension of a 30B LLM, the best embedding initialization is uniform subword averaging with Hindi norm calibration on the input and character-length-weighted averaging on the output, cutting conti...

  3. In-Place Tokenizer Expansion for Pre-trained LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Continuing a model's own BPE merges and training only new embedding rows preserves quality while cutting token counts 2.4–4× for previously under-tokenized languages.

  4. Intersectional Bias in Japanese Large Language Models from a Contextualized Perspective

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Using the new inter-JBBQ benchmark, biased responses of Japanese LLMs vary with scenario context even when similar social attribute combinations are tested.

  5. The Emergence of Abstract Thought in Large Language Models Beyond Any Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Across 20 open LLMs, shared multilingual neurons grow in number and per-neuron importance over release generations, which the authors interpret as evidence of language-agnostic abstract thought and use to guide neuron...

  6. Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An adversarial beam-search method generates over 6,000 bilingual question pairs that reliably make multilingual LLMs perform far worse in non-English languages than in English.

Pith tools