Pith. sign in

REVIEW 1 cited by

Allocating Large Vocabulary Capacity for Cross-lingual Language Model Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.07306 v1 pith:WSX4CFVM submitted 2021-09-15 cs.CL

classification cs.CL
keywords vocabularycross-linguallanguagepre-trainingcapacitymodelsincreasingk-nn-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Compared to monolingual models, cross-lingual models usually require a more expressive vocabulary to represent all languages adequately. We find that many languages are under-represented in recent cross-lingual language models due to the limited vocabulary capacity. To this end, we propose an algorithm VoCap to determine the desired vocabulary capacity of each language. However, increasing the vocabulary size significantly slows down the pre-training speed. In order to address the issues, we propose k-NN-based target sampling to accelerate the expensive softmax. Our experiments show that the multilingual vocabulary learned with VoCap benefits cross-lingual language model pre-training. Moreover, k-NN-based target sampling mitigates the side-effects of increasing the vocabulary size while achieving comparable performance and faster pre-training speed. The code and the pretrained multilingual vocabularies are available at https://github.com/bozheng-hit/VoCapXLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Output token count, observable through response timing, can reveal a user's target language or classification result with 70-87% accuracy in the authors' experiments.

Pith tools