{"id":"a90e68a0-5197-455c-b7f9-2828e76f97f6","arxiv_id":"2412.01271","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 20M-parameter adapter trained only on English image-text pairs adapts frozen diffusion models to generate images from prompts in over 110 languages, using a multilingual image-text encoder as its text encoder.","lead":"MuLan is a small adapter that lets frozen, English-only image diffusion models generate pictures from written prompts in over 110 languages after training only on English data. It could make multilingual text-to-image generation dramatically cheaper and easier for the broader community.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest quantitative evidence for cross-lingual parity uses the model itself as the evaluator; the independent Laion CLIP metric does not reproduce the 110-language claim.","rationale":"The reader's weakest_assumption focuses on InternVL's cross-lingual embedding alignment not being validated for low-resource languages. That is a real risk, and the t-SNE evidence in Appendix A.2 covers only 8 languages and 20 prompts. However, I see the more decisive load-bearing concern as the evaluator being the same model family as the frozen text encoder: even if InternVL's embedding space is perfectly language-aligned, the reported 110-language parity numbers are produced by a metric that is partly optimized for the very encoder being used. The paper does attempt to address this with the Laion CLIP table, and that table does not show the same parity, which is a direct internal contradiction. Since the concern is about the strength and verifiability of the headline numerical claim rather than about a hidden fatal flaw in the method, the reader's CONDITIONAL verdict is appropriate. The concrete test of an independent evaluator (standard CLIP after translation, or multilingual VQA) would settle whether the claim survives. I partially agree with the reader because the embedding-alignment weakness is valid but, in my reading, secondary to the evaluator-confounding issue.","tokens_in":20422,"tokens_out":1618,"duration_ms":13569,"concrete_test":"Recompute image-text similarity for the 85-language COCO2014 set in Table 8 using an evaluator not derived from InternVL or from Laion-CLIP, e.g. NLLB-translate each generated image caption set into English and score with standard OpenAI CLIP ViT-B/32, or use GPT-4o VQA per prompt. If the average non-English score falls significantly below English and below the translation-based PixArt(Google) baseline, then the parity claim is an artifact of the InternVL evaluator and the verdict should be REJECT or CONDITIONAL with the claim weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that an English-only adapter yields native generation parity across 110+ languages — rests on InternVL-LLaMA CLIP scores (Table 3, Table 8, and the abstract's 39.57 vs. 39.61). But InternVL-LLaMA is the same model family used as MuLan's frozen text encoder, so the metric is partially self-referential: the adapter is trained to map InternVL embeddings into the SD conditioning space, and CLIP scores computed with InternVL can inflate agreement between the generated image and the InternVL-encoded prompt. The paper's own independent Laion CLIP results (Table 5) show a far smaller advantage of MuLan over translation baselines, and the English-vs-other-language parity essentially disappears: MuLan-SD15 Avg 23.0 is below SD15(Google) Avg 22.3 and PixArt(Google) 24.3; MuLan-PixArt Avg 24.2 is comparable to PixArt(Google) 24.3 but lower than the English-only same-model score of 24.4. Given that the 110-language claim is the paper's headline quantitative contribution, the load-bearing weakness is that the only metric supporting the headline is computed with the very text encoder whose multilingual embedding space the method assumes is aligned. The FID numbers are mixed (MuLan-PixArt FID 11.8 is strong but AltDiffusion FID 9.1 is better), and VQAScore is only reported for 12 XM languages, not 110. Thus generalization to the 110-language parity claim is not independently confirmed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"MuLan is a multilingual text-to-image adaptation method. The authors train a lightweight adapter (<20M parameters) on English-only LAION image-text pairs while freezing a multilingual text encoder (InternVL-LLaMA) and the diffusion backbone (SD1.5, SD2.1, SDXL, or PixArt-α). The main claim is that because InternVL-LLaMA was contrastively aligned to images on noisy multilingual web data, the adapter trained in English transfers zero-shot to 110+ languages, with CLIP similarity scores of 39.57 (English) vs 39.61 (other languages), at a training cost of roughly 12 hours on 8 A100 GPUs for SD 1.5/2.1. The paper reports comparisons with translation-based baselines, AltDiffusion, GlueGen, language-specific models, and community tools, plus ablations on dataset size and alignment method.","tokens_in":20704,"tokens_out":9630,"duration_ms":79268,"significance":"If the 110-language parity claim were independently confirmed, MuLan would be a valuable practical contribution: it would let a frozen SD-family model be turned into a multilingual generator at a fraction of the cost of training a multilingual model from scratch, and its plug-and-play compatibility with LoRA, ControlNet, LCM, and IP-Adapter is genuinely useful. The paper also makes a substantive comparison of image-centered vs language-centered alignment and provides a clean data-efficiency ablation (Table 7). The central weakness is that the headline result is produced with the same text-encoder family used inside the model, and the independent Laion CLIP scores in Table 5 do not reproduce the parity; fine-grained VQAScore is only reported for 12 languages. The contribution is therefore plausible but currently not established at the scope claimed.","major_comments":[{"comment":"The headline parity (39.57 English vs 39.61 other languages) is computed with InternVL-LLaMA, the same model family that serves as MuLan's frozen text encoder. Because the adapter is trained to map InternVL text embeddings into the image decoder's conditioning space, InternVL-based CLIP scores can be inflated by this internal alignment and are not an independent measure of image-text alignment. This is not merely a hypothetical concern: the independently computed Laion CLIP scores in Table 5 show Mulan-SD15 at 23.0, essentially tied with AltDiffusion (23.1) and only 0.7 above the SD15(Google) translation baseline (22.3), while Mulan-PixArt (24.2) is below the PixArt(Google) translation baseline (24.3) and below its own English-only score (24.4). In addition, the exact values 39.57/39.61 do not appear in any table, so the computation is not reproducible. Please either (a) report independent multilingual CLIP or VQAScore over the full 85/110-language set and show the parity, or (b) restrict the parity claim to XM12 and state clearly which claims depend on the InternVL metric.","section":"§4.1, Tables 3 and 8, abstract"},{"comment":"The per-language CLIP scores reported by the authors themselves contradict a strong reading of 'comparable generation capabilities in over 110 languages.' For example, on the COCO2014 validation set, the same InternVL-MuLan-SD15 model scores 38.85 for Chinese, 38.05 for Japanese, and 38.11 for Russian, but 25.82 for Khmer, 27.56 for Irish, 27.95 for Scottish Gaelic, 28.07 for Telugu, and 30.74 for Pashto. A range of roughly 13 points between high- and low-resource languages is not 'comparable' parity; it is substantial degradation. The abstract and conclusion report only an average over languages, which hides this spread. Please report the distribution (or per-language values) for all claimed languages and either set an explicit tolerance for 'comparable' or amend the claim to 'comparable on average, with large low-resource degradation.'","section":"Table 8"},{"comment":"The method's core premise is that InternVL-LLaMA 'maintains a consistent vector space across languages,' and the paper's only direct evidence for this is the t-SNE visualization and the XM12 results. This evidence is thin and, for low-resource languages, Table 8 raises doubts about whether the premise holds. Moreover, the t-SNE description is internally inconsistent: the text says 20 captions were translated into 8 languages (160 inputs), while the Figure 5 caption says 9 prompts in 20 languages. The authors should provide a quantitative alignment check (e.g., cross-lingual text-image retrieval or nearest-neighbor agreement on a broad language sample) and use it to delineate the set of languages for which the English-trained adapter can be expected to transfer.","section":"§3.2 and Appendix A.2"}],"minor_comments":[{"comment":"The data-efficiency ablation says it was run 'without employing any training tricks,' while Section 4.1 lists 10% text-condition dropout and min-SNR weighting; clarify which setting applies to Table 7.","section":"§A.1 vs §4.1"},{"comment":"Several GlueGen cells contain '%' placeholders rather than numeric values; please replace them with the actual scores or explain the omission.","section":"Table 3"},{"comment":"The language counts should be reconciled: the abstract claims 'over 110 languages,' Table 8 contains 85 COCO languages, and Section 5 says '110 different languages'; please define the exact evaluation set and avoid double counting with XM12.","section":"Abstract, §4.1, §5"},{"comment":"The citations for the translation-based comparison and multilingual-T2I competitors (Sun et al., 2024; Yan et al., 2024) point to papers on 3D shape generation and avatar video, which appear unrelated; please replace them with correct references.","section":"§4.2 references"},{"comment":"The aesthetic-score threshold of 5.8 used for LAION filtering is a free parameter; please report sensitivity to it or justify the choice.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely a useful contribution, but the headline quantitative claim currently depends on a metric computed with the model's own text encoder. I would ask for a dedicated revision focused on independent evaluation; if the authors cannot provide independent scores for the full language set, the claims should be narrowed to XM12. The citation issues noted in the minor comments (Sun et al., Yan et al.) should be checked carefully by the editor, as they may indicate a broader reference-integrity concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is a simple, cheap adapter that gives a frozen SD-family model multilingual generation from English-only training, plus a useful empirical comparison showing that image-centered alignment (InternVL-LLaMA) beats language-centered alignment (AltClip, MultiLang-CLIP) and unaligned multilingual LLMs. That table 6 finding is the most valuable thing here, and the efficiency claim is impressive: 12 hours on 8 A100s for SD1.5/2.1, two days for SDXL/PixArt, with compatibility with LoRA, ControlNet, LCM, and IP-Adapter. The authors also report FID, VQAScore, and aesthetic scores, which is more thorough than typical for this kind of paper.\n\nBut the headline claim — \"comparable generation in 110+ languages, CLIP similarity 39.57 English vs 39.61 others\" — is built on InternVL-LLaMA CLIP scores, and InternVL-LLaMA is the same model family used as the frozen text encoder inside MuLan. That is not an apples-to-apples comparison when the main rival AltDiffusion uses a different text encoder. The independent Laion CLIP numbers in Table 5 do not reproduce the parity: MuLan-PixArt (24.2) is roughly tied with PixArt(Google) (24.3) and below the same model's English-only score (24.4); MuLan-SD15 (23.0) is about level with AltDiffusion (23.1) and below PixArt(Google). Those numbers support \"competitive with translation baselines,\" not \"state-of-the-art across low-resource languages.\" The alignment-method comparison in Table 6 is also scored with InternVL, so it may favor InternVL-LLaMA for the same reason. The low-resource evaluation also leans on Google Translate to produce the prompts, which can confound the results. No code, weights, or data are released, which matters because the central empirical claim needs independent verification. Two references are mismatched to their context, a minor but real sloppiness.\n\nNone of this kills the core idea. The adapter probably does work well for many languages, and the image-centered-alignment insight is likely correct. But the paper overclaims on the evidence presented, and the self-referential metric is a load-bearing weakness. I would send this to serious peer review — the empirical findings, especially Table 6, deserve scrutiny and replication — but the referee should require a version with a non-self-referential evaluation metric and artifact release before accepting the strong claims.","headline":"A practical, cheap multilingual adapter worth taking seriously, but the headline parity claim is carried by a self-referential metric and the independent Laion CLIP scores tell a more modest story.","tokens_in":21307,"tokens_out":3457,"would_cite":true,"duration_ms":33444,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A tiny English-trained adapter can make frozen diffusion models generate images from prompts in more than 110 languages.","keywords":["multilingual text-to-image generation","diffusion models","language adapter","zero-shot cross-lingual transfer","image-centered alignment","English-only training","frozen text encoder","CLIP similarity"],"falsifier":"Take a set of low-resource languages with non-Latin scripts that are rare in the encoder's pretraining data, translate COCO2014 prompts into them, and measure MuLan's CLIP similarity; if scores collapse toward the as-is baseline or fall below a translation-based pipeline, then the English-trained adapter has not actually transferred to those languages.","tokens_in":20187,"feed_emoji":"🌐","tokens_out":7335,"duration_ms":55669,"temperature":0.7,"pith_summary":"This paper sets out to show that native multilingual text-to-image generation does not require multilingual training data. The authors argue that when the text encoder has already been aligned across languages by contrastive training on noisy, web-scale image-text pairs, a single lightweight adapter trained only on English image-text pairs can map prompts in more than 110 languages into a frozen image generator's conditioning space. They report CLIP similarity scores of 39.57 for English and 39.61 on average for other languages, with adapter training for Stable Diffusion-class models costing about 12 hours on eight A100 GPUs using 17 million English samples. If true, multilingual capability becomes a cheap bolt-on for existing diffusion models rather than a data-intensive retraining problem.","feed_headline":"English-only training gives a frozen diffusion model 110+ languages","feed_subtitle":"A 20M-parameter adapter matches English-level CLIP scores across languages after about 12 hours of training.","key_machinery":"The load-bearing object is the image-centered aligned multilingual text encoder, here InternVL-LLaMA, whose contrastive pretraining makes the same meaning expressed in different languages land on nearby vectors. The adapter is a small trainable network of fewer than 20 million parameters, implemented as a one-layer encoder-decoder transformer with learnable queries for Stable Diffusion models, two transformers with an attention pooling layer for SDXL, and a simple MLP for PixArt-α. Its job is to project the frozen encoder's embeddings into the frozen diffusion decoder's conditioning space, so the only learning signal required is English text-to-image pairs. The paper's structural claim is that the encoder's cross-lingual alignment, not the adapter's capacity, is what supplies the multilingual generalization.","core_discovery":"The central claim is that image-centered multilingual alignment inside the text encoder is what lets an English-only adapter transfer across languages. MuLan freezes both the multilingual text encoder (InternVL-LLaMA, a decoder-only language model trained with next-token prediction and contrastive image-text learning on noisy web data) and the diffusion decoder, and trains only a small language adapter that re-projects text embeddings into the decoder's conditioning space. Because semantically equivalent prompts in different languages already sit close together in the encoder's vector space, the adapter learned from English prompts applies to every language the encoder understands. The paper reports generation quality on 12 mainstream benchmark languages and 85 translated COCO languages that matches or exceeds translation-based pipelines and dedicated multilingual models, at a small fraction of the training cost.","pith_inferences":["The method's ceiling is set by the frozen encoder's language coverage: languages the encoder has not seen enough of during contrastive pretraining would receive little or no transfer, so the natural stress test is a set of true low-resource languages with non-Latin scripts.","If the claim generalizes, it suggests multilingual generation is mostly a representation-alignment problem, and the same recipe could be applied to other frozen conditional generators such as video, 3D, or audio models.","A testable extension is to pair the same adapter design with a different image-centered aligned encoder; if performance holds, the alignment property rather than the specific encoder is the causal factor.","The parity in CLIP scores may be partly inherited from the evaluator, since both the generation and the scoring use the same multilingual encoder family; an independent human preference or multilingual VQA evaluation would be a stronger confirmation."],"forward_implications":["Multilingual text-to-image generation can be added to any compatible frozen diffusion model by training only a sub-20M-parameter adapter on English data.","Training cost drops to about 12 hours on eight A100 GPUs for SD-class models and two days for SDXL/PixArt-α, versus thousands of GPU-days for retraining-based multilingual models.","Performance on 12 mainstream languages is comparable to translation-based pipelines and exceeds dedicated Chinese and Japanese models on the XM3600 benchmark.","Because the base diffusion model is untouched, existing community tools such as LoRA, LCM, ControlNet, and IP-Adapter keep working with multilingual prompts.","The 110+ language claim is demonstrated via CLIP scores on translated COCO2014 for 85 languages plus XM3600 for 12, with low-resource languages showing the largest gains over AltDiffusion."],"supporting_citations":[{"why":"Supplies the multilingual text encoder whose image-centered alignment is the basis for cross-lingual transfer.","marker":"Chen et al., 2023b"},{"why":"The noisy web-scale image-text source used for image-centered alignment and for MuLan's English training data.","marker":"Schuhmann et al., 2022"},{"why":"The frozen diffusion backbone that the adapter plugs into for the main experiments.","marker":"Rombach et al., 2022"},{"why":"The strongest multilingual baseline, compared across mainstream and low-resource languages and on training cost.","marker":"Ye et al., 2023a"},{"why":"Provides the XM12 benchmark used for the 12-language quantitative evaluation.","marker":"Thapliyal et al., 2022"},{"why":"Its validation set, machine-translated into 85 languages, is used to demonstrate the 110+ language generalization.","marker":"Lin et al., 2015"},{"why":"Prior plug-and-play adapter approach using language-centered alignment, the main contrast that motivates image-centered alignment.","marker":"Qin et al., 2023"},{"why":"One of the frozen diffusion backbones adapted, requiring a two-transformer adapter design.","marker":"Podell et al., 2023"},{"why":"Frozen DiT backbone used with the simpler MLP adapter.","marker":"Chen et al., 2023a"}],"fun_headline_variants":["English-only training teaches diffusion models 110+ languages","A 20M adapter gives diffusion models 110+ languages","From English alone to 110 languages: MuLan's tiny adapter","110 languages, one tiny adapter, no retraining","Tiny adapter, huge language coverage: MuLan adds 110+"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole multilingual transfer rests on the premise that the text encoder's internal representations of the same prompt are already nearly identical across languages, a property the paper verifies for only 12 mainstream languages and presumes for the rest of the 110+.","fun_headline_variants_meta":{"raw":{"variants":["English-only training teaches diffusion models 110+ languages","A 20M adapter gives diffusion models 110+ languages","From English alone to 110 languages: MuLan's tiny adapter","110 languages, one tiny adapter, no retraining","Tiny adapter, huge language coverage: MuLan adds 110+"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3238,"prompt_tokens":891,"completion_tokens":2347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":2262}},"tokens_in":507,"tokens_out":2347,"duration_ms":17234,"temperature":1.0,"reasoning_tokens":2262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:30:37.259671+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of low-resource languages with non-Latin scripts that are rare in the encoder's pretraining data, translate COCO2014 prompts into them, and measure MuLan's CLIP similarity; if scores collapse toward the as-is baseline or fall below a translation-based pipeline, then the English-trained adapter has not actually transferred to those languages.","supporting_citations":[{"cited_title":"Gluegen: Plug and play multi-modal encoders for x-to-image generation","cited_arxiv_id":null,"evidence_quote":"Prior plug-and-play adapter approach using language-centered alignment, the main contrast that motivates image-centered alignment."}],"review_version":1}