{"id":"c641a602-a7a2-484b-88da-e69930011406","arxiv_id":"2605.26575","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Hubness dominates as the causal driver of cross-lingual retrieval asymmetry in multilingual embeddings, with CSLS correction closing most of the reciprocity gap across five models and a 6518-item parallel corpus.","lead":"Multilingual embedding models exhibit asymmetric cross-lingual retrieval where a translation match in one direction often fails in reverse. The paper tests and supports the claim that hubness in the embedding space, rather than anisotropy or other geometric issues, is the main driver of this asymmetry.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Idiomatic/proverbial corpus may not represent asymmetries in production retrieval; effects could be domain-specific.","rationale":"The reader's weakest assumption explicitly flags the same generalizability risk; the abstract-only review already noted the lack of broader validation. This single assumption is load-bearing because all quantitative support for the 'hubness, not anisotropy' claim rests on one corpus type. A positive replication on literal text would strengthen the claim; failure would confine it to idiomatic phenomena.","tokens_in":1815,"tokens_out":377,"duration_ms":21845,"concrete_test":"Re-run the full pipeline (reciprocity computation, hub-mass/anisotropy/centroid/magnitude extraction, joint regression, and CSLS vs ablation) on a matched-size parallel corpus of literal sentences (e.g., 6,500 Tatoeba or Europarl pairs in the same four languages) using the identical five encoders; if hub-mass dominance share falls below 30% or CSLS gap closure drops below 40%, the representativeness assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central mechanistic claim (hub mass dominance in joint regression, CSLS closing 63.5% of reciprocity gap) is derived entirely from 6,518 parallel idiomatic and proverbial expressions. These items frequently exhibit non-compositional semantics, fixed collocations, and lower lexical diversity than the literal sentences or documents typical in cross-lingual retrieval pipelines. Consequently, the observed statistical dissociation between hubness and anisotropy, the partial R² values (0.302 vs 0.003), and the ablation-vs-CSLS contrast may reflect properties of this narrow domain rather than a general geometric pathology of multilingual embedding spaces. The pre-registered falsification conditions address internal isolation of predictors but do not test external validity across corpus types.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that hubness (specifically hub mass), rather than anisotropy, centroid drift, or magnitude, is the dominant driver of asymmetric cross-lingual retrieval (measured as lack of mutual nearest-neighbor reciprocity) in multilingual embedding models. Using a parallel corpus of 6,518 idiomatic and proverbial expressions in English, Bangla, Hindi, and Arabic embedded by five production encoders (Gemini, Mistral, OpenAI-L, OpenAI-S, Qwen), five pre-registered experiments with a priori falsification conditions show hub mass dominating a joint regression on reciprocity (49.5% dominance share, 1.68x the next predictor; partial R²=0.302 vs. 0.003 for anisotropy), while the hub-aware CSLS correction closes 63.5% of the worst-to-best reciprocity gap (with mean within-model effect size 130x larger than surgical hub-vector ablation). The work also claims to resolve the anisotropy-hubness paradox by statistical dissociation and recommends replacing cosine with CSLS as the default retrieval metric.","tokens_in":1991,"tokens_out":601,"duration_ms":21852,"significance":"If the central mechanistic dissociation and the superiority of CSLS hold beyond the tested corpus, the result would be significant for multilingual retrieval pipelines, as it identifies a metric-level pathology rather than a vector-level one and supplies a simple, previously published correction with large reported effect sizes. The pre-registered design, explicit falsification conditions, and use of an external correction (CSLS) rather than a fitted quantity are methodological strengths that support internal validity of the regression and ablation contrasts.","major_comments":[{"comment":"The central claim that hubness dominates anisotropy as a driver of reciprocity (49.5% dominance share, partial R²=0.302 vs. 0.003) and that CSLS closes 63.5% of the gap is derived entirely from the 6,518 idiomatic/proverbial expressions. These items exhibit non-compositional semantics and lower lexical diversity than the literal sentences or documents typical in production cross-lingual retrieval; the pre-registered falsification conditions isolate predictors internally but do not test external validity across corpus types. This is load-bearing for the general recommendation to replace cosine with CSLS in multilingual embedding pipelines.","section":"Data and Experiments sections (corpus of 6,518 idiomatic expressions)"},{"comment":"The abstract reports dominance shares and partial R² values without error bars or details of the joint regression specification (e.g., exact predictors, multicollinearity checks, or held-out split procedure). Given that the 1.68x dominance claim and the 0.302 vs. 0.003 contrast are central to the mechanistic argument, these omissions make it difficult to assess stability of the reported effect sizes.","section":"Abstract and regression results"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments and the positive assessment of the methodological strengths. We address the major comments point by point below.","responses":[{"response":"We selected the idiomatic corpus to ensure high-quality parallel translations with controlled semantics, enabling precise measurement of reciprocity. While we agree that this corpus type may limit generalizability to literal text, the geometric properties tested (hubness, anisotropy) are intrinsic to the embedding spaces and not corpus-specific. Nevertheless, to address external validity concerns, we will revise the discussion and conclusion to explicitly note the corpus limitation and recommend validation on additional corpora before broad adoption of CSLS. This constitutes a partial revision.","revision_made":"partial","referee_comment":"The central claim that hubness dominates anisotropy as a driver of reciprocity (49.5% dominance share, partial R²=0.302 vs. 0.003) and that CSLS closes 63.5% of the gap is derived entirely from the 6,518 idiomatic/proverbial expressions. These items exhibit non-compositional semantics and lower lexical diversity than the literal sentences or documents typical in production cross-lingual retrieval; the pre-registered falsification conditions isolate predictors internally but do not test external validity across corpus types. This is load-bearing for the general recommendation to replace cosine with CSLS in multilingual embedding pipelines."},{"response":"We accept this point. The full regression details, including predictors, VIF checks for multicollinearity, and cross-validation procedure, are provided in the Methods section. For the abstract, we will add a brief mention of the regression approach and report standard errors or confidence intervals for the key statistics in a revised version. This will improve transparency without changing the results.","revision_made":"yes","referee_comment":"The abstract reports dominance shares and partial R² values without error bars or details of the joint regression specification (e.g., exact predictors, multicollinearity checks, or held-out split procedure). Given that the 1.68x dominance claim and the 0.302 vs. 0.003 contrast are central to the mechanistic argument, these omissions make it difficult to assess stability of the reported effect sizes."}],"tokens_in":1660,"tokens_out":476,"duration_ms":33014,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that on their 6518 parallel idioms and proverbs, hub mass explains far more of the reciprocity deficit than anisotropy does, and CSLS recovers most of the gap while simple hub ablation does not.\n\nThe paper runs five pre-registered experiments across Gemini, Mistral, OpenAI-L, OpenAI-S, and Qwen. It reports hub mass at 49.5% dominance share in the joint regression, 1.68 times the next predictor, with partial R² of 0.302 versus 0.003 for anisotropy. CSLS closes 63.5% of the worst-to-best reciprocity gap and produces a within-model effect 130 times larger than ablation. They also show the two geometric properties are statistically separable, which addresses the old paradox without forcing one into the other. The design uses held-out splits and external corrections, so the central numbers are not tautological.\n\nThe soft spot is the corpus itself. Idioms and proverbs are non-compositional and lexically constrained; they are not the literal sentences or documents that dominate production cross-lingual retrieval. The observed dissociation and the size of the CSLS improvement could be tied to that domain. The paper does not report tests on other text types, so the claim that hubness is the dominant driver in multilingual spaces rests on how representative these expressions are.\n\nThis work is for people who already care about geometric fixes in multilingual embeddings. The structured experiments and quantitative head-to-head comparison are clear enough to justify sending it to referees, though any review would likely press for checks on broader data.","headline":"Hubness dominates the regression on this idiomatic corpus while anisotropy barely registers, but the narrow domain leaves the general claim open.","tokens_in":2508,"tokens_out":394,"would_cite":false,"duration_ms":24625,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Hubness in multilingual embedding spaces drives asymmetric cross-lingual retrieval, while anisotropy does not.","keywords":["hubness","multilingual embeddings","cross-lingual retrieval","reciprocity","CSLS","anisotropy","embedding asymmetry"],"falsifier":"A replication in which hub mass fails to show the highest dominance share in the joint regression on reciprocity or in which CSLS fails to close a substantial share of the reciprocity gap.","tokens_in":2713,"feed_emoji":"","tokens_out":624,"duration_ms":22818,"temperature":0.7,"pith_summary":"The paper tests why cross-lingual retrieval fails to be symmetric in multilingual embedding models despite the assumption that it should be. It embeds a parallel corpus of 6518 idiomatic expressions across English, Bangla, Hindi, and Arabic using five production encoders and measures failures via mutual nearest-neighbour reciprocity. The central claim is that hubness dominates other geometric properties as the causal driver of asymmetry. A hub-aware correction substantially reduces the observed gaps, indicating the issue resides in the similarity metric itself rather than vector properties.","feed_headline":"Hubness, not anisotropy, drives cross-lingual retrieval asymmetry","feed_subtitle":"Hub-aware correction closes 63.5 percent of the reciprocity gap across five encoders.","key_machinery":"Hub mass as the dominant predictor in regressions on reciprocity, contrasted against anisotropy, centroid drift, and magnitude; the CSLS score correction that accounts for hubness in the similarity computation.","core_discovery":"Across five pre-registered experiments, hub mass dominates a joint regression on reciprocity with a 49.5% dominance share and partial R² of 0.302, compared to 0.003 for anisotropy. The CSLS correction closes 63.5% of the worst-to-best reciprocity gap and produces a mean within-model effect size 130 times larger than surgical hub-vector ablation, establishing that hubness is a pathology of the similarity metric. The experiments also show that anisotropy and hubness are statistically dissociable.","pith_inferences":["Similar metric-driven asymmetries could appear in high-hubness monolingual retrieval settings.","Model training objectives could incorporate explicit penalties for hub formation to reduce downstream asymmetry.","Extending the tests to sentence-level or document-level retrieval would show whether the hubness effect scales beyond short idiomatic phrases."],"forward_implications":["Replacing cosine similarity with CSLS as the default metric improves reciprocity in multilingual retrieval pipelines.","Hubness and anisotropy are statistically separable, so interventions can target one without the other.","The asymmetry arises from the similarity computation rather than properties of individual hub vectors.","Production systems should apply hub-aware corrections before addressing other geometric properties."],"fun_headline_variants":["Hubness dominates cross-lingual reciprocity over anisotropy","Hub-aware correction closes 63.5 percent reciprocity gap","Hub mass explains 49.5 percent dominance in reciprocity regression","CSLS outperforms hub ablation in fixing embedding asymmetry","Anisotropy and hubness are statistically dissociable in embeddings"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 6518 idiomatic and proverbial expressions form a representative sample of the phenomena that drive asymmetry in production multilingual retrieval systems.","fun_headline_variants_meta":{"raw":{"variants":["Hubness dominates cross-lingual reciprocity over anisotropy","Hub-aware correction closes 63.5 percent reciprocity gap","Hub mass explains 49.5 percent dominance in reciprocity regression","CSLS outperforms hub ablation in fixing embedding asymmetry","Anisotropy and hubness are statistically dissociable in embeddings"]},"model":"grok-4.3","cost_usd":0.005915,"raw_usage":{"total_tokens":2849,"prompt_tokens":751,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":59149500,"prompt_tokens_details":{"text_tokens":751,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2028,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":751,"tokens_out":70,"duration_ms":15549,"temperature":1.0,"reasoning_tokens":2028,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:37:22.823381+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication in which hub mass fails to show the highest dominance share in the joint regression on reciprocity or in which CSLS fails to close a substantial share of the reciprocity gap.","supporting_citations":[],"review_version":1}