{"id":"62e81c1d-466c-43fb-a811-b14487d89c7a","arxiv_id":"2505.14729","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across six VLMs and 211 countries, country-recognition accuracy varies widely and errors concentrate on a few overpredicted countries (USA, India, Brazil), not on a uniform Western bias.","lead":"This paper tests six vision-language models on identifying the country of photos from 211 countries, under different question formats, languages, and image distortions. It finds accuracy varies strongly by country and that models overpredict a few countries such as the USA, India, and Brazil, rather than showing a simple Western bias.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The conclusion that overprediction of USA/India/Brazil reflects training-data overrepresentation is unsupported: Section 5.5 hedges with 'or benefit from more visually distinctive cues', yet Section 7 asserts it as fact, with no measurement of training data or control for image distinctiveness.","rationale":"The reader's weakest assumption identified the same core issue: the attribution of prediction patterns to training-data overrepresentation is unmeasured. I agree, and I sharpen it by pointing to an internal inconsistency in the manuscript: Section 5.5 presents the overrepresentation explanation as one of two hypotheses ('overrepresented in the models' pretraining data or benefit from more visually distinctive cues'), while Section 7 asserts it as an established finding. This is not a matter of external consensus; the paper's own data can discriminate between the two hypotheses through the category labels it has already created. A category-controlled re-analysis is inexpensive and directly tests whether the effect is due to the visual content composition of Country211. If the effect survives the control, the conclusion remains plausible but still requires direct training-data evidence; if it does not, the central claim is unsupported. I therefore see no reason to move the reader's verdict: CONDITIONAL is the right level, as the paper should be asked to provide this analysis (or a human baseline and/or training-data statistics) before the causal reading is accepted. One additional caveat: the open-ended scoring procedure is not described, and this could affect the prediction distributions in Figure 7; that is a secondary reproducibility concern but not the central interpretive problem.","tokens_in":19916,"tokens_out":10234,"duration_ms":99752,"concrete_test":"Re-analyze the open-ended mispredictions using the paper's own nine image-category labels (Table 4). For each country, compute how often the models predict USA, India, or Brazil, both overall and within each image category. Then reweight or stratify per category so that each country contributes equally to every category, and recompute the overprediction rates. If the overprediction of USA/India/Brazil largely disappears when category composition is controlled, the central conclusion is an artifact of Country211's unequal category mix (a form of visual distinctiveness) rather than a training-data frequency effect. If it persists within category-balanced samples, the distinctiveness-as-category-imbalance explanation is weakened, and the training-data attribution requires direct measurement of pretraining data composition to be supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in Section 7, is that model biases 'reflect over representation of certain countries in training data', specifically the overprediction of the USA, India, and Brazil seen in open-ended responses (Section 5.5, Figure 7). This is an inference, not a measurement. No training-data composition is ever examined: the only evidence cited is the response distribution itself and, in footnote 2, a result about text-to-image models (Basu et al., 2023), which does not speak to VLM discriminative behavior. Section 5.5 itself states only a hypothesis: 'We hypothesize that these countries are likely overrepresented in the models' pretraining data or benefit from more visually distinctive cues.' The alternative, visual distinctiveness, is not tested, and the Section 7 conclusion silently drops it. The paper even shows (Section 5.4, Figure 13) that image category strongly affects accuracy, and Country211 is balanced by country count but not by category composition, so category imbalance alone could generate the USA/India/Brazil overprediction (e.g., India's many exterior-architecture images creating a default). No human baseline or distinctiveness control is provided. Because the conclusion is the paper's headline contribution (the 'not uniformly Western' framing), the unsupported causal attribution is load-bearing: if overprediction reflects distinctiveness or test-set composition rather than pretraining-data frequency, the paper's call for dataset-composition transparency loses its basis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates six vision-language models (Gemini-2.5-Flash, Gemma-3-27B, Gemma-3-12B, Aya-Vision-8B/32B, GPT-4o-Mini) on the Country211 image-based country identification benchmark, covering 211 countries under open-ended, multiple-choice (with random and similar distractors), multilingual (English, Hindi, Chinese, Portuguese, Spanish), and image-perturbation conditions. It reports large country-level accuracy disparities, consistent overprediction of the USA, India, and Brazil in open-ended responses, and sensitivity patterns across image categories and perturbations. The paper concludes that VLM cultural bias is not uniformly Western but reflects overrepresentation of certain countries in training data, and argues for greater dataset-composition transparency.","tokens_in":20194,"tokens_out":4536,"duration_ms":42905,"significance":"If the central finding holds, the paper is a useful corrective to the common 'uniform Western bias' narrative and provides a broad, reproducible evaluation matrix: 168.8K samples, 6 models, 5 prompt languages, 9 image categories, and perturbation analysis, with country-level accuracy tables in the appendix and datasets/code released. The paper's most distinctive contribution is the claim that VLM biases concentrate on a few overrepresented or visually salient countries (USA, India, Brazil) rather than being broadly Western. However, that claim currently rests on an unsupported causal inference from model output frequencies to pretraining-data composition, and the open-ended answer-normalization procedure is underspecified, so the headline result needs strengthening before it can be accepted as stated.","major_comments":[{"comment":"The central conclusion in Section 7, that overprediction of the USA, India, and Brazil 'reflect[s] over representation of certain countries in training data', is not supported by the evidence presented. Section 5.5 itself offers only a hypothesis ('likely overrepresented in the models' pretraining data or benefit from more visually distinctive cues'), and Section 6 restates it as 'reinforcing the role of training data bias' without any measurement of training data. The only external support cited, footnote 2 (Basu et al., 2023), concerns text-to-image generation, not discriminative VLM behavior. Given that Section 5.4 (Figure 13) shows image category strongly affects accuracy, and Country211 is balanced by country count but not controlled for category composition, the observed overprediction could plausibly stem from category imbalance or visual distinctiveness rather than training-data frequency. Please either (a) soften the conclusion to a hypothesis, or (b) add targeted controls, such as a human or zero-shot CLIP baseline, a per-image-category overprediction analysis, or direct evidence about the country distribution in the models' pretraining data. This is load-bearing for the paper's headline contribution.","section":"Sections 5.5, 6, 7 (and footnote 2)"},{"comment":"The open-ended evaluation requires mapping free-form country names from model responses to ISO-3166 labels, but the manuscript never specifies the normalization or matching procedure. Section 3's note that 'the list of tags and corresponding country names led to the models responding consistently' is not a specification: it does not state whether exact string matching, fuzzy matching, an LLM-based resolver, or manual adjudication was used, nor how non-English or synonymous answers (e.g., 'Britain' vs. 'United Kingdom') were handled. Since the response distribution in Figure 7 and all open-ended accuracy numbers depend on this mapping, the results are not reproducible without this detail. Please document the answer-to-label pipeline and include a brief error analysis.","section":"Sections 3, 4.1, Appendix D"},{"comment":"The reported Cochran's Q statistics are negative (e.g., -0.727, -0.952), but Cochran's Q is defined as a nonnegative statistic; negative values, together with p-values of ~1.0 reported to several decimal places, indicate a misapplication of the test or of the p-value computation. The Pearson chi-square rows also lack a clear description of the contingency tables being tested (across models? across countries?). Please clarify the exact statistical procedures and re-report the results. This does not affect the country-level accuracy tables, but it undermines the claims in Section 5 that perturbations and language 'do not meaningfully alter' overall accuracy.","section":"Table 2"}],"minor_comments":[{"comment":"The sentence 'proving the list of tags and corresponding country names led to the models responding consistently' appears to contain a typo (likely 'providing' or 'promising') and should be rephrased for clarity.","section":"Section 3"},{"comment":"The selection of similar distractors is described only as 'chosen from among the bordering nations' with manual additions for culturally similar countries; please specify how many distractors were used, how they were sampled, and whether the selection was balanced across countries and conditions.","section":"Section 4.3"},{"comment":"Please clarify whether the response distribution in Figure 7 is pooled across models or per-model, and whether it is normalized by the number of samples per country; the current caption does not make this explicit.","section":"Figure 7"},{"comment":"The reproducibility section reports seed values and API providers but not the specific model versions or access dates for the proprietary models; adding version identifiers would further support reproducibility.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely question and the empirical scope is impressive, but the headline causal claim goes beyond the evidence. The main revision should focus on either reframing the conclusion as a hypothesis or adding the missing controls. I would also suggest the editor ask the authors to clarify the open-ended answer normalization, as this is a reproducibility issue that affects the paper's central figure. The statistical reporting in Table 2 should be corrected independently."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper is a genuinely broad empirical evaluation—six VLMs, 211 countries, open-ended and MCQ (random vs. similar distractors), five languages, four perturbations, nine image categories. The country-level accuracy tables alone are worth having. The finding that errors are not a simple Western/non-Western split, but cluster on USA, India, and Brazil, is real and consistent across models. That is the paper's best contribution.\n\nWhat is new: scale, and the combination of random vs. similar distractors with multilingual prompts. The result that language matters very little (except for the target-language country itself) and that similar distractors produce different error patterns is useful. The perturbation analysis and category breakdown also add texture. The paper is honest enough to release code and data, and the appendix shows reproducibility attempts (three runs, seed fixed).\n\nThe soft spots are in the interpretation. Section 5.5 says the USA/India/Brazil overprediction is \"likely overrepresented in the models' pretraining data or benefit from more visually distinctive cues.\" Section 7 then asserts the overrepresentation explanation as fact. That is a jump. The paper never measures training data composition, and the alternative (visual distinctiveness, or category imbalance in Country211) is not controlled. Given that image category strongly affects accuracy, and Country211 is balanced by country but not by category mix, this is a real confound. The stress-test note has this right: the causal claim is load-bearing for the \"not uniformly Western\" framing, and it currently rests on an untested assumption.\n\nAlso, the open-ended scoring procedure is underspecified. How were free-form country names mapped to ISO codes? The text hints at manual mapping but doesn't give the rules, and open-ended accuracy is a headline number. Error bars are missing from most figures; appendix C says runs varied by 1–1.2% overall, but the tables report single values. There are also small sample-count inconsistencies (63,300 vs 168.8K) and a few broken citations. These are fixable.\n\nWho this is for: anyone working on cultural or geographic bias in VLMs. The benchmark itself is a useful diagnostic resource. The central interpretation needs either evidence (training-data composition, distinctiveness controls) or a downgrade to a descriptive finding. I would send it to review—it deserves referee time—but I'd expect the authors to address the attribution before publication.\n\nRecommendation: engage with it, but read the 5.5–7 transition carefully; that is where the paper overreaches.","headline":"Broad, useful benchmark of VLM country recognition; the central causal claim about training data is an inference the paper doesn't back up.","tokens_in":20745,"tokens_out":3586,"would_cite":true,"duration_ms":32825,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that vision-language models' cultural bias is concentrated in a few overrepresented countries rather than being uniformly Western, and shows country-level evaluation is required to expose it.","keywords":["vision-language models","cultural bias","country identification","Country211","geographic representation","multilingual prompting","multiple-choice evaluation","pretraining data bias"],"falsifier":"Measure the geotagged country frequencies in the actual pretraining data behind these models and compare them with per-country overprediction rates; if a low-frequency country is overpredicted as often as India, the data-overrepresentation explanation fails. Training or fine-tuning a vision-language model on country-balanced data provides the control: if USA, India, and Brazil overprediction persists, visual salience, not data volume, drives the bias.","tokens_in":19714,"feed_emoji":"🌍","tokens_out":7303,"duration_ms":64689,"temperature":0.7,"pith_summary":"This paper tries to establish where cultural bias in vision-language models actually lives. Using image-based country identification as a proxy for cultural recognition, it tests six models on 100 images from each of 211 countries and finds that accuracy varies enormously by country, with near-chance open-ended performance across much of Africa and Central America. The central claim is that this disparity is not a uniform Western bias: models consistently overpredict the USA, India, and Brazil regardless of the true country, a pattern the authors attribute to overrepresentation in pretraining data. If correct, the practical implication is that country-level evaluation and training-data transparency, not generic debiasing, are the levers that would make vision-language models more culturally equitable.","feed_headline":"Cultural bias in VLMs is concentrated, not Western","feed_subtitle":"Six models, 211 countries: recognition errors funnel toward the USA, India, and Brazil.","key_machinery":"The central instrument is the country-identification probe built on the Country211 dataset, which contains 100 GPS-labelled images from each of 211 countries, so every country is equally represented in the test set. The paper combines this balanced probe with three prompting formats, namely open-ended questions, multiple-choice with random distractors, and multiple-choice with culturally similar distractors, plus five prompt languages and image perturbations. The mechanism that carries the argument is the response-distribution and misclassification-map analysis: instead of only reporting accuracy, the paper traces where wrong answers go, which is what exposes the USA-India-Brazil overprediction cluster and the frequent collapse of African and South American images onto India.","core_discovery":"On the paper's own terms, the discovery is that cultural bias in vision-language models is concentrated rather than uniform. When asked to name the country of an image, all six models, proprietary and open-weight alike, overpredict a small set of nations, namely the USA, India, and Brazil, no matter what the ground-truth label is, while countries such as Angola, the Central African Republic, and Eswatini are recognized at rates near or at zero in open-ended questioning. This pattern persists across prompt languages and is made worse by image rotation and grayscale conversion. The paper concludes that the biases \"are not uniformly Western but instead reflect over representation of certain countries in training data,\" and argues that country-level evaluation is required to surface disparities that regional averages hide.","pith_inferences":["A direct test of the paper's central mechanism would be to measure geotagged country frequencies in the actual pretraining corpora and correlate them with each model's overprediction rates; the paper only infers the data from model outputs.","If overprediction is driven by visual distinctiveness rather than data volume, countries with highly recognizable flags, architecture, or attire would be overpredicted even by models trained on perfectly balanced data; the two explanations are separable and testable.","The same probe could be turned into an intervention: fine-tuning with hard negatives drawn from culturally similar countries might selectively improve recognition for underrepresented countries.","Extending the country proxy to sub-national or regional labels would test whether the concentration pattern persists at finer cultural granularity, which the paper's own limitation note suggests it might not."],"forward_implications":["If the concentration pattern is real, country-level accuracy must replace regional accuracy as the reporting unit; the regional tables in the paper mask that some countries are recognized at near-zero rates.","If pretraining overrepresentation drives overprediction, then documenting the geographic composition of vision-language-model training corpora becomes a necessary step for any fairness claim about these models.","If prompt language barely moves accuracy, with under 2 percent difference across English, Hindi, Chinese, Portuguese, and Spanish, then multilingual prompting alone is not a remedy for cultural bias.","If multiple-choice questions with culturally similar distractors reveal confusions that random-distractor questions hide, then evaluation suites that only use easy multiple-choice formats will overstate cultural competence.","If rotation and grayscale hurt some countries and models far more than others, then deployments on imperfect real-world images will inherit and likely widen the same disparities."],"supporting_citations":[{"why":"Supplies the Country211 dataset, the 211-country balanced image set that defines the evaluation task.","marker":"Radford et al., 2021"},{"why":"Provides YFCC100M, the GPS-tagged image collection from which Country211 images are drawn.","marker":"Thomee et al., 2016"},{"why":"Documents the geographic skew of web image datasets, where 32 percent of geolocatable OpenImages samples come from the US and 60 percent from six Western countries, the external evidence that pretraining data overrepresent certain places.","marker":"Shankar et al., 2017"},{"why":"Shows text-to-image models default to US and Indian surroundings, supporting the paper's claim that specific countries are overrepresented in training data.","marker":"Basu et al., 2023"},{"why":"Prior result that geolocation performance tracks regional training prevalence, motivating country identification as a bias probe.","marker":"Pouget et al., 2024"},{"why":"Establishes the Western-bias baseline that this paper's concentrated-overrepresentation finding qualifies.","marker":"de Vries et al., 2019"}],"fun_headline_variants":["VLM bias: errors cluster on USA, India, Brazil","Cultural bias in VLMs is concentrated, not Western","Country-level test reveals uneven VLM recognition","Vision models overpredict a few nations globally","VLMs' cultural blind spots are not what you'd think"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on the untested assumption that a model's overprediction frequencies mirror the geographic composition of its pretraining data, and that GPS-derived country labels are a valid proxy for cultural origin; the paper never measures the training corpora themselves.","fun_headline_variants_meta":{"raw":{"variants":["VLM bias: errors cluster on USA, India, Brazil","Cultural bias in VLMs is concentrated, not Western","Country-level test reveals uneven VLM recognition","Vision models overpredict a few nations globally","VLMs' cultural blind spots are not what you'd think"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1613,"prompt_tokens":834,"completion_tokens":779,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":703}},"tokens_in":450,"tokens_out":779,"duration_ms":7345,"temperature":1.0,"reasoning_tokens":703,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:09:36.325577+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the geotagged country frequencies in the actual pretraining data behind these models and compare them with per-country overprediction rates; if a low-frequency country is overpredicted as often as India, the data-overrepresentation explanation fails. Training or fine-tuning a vision-language model on country-balanced data provides the control: if USA, India, and Brazil overprediction persists, visual salience, not data volume, drives the bias.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides YFCC100M, the GPS-tagged image collection from which Country211 images are drawn."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the geographic skew of web image datasets, where 32 percent of geolocatable OpenImages samples come from the US and 60 percent from six Western countries, the external evidence that pretraining data overrepresent certain places."},{"cited_title":"Inspecting the Geographical Representativeness of Images from Text-to-Image Models","cited_arxiv_id":"2305.11080","evidence_quote":"Shows text-to-image models default to US and Indian surroundings, supporting the paper's claim that specific countries are overrepresented in training data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior result that geolocation performance tracks regional training prevalence, motivating country identification as a bias probe."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the Western-bias baseline that this paper's concentrated-overrepresentation finding qualifies."}],"review_version":1}