{"id":"f0abc9f7-dcbe-4a0f-973f-cadb9f002493","arxiv_id":"2607.07895","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A human-LLM collaborative pipeline yields EspanStereo, a multi-country Spanish stereotype dataset that exposes region-specific biases in Spanish LLMs and diverges sharply from English-centric resources.","lead":"The paper builds EspanStereo, a Spanish stereotype dataset for five countries, by having LLMs propose candidates that in-culture annotators then validate and instantiate. It shows Spanish-supporting LLMs encode country-specific stereotypes that English or translated datasets miss, and that a human-LLM pipeline can scale such resources.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"LLM-generated candidate pool may systematically under-sample culture-specific stereotypes, so high validation rates do not establish completeness of EspanStereo.","rationale":"The reader correctly isolates the completeness assumption in Section 3.1 / Limitations as the weakest link. High validation rates establish that the LLM-generated candidates that survive majority vote are culturally recognized; they do not establish that the generation step surfaces a representative sample of the stereotypes that exist. The Nicaragua race result and the authors’ own Limitations paragraph already flag the risk. Because the paper never measures recall against an independent human free-list or ethnographic baseline, the claims of cultural specificity and of country-dependent encoding rest on an untested sampling premise. The concrete free-list recall test would settle the issue without requiring a full re-annotation of the dataset. The rest of the empirical package (released data, low English overlap, pruning results) remains useful, so the verdict stays CONDITIONAL rather than moving to REJECT; the concern simply confirms that the conditionality is load-bearing.","tokens_in":24375,"tokens_out":601,"duration_ms":8851,"concrete_test":"For one country (e.g., Colombia), recruit a fresh panel of 10–15 in-culture annotators who have never seen the LLM list and ask them to free-list 15–20 race/religion stereotypes they consider common. Compute the fraction of these free-listed stereotypes that already appear (after synonym matching) among the validated EspanStereo entries for that country. If recall is substantially below ~70 %, the completeness assumption fails and the strongest claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the framework yields high-quality, country-specific stereotypes that are largely absent from English resources and literature rests on the assumption that injection-prompted LLM generations (Section 3.1) produce a sufficiently complete and unbiased candidate pool. Human validation (Section 3.2) then mainly filters noise (median Likert ≤2 discarded). High validation rates (Table 7, mostly >85 %) and low overlap with StereoSet/CrowS-Pairs (Tables 11–12) and literature (Table 10) are consistent with this, yet they only measure precision of the emitted set. They do not measure recall against the stereotypes actually circulating in each culture. The Limitations section itself notes weaker coverage of less-prominent or emerging stereotypes; Nicaragua race validation drops to 36 % precisely where the LLM emitted many immigration stereotypes that annotators rejected. If entire classes of local stereotypes never appear in the LLM outputs (because of training-data gaps, safety filters, or the six fixed points-of-view), subsequent validation cannot recover them. Consequently the claims of cultural specificity, low English overlap, and country-dependent encoding patterns (Figure 2) are only as strong as the untested completeness of the candidate pool.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces a human-LLM collaborative framework that uses injection-style prompting of LLMs (from six points of view, repeated to saturation) to generate candidate stereotypes, followed by validation and StereoSet-style instantiation by in-culture annotators. It applies the framework to construct EspanStereo, covering race, religion, gender, sexual orientation and age stereotypes for Spain, Mexico, Argentina, Colombia and Nicaragua (538 validated stereotypes, 2 690 triples). Validation rates are high for most country-category cells (Table 7), overlap with StereoSet/CrowS-Pairs is low (Tables 11-12), many stereotypes are absent from existing sociological literature (Table 10), and country-specific cultural grounding is illustrated (Table 6). Shapley-value probing and attention-head pruning on BETO and XLM-R (following Ma et al. 2023b) show country-dependent contribution patterns (Figure 2) and that top-down pruning moves stereotype scores toward 50 while largely preserving language-modeling scores.","tokens_in":24734,"tokens_out":1059,"duration_ms":38849,"significance":"If the results hold, the work supplies both a concrete multi-country Spanish stereotype benchmark and a language-agnostic, lower-cost construction pipeline that can reduce the annotation burden that has limited non-English resources. Explicit strengths include the public MIT-licensed release of EspanStereo, transparent reporting of validation rates, inter-annotator vote ratios (Tables E1-E5), literature and English-dataset overlaps, and faithful reproduction of a published probing protocol that yields the expected top-down versus bottom-up ablation curves. These elements make the contribution immediately usable for culturally grounded evaluation of Spanish-supporting models and for scaling similar resources to other languages.","major_comments":[{"comment":"The central claim that the framework yields high-quality, country-specific stereotypes largely absent from English resources and literature rests on the untested completeness of the LLM-generated candidate pool (Section 3.1). High validation rates (Table 7, mostly >85 %) and low overlaps (Tables 10-12) establish precision of the emitted set after majority-vote filtering (median Likert ≤ 2 discarded), but supply no recall measure against stereotypes actually circulating in each culture. The Nicaragua race cell (36 % validation) shows that annotators successfully reject invalid immigration stereotypes, yet the Limitations section itself notes weaker coverage of less-prominent or emerging stereotypes. Without a complementary human-elicitation baseline or other completeness check for at least one country, the claims of cultural specificity, low English overlap, and country-dependent encoding","section":"Section 3.1, Table 7, Limitations"},{"comment":"The probing and pruning experiments that support the country-variation claim (Section 6, Figure 2) are conducted only on two encoder models (BETO, XLM-R). While the protocol is correctly followed and the ablation curves behave as expected, the paper repeatedly frames its contribution in terms of LLMs more broadly; the absence of even one modern decoder-only Spanish-supporting model leaves open whether the observed country-dependent attention-head patterns generalize beyond the two tested architectures.","section":"Section 6, Figure 2"}],"minor_comments":[{"comment":"Several author names appear with anomalous spacing (e.g., \"V osoughi\", \"Soroush V osoughi\"); these should be corrected throughout.","section":"Title page and references"},{"comment":"Figures F1-F6 and the correlation heatmaps in Figure 2 would benefit from larger fonts and explicit color-bar legends so that positive versus negative Shapley values remain legible in print.","section":"Figures 2, F1-F6"},{"comment":"The term \"injection attack\" is used for the prompting strategy; a more neutral description (e.g., \"adversarial role-play prompting\") would better match the ethics discussion and avoid unnecessary security connotations.","section":"Section 3.1, Ethics Statement"},{"comment":"Table 1-5 captions refer to \"Proportion of Mexican stereotypes shared by other countries\" etc.; a short note clarifying that the percentages are computed after validation would remove ambiguity.","section":"Tables 1-5"}],"recommendation":"major_revision","confidential_remarks":"The completeness/recall gap is the single load-bearing weakness; if the authors can add even a modest human-elicitation comparison for one country or a clearer quantitative bound, the paper becomes a strong accept for a CL venue. The injection-prompt technique is clever but may become brittle as closed-model safety filters evolve; that is a practical rather than scientific concern."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that they actually ship EspanStereo: 538 validated stereotypes across five Spanish-speaking countries, instantiated into 2690 triples, with transparent validation rates, literature-overlap counts, and English-dataset overlap tables. That alone is more useful than most bias papers that stop at “we translated StereoSet.”\n\nWhat is new is the concrete resource plus the generation-validation loop. They prompt GPT-4o (and show Gemini/Llama variants) with injection-style prompts from six points of view, get candidates, then have five in-culture annotators per country rate them on a 5-point Likert and keep the median ≥3. Instantiation follows StereoSet inter-sentence format. Validation is mostly >85 % (Nicaragua race is the clear outlier at 36 %). Overlap with StereoSet/CrowS-Pairs is only 9–13 %, and 77 % of their stereotypes are not in the sparse Spanish-language sociological literature they cite. The probing/pruning experiments on XLM-R and BETO (following Ma et al. 2023b) cleanly show country-dependent attention-head rankings and that top-down pruning moves ss toward 50 with little LMS damage. Dataset is public under MIT.\n\nThe soft spot is exactly the one the stress-test flags: high validation rates measure precision of the emitted set, not recall against the stereotypes that actually circulate. The Limitations section admits weaker coverage of less-prominent or emerging stereotypes, and the Nicaragua race drop shows the LLM can emit irrelevant immigration stereotypes that annotators correctly reject. Six fixed points of view and closed-model injection prompts are brittle. That said, they do not hide it, the rest of the numbers are consistent, and for underrepresented cultures this is still better than pure translation or tiny hand-curated sets. Free parameters (Likert threshold, number of annotators, Shapley samples) are ordinary and reported.\n\nThis is for people building or evaluating multilingual bias benchmarks, especially Spanish or Latin-American work. It is a resource paper with clean methods and honest caveats, not a theory paper. I would send it to peer review; the contribution is real enough to deserve referee time even if reviewers push hard on completeness and prompt brittleness.","headline":"Useful released multi-country Spanish stereotype dataset plus a practical LLM-generation + human-validation pipeline; completeness of the candidate pool is the real soft spot, but the work is honest and the empirical results hold up.","tokens_in":25248,"tokens_out":567,"would_cite":true,"duration_ms":16129,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A human-LLM collaboration framework builds country-specific Spanish stereotype datasets at low cost, exposing large regional differences in how models encode bias.","keywords":["stereotype datasets","human-LLM collaboration","cultural bias","Spanish NLP","cross-cultural evaluation","attention-head pruning","multilingual fairness"],"falsifier":"Recruit independent native speakers from the same five countries who have never seen the LLM outputs and ask them to list stereotypes freely; if large numbers of frequently mentioned stereotypes are absent from EspanStereo, the retrieval step is incomplete.","tokens_in":25278,"feed_emoji":"🌍","tokens_out":841,"duration_ms":13144,"temperature":0.7,"pith_summary":"English-only stereotype benchmarks leave non-English cultures understudied because full manual collection is expensive. This paper shows that large language models can first propose candidate stereotypes for a target country, after which local annotators simply validate and turn them into test examples. The resulting EspanStereo dataset covers five Spanish-speaking countries and contains both familiar stereotypes and many that never appear in English resources or prior sociological lists. When Spanish-supporting models are probed with these examples, the attention heads that drive stereotyping differ markedly by country, and pruning the right heads reduces the bias with little harm to language modeling. The same pipeline is language-agnostic, so it offers a practical route to multilingual, culture-grounded bias evaluation.","feed_headline":"LLMs plus locals build Spanish stereotype data by country","feed_subtitle":"EspanStereo shows models encode bias differently across Spain and Latin America, and pruning can cut it.","key_machinery":"Human-LLM collaborative annotation: an LLM first generates candidate stereotypes under constrained, multi-viewpoint prompts; in-culture annotators then rate prevalence on a Likert scale and write context/stereotype/anti-stereotype triples. The validated set becomes EspanStereo.","core_discovery":"Large language models, when prompted with carefully designed injection attacks, already contain enough cultural knowledge to surface high-quality, country-specific stereotypes; human validation then filters them into a reliable test set. Applied to Spanish, this process yields EspanStereo, whose stereotypes overlap little with English datasets and vary substantially across Spain, Mexico, Argentina, Colombia and Nicaragua. Probing and pruning experiments confirm that the models encode these stereotypes in country-dependent patterns of attention heads.","pith_inferences":["If the method works for Spanish, it should also surface under-documented stereotypes in other mid-resource languages that share training data with major LLMs, such as Portuguese or Turkish.","Country-level differences in attention-head rankings suggest that a single global debiasing recipe may be suboptimal; fine-grained cultural adapters could be more effective.","The high validation rates imply that LLM pre-training already encodes many local stereotypes, so safety filters that simply block stereotype generation may also hide useful cultural knowledge.","Repeating the pipeline every few years with newer models could track how quickly emerging social stereotypes enter model weights."],"forward_implications":["Spanish-supporting models can now be audited and debiased on a per-country basis rather than with translated English lists.","The same generate-then-validate pipeline can be run for any language that has at least modest LLM coverage, producing new culture-specific benchmarks at far lower cost than pure manual collection.","Attention-head pruning guided by EspanStereo-style data becomes a practical mitigation tool for regional stereotypes.","Future multilingual bias suites can be assembled incrementally, country by country, instead of waiting for large-scale sociological surveys."],"fun_headline_variants":["Human-LLM collab builds EspanStereo by Spanish-speaking country","LLMs generate Spanish stereotypes, locals validate country biases","EspanStereo maps region-specific stereotypes across Spanish LLMs","Scalable human-LLM framework yields multi-country Spanish bias data","Spanish LLM stereotypes vary by nation in new collaborative dataset"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The stereotypes an LLM can be induced to emit form a sufficiently complete and unbiased sample of the stereotypes that actually circulate in each culture, so human validation mainly removes noise rather than missing whole classes of local bias.","fun_headline_variants_meta":{"raw":{"variants":["Human-LLM collab builds EspanStereo by Spanish-speaking country","LLMs generate Spanish stereotypes, locals validate country biases","EspanStereo maps region-specific stereotypes across Spanish LLMs","Scalable human-LLM framework yields multi-country Spanish bias data","Spanish LLM stereotypes vary by nation in new collaborative dataset"]},"model":"grok-4.5","effort":"low","cost_usd":0.008298,"raw_usage":{"total_tokens":1949,"prompt_tokens":753,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":82980000,"prompt_tokens_details":{"text_tokens":753,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1110,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":753,"tokens_out":86,"duration_ms":9940,"temperature":1.0,"reasoning_tokens":1110,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T15:50:21.845595+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recruit independent native speakers from the same five countries who have never seen the LLM outputs and ask them to list stereotypes freely; if large numbers of frequently mentioned stereotypes are absent from EspanStereo, the retrieval step is incomplete.","supporting_citations":[],"review_version":1}