{"id":"d4c5e20d-86fb-40d7-9e68-8d7ea16ba2d9","arxiv_id":"2605.30214","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"GRUFF dataset shows LLMs agree well with masculine and feminine German pronouns but fail on neopronouns and distractors, with occupational stereotypes poorly correlated across cases.","lead":"This paper introduces the GRUFF dataset to test how LLMs handle pronoun reuse and grammatical gender agreement in German. A smart generalist might read it to learn about biases and reasoning limits when extending English-focused AI tests to languages with richer grammar.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest_assumption correctly flags the dataset-construction step as the least-secured premise. With full text now available, no additional load-bearing gap (e.g., hidden assumption in an equation or unstated confound in the metric) appears in the argument structure itself. The concern therefore remains exactly where the reader located it; no adjustment to UNVERDICTED is warranted.","tokens_in":1732,"tokens_out":268,"duration_ms":15833,"concrete_test":"Re-run the main pronoun-fidelity and distractor-robustness analyses on the released GRUFF test split after randomly permuting occupational stereotype labels within each grammatical case; if the reported masculine/feminine vs. neopronoun gap or the encoder-only robustness advantage disappears, the isolation assumption is violated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims rest on GRUFF items isolating pronoun fidelity via controlled distractors and occupational stereotypes. The abstract and high-level description provide no internal inconsistency or circularity once the dataset is taken as given; the reported differences between pronoun classes and between model families follow directly from the stated experimental contrast. No parameter-free derivation or machine-checked element is claimed, but the design is falsifiable via the released data.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the GRUFF dataset to measure pronoun fidelity in German LLMs across four gender agreement systems in nouns and four pronoun sets. It claims LLMs exhibit strong grammatical agreement for masculine and feminine entities without explicit context but not for neopronouns xier and en; models are generally not robust to distractors, though encoder-only models show greater robustness in German than English; and occupational stereotypes are poorly correlated across grammatical cases and across most models except those with closely related architectures. All code and data are released.","tokens_in":1771,"tokens_out":365,"duration_ms":36137,"significance":"If the empirical results hold, this extends English-centric pronoun fidelity and bias research to a language with rich grammatical gender, providing evidence on how grammatical features affect referential reasoning and neopronoun handling in LLMs. The release of code and data is a clear strength supporting reproducibility and further work on gender-inclusive language. The model-family contrast offers a falsifiable benchmark for multilingual evaluation.","major_comments":[],"minor_comments":[{"comment":"Abstract: the four pronoun sets are referenced but not named; listing them (or adding a short table) would improve immediate clarity for readers.","section":null},{"comment":"Dataset description: include one or two concrete GRUFF example items in the main text (rather than only in supplementary material) to illustrate how distractors and occupational stereotypes are instantiated.","section":null},{"comment":"Results: the claim of 'poorly correlated' stereotypes across cases would be strengthened by reporting the exact correlation coefficients and any statistical tests used.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The reader's low soundness score stems from an abstract-only review; the full manuscript appears to contain the necessary experimental details and released artifacts. No concerns about citation patterns or scope."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary of the GRUFF dataset and its contributions to multilingual pronoun fidelity research, as well as for recommending minor revision. No specific major comments were raised in the report.","responses":[],"tokens_in":1240,"tokens_out":59,"duration_ms":14211,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core move is releasing GRUFF, a large dataset that applies pronoun fidelity testing to German across four noun gender agreement systems and four pronoun sets, including the neopronouns xier and en. That extension from English-only work is the actual addition.\n\nIt does a few things cleanly. The setup lets them compare grammatical agreement strength for masculine and feminine referents against neopronouns, and it checks robustness when distractors are inserted. The finding that encoder-only models hold up better in German than prior English results is worth noting, since German has richer agreement. Releasing code and data is the right step for follow-up work.\n\nThe soft spots are in the experimental controls. The abstract states that items isolate fidelity from other discourse factors, but without the actual item templates or distractor selection criteria it is hard to judge whether occupational stereotypes and case variations are representative or just convenient. The claim that stereotypes correlate poorly across grammatical cases is reported, yet it could shift with different occupation lists or model families. These are not fatal, but they mean the patterns are tied to the specific dataset choices.\n\nThis paper is for people already working on multilingual LLM evaluation and gender-inclusive language tasks. The dataset itself is the part that could get used. It is coherent enough on its own terms to go to referees rather than desk reject, though any review would need to press on the item construction details and whether the German-English model comparison controls for architecture differences.","headline":"GRUFF gives a usable new dataset for German pronoun tests but the robustness and stereotype claims rest on unshown details of item construction.","tokens_in":2230,"tokens_out":367,"would_cite":false,"duration_ms":14403,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs maintain strong grammatical agreement for masculine and feminine pronouns in German but fail with neopronouns xier and en, and most models lose robustness when distractors appear.","keywords":["pronoun fidelity","German language models","grammatical gender","neopronouns","referential reasoning","gender bias","LLM evaluation"],"falsifier":"A controlled test in which current LLMs reuse xier and en at rates matching er and sie in sentences without distractors would falsify the claim of differential agreement for neopronouns.","tokens_in":2626,"feed_emoji":"🇩🇪","tokens_out":772,"duration_ms":26217,"temperature":0.7,"pith_summary":"The paper introduces the GRUFF dataset to test pronoun fidelity, which measures whether models reuse a previously specified pronoun for an entity even when other entities intervene. It finds that models follow masculine and feminine pronouns according to German's grammatical gender rules without extra context, yet do not do so for the neopronouns xier and en. Models generally lose track when distractors are added, though encoder-only models hold up better in German than they do in English. Occupational stereotypes show little consistency across grammatical cases and across different models. These results matter because German requires more gender agreement than English, so English-only studies of bias and reference may miss language-specific patterns in reasoning and inclusion.","feed_headline":"German LLMs keep masculine and feminine pronouns but drop neopronouns","feed_subtitle":"Encoder-only models resist distractors better than in English thanks to grammatical gender, yet stereotypes fail to align across cases.","key_machinery":"The GRUFF dataset, which isolates the task of correctly reusing a specified pronoun for a discourse entity despite intervening distractors across different noun gender classes and pronoun types including neopronouns.","core_discovery":"We present GRUFF, a large-scale dataset covering four gender agreement systems in nouns and four pronoun sets to measure pronoun fidelity in German. Using this dataset we show that LLMs exhibit strong grammatical agreement for masculine and feminine entities in the absence of explicit context, but not for neopronouns xier and en. Models are generally not robust to distractors, but encoder-only models are more robust in German than in English, reflecting the importance of grammatical gender. Occupational stereotypes in this context are poorly correlated across grammatical cases and across most models, except ones with closely related architectures.","pith_inferences":["The greater robustness of encoder-only models in German suggests that explicit grammatical gender marking during training improves handling of referential distractors.","The dataset could support development of fine-tuning methods that raise neopronoun fidelity while preserving agreement on traditional pronouns.","Because stereotypes correlate poorly across cases, bias audits in German should test multiple grammatical forms rather than nominative alone.","Similar fidelity gaps may appear in other languages with grammatical gender such as French or Spanish, warranting parallel datasets."],"forward_implications":["Models will correctly reuse masculine and feminine pronouns in German text without distractors.","Neopronouns xier and en will not be maintained reliably by current LLMs in German discourse.","Encoder-only models will show greater robustness to distractors than decoder-only models when processing German pronouns.","Occupational stereotypes will show low correlation across different grammatical cases for most models.","Bias patterns measured in German will differ from those in English because of richer grammatical gender."],"fun_headline_variants":["LLMs agree on German masc fem pronouns but not neopronouns xier en","Encoder only models resist distractors better in German than English","German occupational stereotypes misalign across grammatical cases","GRUFF dataset covers four noun genders and four pronoun sets"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The GRUFF dataset items isolate pronoun fidelity from other discourse factors and that the chosen occupational stereotypes and distractors are representative of real usage patterns in German.","fun_headline_variants_meta":{"raw":{"variants":["LLMs agree on German masc fem pronouns but not neopronouns xier en","Encoder only models resist distractors better in German than English","German occupational stereotypes misalign across grammatical cases","GRUFF dataset covers four noun genders and four pronoun sets"]},"model":"grok-4.3","cost_usd":0.005345,"raw_usage":{"total_tokens":2603,"prompt_tokens":715,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":53449500,"prompt_tokens_details":{"text_tokens":715,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1820,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":715,"tokens_out":68,"duration_ms":13939,"temperature":1.0,"reasoning_tokens":1820,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:37:05.609193+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which current LLMs reuse xier and en at rates matching er and sie in sentences without distractors would falsify the claim of differential agreement for neopronouns.","supporting_citations":[],"review_version":1}