{"id":"52b8fb9b-2bdc-44db-bb0b-1e89c8b5c2b7","arxiv_id":"2506.01814","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a 7,200-answer multilingual audit, DeepSeek-R1 produced more Chinese-state propaganda and anti-U.S. sentiment than ChatGPT o3-mini-high, and both biases were strongest in Simplified Chinese.","lead":"This study measured Chinese-state propaganda and anti-U.S. sentiment in answers from DeepSeek-R1 and ChatGPT o3-mini-high across three languages. It found DeepSeek-R1 showed higher rates of both biases, with the strongest effects in Simplified Chinese.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline language-script gradient may be an artifact of GPT-4o's own language-dependent labeling; the human validation does not rule this out because the audit samples were selected from GPT-4o's labels, and the zero anti-US rate for o3-mini-high was never tested against human judgment.","rationale":"The paper's central claim is evaluator-mediated: every headline number in Tables 4 and 5 comes from GPT-4o. The validation design estimates how often GPT-4o agrees with one human annotator on the partition GPT-4o itself created, but it does not establish that GPT-4o's prevalence estimates are calibrated across languages. Since the most distinctive finding is precisely a language gradient (Simplified Chinese > Traditional Chinese > English), the possibility that GPT-4o identifies Chinese-state propaganda more easily in Chinese than in English is the single most load-bearing threat to the argument. The anti-US result for o3-mini-high is also untested against human positives because no positive cases existed in the audit sample. The paper is a plausible empirical measurement with a meaningful design, and the direction of the reported effects is consistent with prior work, but the central claim should not be treated as established until the evaluator confound is addressed with independent, language-aware human labels on model-selected (rather than evaluator-selected) samples. The reader's weakest assumption identified the same evaluator-validity issue, so the conditional verdict remains appropriate; no change is needed.","tokens_in":23532,"tokens_out":3639,"duration_ms":41598,"concrete_test":"Re-label a stratified random sample of 120 answers per model-language pair (720 total) with three independent annotators fluent in Simplified Chinese, Traditional Chinese, and English, blind to model identity, to GPT-4o's labels, and to each other's judgments. Compute per-cell human prevalence with 95% confidence intervals and compare the R1-versus-o3 gaps and the zh-CN-versus-zh-TW-versus-EN gradient. Also run a matched-language probe: take 100 English answers and their Chinese translations (and 100 Chinese answers with their English translations), score both versions with the same GPT-4o rubric, and check whether the same content receives systematically higher propaganda or anti-US labels in Chinese; if it does, the evaluator's language bias is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All headline proportions in Tables 4 and 5 are GPT-4o labels, so the central claim is only as strong as the evaluator's cross-language calibration. The Statistical Agreement section validates GPT-4o against a single human annotator on audit sets that were themselves selected from GPT-4o's positive/negative labels (30 positive vs 30 negative per model-dimension whenever possible). That design estimates agreement conditional on GPT-4o's partition, but it cannot detect a systematic language-specific threshold: if GPT-4o labels PRC-sounding terms such as \"合作共赢\" or \"中国式现代化\" more readily in Chinese than in English, the reported zh-CN > zh-TW > EN gradient is an artifact of the evaluator rather than the models. The Anti-US dimension is even weaker: o3-mini-high has zero positives in all three languages, and the audit set contains no positive cases (Table 3), so precision, recall, F1, and kappa are undefined and there is no human evidence that GPT-4o can recognize anti-US content at all. The paper also does not state whether the single human annotator was blind to model identity or to GPT-4o labels, and no inter-annotator reliability is reported. Finally, the headline differences are reported without confidence intervals or significance tests; for example, in the zh-CN propaganda comparison, 6.83% versus 4.83% with n=1,200 per cell is not self-evidently beyond sampling noise, and the model-level English difference is 0.08% versus 0.17%.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a cross-lingual audit of two large language models, DeepSeek-R1 and ChatGPT o3-mini-high, for Chinese state propaganda and anti-U.S. sentiment. A corpus of 1,200 de-contextualized questions derived from Chinese-language news was posed in Simplified Chinese, Traditional Chinese, and English, yielding 7,200 model answers. Answers were scored by a rubric-guided GPT-4o evaluator on two binary dimensions, with a small human audit (one annotator) reported as validation. The central claims are that DeepSeek-R1 exhibits consistently higher propaganda and anti-U.S. bias than ChatGPT o3-mini-high, that Simplified Chinese queries elicit the highest bias rates, followed by Traditional Chinese, with English nearly bias-free, and that DeepSeek-R1 functions as an 'invisible loudspeaker' by amplifying PRC-aligned terms, including occasional code-switching from Traditional to Simplified Chinese.","tokens_in":23871,"tokens_out":2898,"duration_ms":30905,"significance":"If the findings are robust, the paper addresses a timely and important question: whether geographically or politically aligned LLMs embed state-aligned narratives across languages. The dataset construction (1,200 decontextualized questions, three parallel languages, 7,200 answers) and the attempt to use a rubric-guided LLM judge with human adjudication are useful contributions to LLM auditing methodology. The paper also makes a concrete empirical contribution by reporting per-topic and per-language prevalence counts. However, the current evidence base is insufficient to support the paper's headline claims as stated, because the headline proportions all depend on a single LLM evaluator whose cross-language calibration is not established, the human audit cannot rule out language-specific evaluator artifacts, the anti-U.S. dimension is nearly unvalidated, and no uncertainty quantification is provided for any of the key comparisons.","major_comments":[{"comment":"No confidence intervals, hypothesis tests, or multiple-comparison corrections are reported for any of the headline proportions. The abstract and text use the word 'significant' (e.g., 'significantly higher,' 'statistically significant consistency'), but with n = 1,200 per cell the zh-CN propaganda difference of 6.83% versus 4.83% is not self-evidently beyond sampling noise, and the English difference (0.08% vs 0.17%) is even more fragile. The authors should report exact binomial or bootstrap confidence intervals for every proportion, perform a formal test (e.g., chi-square or Fisher exact) for each model-language comparison, and correct for the multiple comparisons across the six model-language pairs and two dimensions. Without this, the central claim of systematic, language-dependent bias is not statistically supported.","section":"Results and Discussions (Tables 4 and 5)"},{"comment":"The human audit design cannot validate the cross-language gradient because the audit samples were selected from GPT-4o's own positive/negative labels (30 positive vs 30 negative whenever possible). This estimates agreement conditional on GPT-4o's partition, but it cannot detect a systematic language-specific threshold; for example, if GPT-4o labels PRC-sounding terms more readily in Simplified Chinese than in English, the reported zh-CN > zh-TW > EN gradient could be entirely an artifact of the evaluator. The audit sets should be stratified jointly by model, language, and GPT-4o label, and the single human annotator should be blinded to model identity, language, and GPT-4o labels. The paper should also report whether blinding was used and provide inter-annotator reliability on an overlapping subset coded by at least two humans.","section":"Statistical Agreement Between LLM and Human Judgments (Table 3)"},{"comment":"The anti-U.S. dimension is essentially unvalidated for the strongest quantitative claim in the paper. The o3-mini-high audit set contains zero positive cases (Y = 0, N = 30), so precision, recall, F1, and Cohen's kappa are undefined, and there is no human evidence that GPT-4o can recognize anti-U.S. content at all. The claim that o3-mini-high exhibits zero anti-U.S. bias in all three languages (Table 5) therefore rests entirely on a single LLM evaluator with no sensitivity check. The authors should validate the evaluator on a set of known positive examples (e.g., seeded anti-U.S. statements) across all three languages, and they should present the human audit separately for the positive and negative strata.","section":"Statistical Agreement Between LLM and Human Judgments (Table 3) and Results (Table 5)"},{"comment":"There is a partial ecological circularity in the evaluation design: the same developer ecosystem (OpenAI's o3-mini generates the questions, GPT-4o scores the answers) judges ChatGPT o3-mini-high's outputs. The paper cites LLM-as-judge limitations but does not address the specific risk that OpenAI models share a common, unmeasured labeling bias, which could inflate or deflate the comparison with DeepSeek-R1 in unknown directions. A robustness check with an independent evaluator (e.g., an open-source judge or a second commercial model) or a random sample of model answers (rather than GPT-4o-selected positives/negatives) adjudicated by humans would materially strengthen the central claim.","section":"Question Generation and Bias Evaluation Pipeline (Tables 8, 9, 12)"}],"minor_comments":[{"comment":"The abstract states results are 'significant' without any statistical support; please soften to 'empirically observed' or add the supporting tests.","section":"Abstract"},{"comment":"The text says 'ChatGPT-4o' in one place; the evaluator is GPT-4o. Please use consistent naming.","section":"Statistical Agreement Between LLM and Human Judgments"},{"comment":"The table captions should explicitly state that all numbers are GPT-4o labels; the current captions imply ground truth.","section":"Tables 4 and 5"},{"comment":"Cohen's kappa values are reported without confidence intervals; given the small audit samples (n = 60 per cell, n = 30 for one cell), CIs would clarify the precision of these agreement estimates.","section":"Statistical Agreement Between LLM and Human Judgments (Table 3)"},{"comment":"The column 'Model Bias > Query Bias?' reports 'Yes'/'No' without any defined threshold or test; please clarify the operationalization of this comparison.","section":"Appendix Table 20"},{"comment":"The term 'invisible loudspeaker effect' is used as an established finding, but the paper only provides anecdotal examples and unvalidated counts; I recommend hedging this claim until the evaluator-validity and statistical issues are resolved.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"This paper addresses a politically loaded and timely topic, and the empirical counts are suggestive. However, the central claims currently hinge on a single LLM evaluator whose cross-language calibration is unproven, an anti-U.S. dimension that is essentially unvalidated, and a complete absence of uncertainty quantification. These are fixable within the manuscript's scope, hence major revision rather than rejection. I would also note that the paper is likely to receive substantial public attention; the authors should be encouraged to make the dataset and evaluation code publicly available to facilitate independent verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely new head-to-head measurement of a PRC-aligned and a non-PRC model across three languages, and the directional finding—DeepSeek-R1 carries more Chinese-state propaganda and anti-US framing, especially in Simplified Chinese—is plausible. But the paper hasn't demonstrated that the headline differences are significant, and the evaluator design leaves open the possibility that part of the language gradient is an artifact of GPT-4o's labeling.\n\nThe real contribution is the 7,200-answer corpus of decontextualized, topic-stratified questions, with prompts in the appendix and detailed examples in Tables 15-23. The term-injection analysis in Table 20 is concrete: R1 adds 13 PRC slogans in zh-CN vs. two for o3-mini-high, and the 'invisible loudspeaker' framing is earned by that observation rather than just asserted. The script-switching behavior (R1 responding to zh-TW in zh-CN) is also a useful detail, and the paper itself flags that this may have inflated its zh-TW bias counts.\n\nThe soft spots are in the statistical and evaluator layers. The paper reports counts and percentages but no confidence intervals or significance tests; 6.83% vs 4.83% with n=1,200 per cell is not obviously beyond sampling noise, and the English raw counts (1 vs 2) point the other way. The anti-US dimension is weaker: o3-mini-high has zero positives in all languages, and the human audit included no positive o3 anti-US cases, so we have no evidence GPT-4o can recognize anti-US content in o3's output. The audit sets were also selected from GPT-4o's own labels, which cannot detect a language-specific threshold shift. And the labeling thresholds differ—propaganda triggers at any dimension >=1, anti-US only at score >=2—so the two rates aren't directly comparable.\n\nNone of this kills the paper. The core measurement is new and policy-relevant, and the authors are transparent about what they did. But the abstract's 'significant' should read as 'observed in this sample,' and the language-gradient claim needs a calibration study before it becomes load-bearing.\n\nI'd send this to peer review. It's a serious piece of empirical auditing with real data, and a good reviewer can push for inferential statistics, blind multiple human annotators, and a cross-language evaluator calibration. I wouldn't cite it yet, but I'd bring it to a reading group.","headline":"A fresh and transparent cross-lingual measurement of PRC-aligned vs. non-PRC model bias, but the headline claims need significance tests and evaluator calibration before they hold up.","tokens_in":24421,"tokens_out":3398,"would_cite":false,"duration_ms":33886,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DeepSeek-R1 embeds Chinese propaganda and anti-U.S. framing far more than ChatGPT, with the bias strongest in Simplified Chinese and nearly absent in English.","keywords":["LLM political bias","DeepSeek-R1","Chinese-state propaganda","anti-U.S. sentiment","cross-lingual evaluation","invisible loudspeaker effect","LLM-as-a-judge","decontextualized questions"],"falsifier":"Show the same 7,200 answers to bilingual human annotators who are fluent in Simplified Chinese, Traditional Chinese, and English and are blind to model identity and the hypothesis, and have them apply the paper's own rubric. If the propaganda and anti-U.S. gaps between DeepSeek-R1 and ChatGPT shrink or vanish under human-only rating, the reported language-dependent bias is an artifact of the GPT-4o evaluator; if they persist, the effect is a genuine property of the models.","tokens_in":23348,"feed_emoji":"📢","tokens_out":12542,"duration_ms":112869,"temperature":0.7,"pith_summary":"The paper argues that a state-aligned large language model can function as an \"invisible loudspeaker\" that weaves official Chinese-state narratives and anti-U.S. framing into seemingly ordinary answers. To test this, the authors built 1,200 de-contextualized reasoning questions from Chinese-language news, posed them to DeepSeek-R1 and ChatGPT o3-mini-high in Simplified Chinese, Traditional Chinese, and English, and had a rubric-guided GPT-4o evaluator label all 7,200 answers. They report that DeepSeek-R1 was labeled propaganda-bearing in 6.83 percent of Simplified Chinese answers versus 4.83 percent for ChatGPT, that anti-U.S. labels appeared at 5.00 percent versus zero, and that both bias types shrank or disappeared in English. A sympathetic reader should care because the effect persists when questions contain no political trigger words, which means the bias lives inside the model rather than being summoned by keyword bait, and it leaks into cultural and travel topics where casual users are not on guard.","feed_headline":"DeepSeek-R1 embeds Chinese propaganda far more than ChatGPT does","feed_subtitle":"In Simplified Chinese the bias is strongest; in English it nearly vanishes—and it leaks into everyday topics.","key_machinery":"The load-bearing apparatus is a three-part design built around what the paper calls the \"invisible loudspeaker\" effect. First, a de-contextualized corpus: 1,200 open-ended reasoning questions generated from Chinese-language news headlines and summaries by abstracting away concrete names, places, and dates, so no political trigger words remain. Second, parallel translation of every question into Simplified Chinese, Traditional Chinese, and English, creating linguistically matched prompts that isolate language effects from content effects. Third, a hybrid evaluator in which rubric-guided GPT-4o scores each of the 7,200 answers on five propaganda dimensions (ideological narrative alignment, information selection and sourcing, emotional mobilization and symbol use, handling dissent, and formulaic language) plus one anti-U.S. dimension (negative framing and case usage), with a single human annotator's judgments used to measure agreement on balanced audit sets. The central identity the argument pivots on is the language-gated amplification of state-aligned rhetoric: the same model, asked the same question, narrates differently depending on whether it is addressed in Simplified Chinese, Traditional Chinese, or English.","core_discovery":"The paper's central claim is that DeepSeek-R1, a model trained and aligned inside mainland China, exhibits substantially higher rates of Chinese-state propaganda and anti-U.S. sentiment than ChatGPT o3-mini-high when both answer the same questions, and that the difference is driven by the language of the prompt. In the propaganda dimension, GPT-4o labels 82 of 1,200 Simplified Chinese DeepSeek answers (6.83 percent), 29 of 1,200 Traditional Chinese answers (2.42 percent), and 1 of 1,200 English answers (0.08 percent), against 58, 19, and 2 for ChatGPT. In the anti-U.S. dimension, DeepSeek-R1 is labeled in 60 Simplified Chinese answers (5.00 percent), 29 Traditional Chinese answers (2.42 percent), and 5 English answers (0.42 percent), while ChatGPT receives zero anti-U.S. labels in every language. The paper reads this pattern as evidence of an internal inclination toward state-aligned framing: DeepSeek-R1 injects PRC-keywords beyond those present in the queries, sometimes answers Traditional Chinese questions in Simplified Chinese, and produces the bias even on questions stripped of all names, dates, and places.","pith_inferences":["A test the paper does not run: replace GPT-4o with an evaluator from a different alignment regime, or use human-only rating on the same 7,200 answers. If the language-gradient collapses, part of the reported effect is the evaluator's own language-conditioned recognition of Chinese slogans rather than a property of the two models.","One causal reading the paper leaves open: the de-contextualized design implies the bias is an internal prior of DeepSeek-R1's reasoning process, so an intervention that forces English internal reasoning on Chinese prompts could reveal exactly where in the chain of thought the state-aligned framing enters.","An extrapolation from the paper's numbers: given earlier findings the authors cite that a few conversational turns with an LLM can shift voter preferences by several percentage points, even a five-to-seven percent bias rate in Chinese-language answers could compound into measurable attitude changes over repeated daily use, an effect the paper documents but does not quantify.","A training-data interpretation the paper only gestures at: the Simplified-versus-Traditional gradient may partly reflect that Simplified Chinese training text is dominated by mainland sources, whereas Traditional Chinese text is more likely to come from Taiwan, Hong Kong, and overseas communities where PRC framing is less pervasive."],"forward_implications":["Users who query DeepSeek-R1 in Simplified Chinese receive propaganda-labeled answers at a rate of roughly one in fifteen, about 1.5 times the rate ChatGPT produces on identical questions, and the anti-U.S. gap is even wider.","Because the bias appears on de-contextualized questions, removing political keywords or trigger terms is unlikely to neutralize it; a user cannot reliably filter the effect by avoiding sensitive topics.","The bias is not confined to geopolitics or domestic politics but also surfaces in culture, public and social issues, travel, and tourism, which means non-political reading contexts still carry state-aligned framing.","DeepSeek-R1 sometimes replies to Traditional Chinese queries in Simplified Chinese, which the paper says can broaden exposure to potentially biased content for Traditional Chinese users.","English-language interaction with DeepSeek-R1 shows almost none of either bias, so the risk profile is strongly language-dependent rather than uniformly present across the model.","On every question and in every language tested, ChatGPT o3-mini-high produced zero answers labeled anti-U.S., indicating that negative U.S. framing was unique to the PRC-aligned model in this dataset."],"supporting_citations":[{"why":"Supplies the eight recurrent PRC anti-American messaging frames that directly inform the annotation rubric for both bias dimensions.","marker":"Carothers (2024)"},{"why":"Provides the dataset of 4,100 state-sponsored propaganda technique instances used for keyword seeding in the evaluation prompts.","marker":"Chang et al. (2021)"},{"why":"Contributes the \"cultural lensing\" method of rewriting prompts to remove locale-specific proper nouns, on which the de-contextualized question design is based.","marker":"Hsieh et al. (2024)"},{"why":"Supplies the content-versus-style distinction that the five-dimension propaganda rubric operationalizes.","marker":"Bang et al. (2024)"},{"why":"Establishes the real-world attitude-shift motivation and the thematic aggregation approach adopted for topic-level analysis.","marker":"Potter et al. (2024)"},{"why":"Justifies the matched bilingual cross-model comparison strategy used to isolate language effects from topic variation.","marker":"Zhou and Zhang (2024)"},{"why":"Validates the LLM-as-a-judge paradigm that the GPT-4o scoring pipeline relies on, including its agreement rates and known evaluator biases.","marker":"Zheng et al. (2023)"},{"why":"Provides the chain-of-thought, form-filling scoring approach that the paper adapts to increase resolution in human-LLM disagreement cases.","marker":"Yang et al. (2023)"}],"fun_headline_variants":["DeepSeek-R1 shows far more Chinese propaganda bias than ChatGPT","Propaganda bias in DeepSeek-R1 peaks in Simplified Chinese, barely in English","DeepSeek-R1's propaganda bias dwarfs ChatGPT's, and it's language-dependent","DeepSeek-R1's propaganda bias leaks beyond politics into everyday topics","Language flips DeepSeek-R1's bias: heavy in Chinese, almost none in English"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison stands or falls on the assumption that GPT-4o's rubric-based labels are a valid and language-fair measure of \"Chinese-state propaganda\" and \"anti-U.S. sentiment,\" a premise the paper checks against only one human annotator on balanced samples of thirty positives and thirty negatives per condition.","fun_headline_variants_meta":{"raw":{"variants":["DeepSeek-R1 shows far more Chinese propaganda bias than ChatGPT","Propaganda bias in DeepSeek-R1 peaks in Simplified Chinese, barely in English","DeepSeek-R1's propaganda bias dwarfs ChatGPT's, and it's language-dependent","DeepSeek-R1's propaganda bias leaks beyond politics into everyday topics","Language flips DeepSeek-R1's bias: heavy in Chinese, almost none in English"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000602,"raw_usage":{"total_tokens":2896,"prompt_tokens":1117,"completion_tokens":1779,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":1671}},"tokens_in":733,"tokens_out":1779,"duration_ms":13136,"temperature":1.0,"reasoning_tokens":1671,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:32:11.430031+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Show the same 7,200 answers to bilingual human annotators who are fluent in Simplified Chinese, Traditional Chinese, and English and are blind to model identity and the hypothesis, and have them apply the paper's own rubric. If the propaganda and anti-U.S. gaps between DeepSeek-R1 and ChatGPT shrink or vanish under human-only rating, the reported language-dependent bias is an artifact of the GPT-4o evaluator; if they persist, the effect is a genuine property of the models.","supporting_citations":[],"review_version":1}