REVIEW 3 major objections 5 minor 62 references
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Across four Qwen generations, censorship migrates from visible refusal to invisible reframing.
desk verdict The Chinese-language framing gate is a solid, robust finding, but the headline refusal-to-reframing trend rests on an unverified assumption that judge sensitivity is stable across generations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the six-dimension audit rubric in which D1 explicit refusal and D4 state-aligned framing are scored independently, so a model's visible refusal and its invisible reframing can move in opposite directions. State-aligned framing is defined through a three-axis discourse taxonomy (overt endorsement, substitution/euphemism, deflection) and scored per trial by two frontier LLM judges whose labels are validated against three human raters; a short derivation shows the judges' imperfect recall attenuates absolute rates but leaves odds ratios approximately unbiased. The visual-abstraction probe (original, crop, grayscale, edge, binary, low-pass, silhouette) supplies the evidence that subject recognition, not pixel detail, triggers framing.
What would settle it
Take the 200-trial validation sample, or a larger stratified sample across the four Qwen generations and both languages, and have human raters label state-aligned framing directly; if the monotonic rise from 4.8% to 33.0% reverses or disappears under human labels while judge labels keep the trend, the refusal-to-reframing claim is an artifact of judge bias rather than a property of the models.
Extended reading notes
Core claim
The central discovery, as the authors state it, is that refusal and reframing move in opposite directions across model generations: state-aligned framing rises monotonically (4.8%, 7.2%, 21.4%, 33.0%) while explicit refusal declines overall (9.0%, 5.4%, 2.0%, 4.6%), with the two crossing at the second generation. Because refusal and framing are measured as independent dimensions, the paper can show that a model can stop refusing while still reframing; the new form is a fluent, substitution-dominated description that advances the official narrative, such as calling a detention facility a vocational training center or a tank column a military parade. The paper also finds this framing is gated by recognition of the depicted subject rather than pixel detail, persists across prompt-language and abstraction conditions, and is concentrated on politically sensitive content rather than being generic.
Load-bearing premise
The whole trend rests on the assumption that the LLM judges' tendency to miss state-aligned framing—about half of the cases humans flag, validated on only 200 trials—is roughly constant across prompt languages and the four model generations.
Editorial extensions
If this is right
- Refusal rate alone is an incomplete safety metric, because a model can appear more open while still steering users toward a distorted account.
- Users who prompt in Chinese face roughly three times the odds of state-aligned framing, within every model audited.
- Text-only queries about sensitive subjects trigger far more reframing than image queries (36.5% vs. 9.8%), so framing is driven mainly by the model's textual prior.
- The newest Qwen generation is the most likely to reframe silently, and the pattern can be inherited by downstream fine-tuned systems built on these open-weight checkpoints.
- Keyword and refusal detectors miss most reframing: the union of three lexical and length detectors leaves 83.5% of state-aligned framing undetected.
Reading between the lines
- If the trend continues, future alignments may eliminate refusal entirely, leaving fluent reframing as the only censorship channel; audits should therefore monitor framing rates, not refusal rates, as the primary signal.
- The language gate suggests the behavior is triggered by semantic and political context rather than content policy; one could test whether other aligned models show similar language-modulated framing.
- Because reframing is invisible to the user, interface-level disclosure, such as provenance notices or side-by-side sourcing, may be the only practical mitigation; this is a testable design question outside the paper's scope.
- The origin effect being concentrated on sensitive content suggests that downstream developers who fine-tune China-origin open-weight models may inherit governance-shaped alignment without knowing it; checking framing on sensitive imagery before deployment would quantify that risk.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a large-scale audit of political censorship in nine open-weight vision-language models (seven China-origin, two non-China) across 200 sensitive images, four elicitation paradigms, two prompt languages, and three random seeds, yielding 21,708 trials. Each response is labeled on six dimensions by two LLM judges (Claude Opus 4.7 and GPT-5.5), validated against three human experts on a 200-trial sample. The main findings are: (i) Chinese-language prompting roughly triples the odds of state-aligned framing within every model; (ii) China-origin models reframe more than non-China models, with a judge-dependent magnitude; (iii) framing is strongest in text-only political commentary and persists even at silhouette for politically iconic images; and (iv) across four Qwen generations, state-aligned framing rises monotonically (4.8%, 7.2%, 21.4%, 33.0%) while explicit refusal declines overall (9.0%, 5.4%, 2.0%, 4.6%), which the authors interpret as a migration from visible refusal to invisible reframing. The paper contributes a six-dimension measurement framework, a full-corpus dual-judge audit, a human-validation protocol, and an attenuation argument for odds-ratio robustness (Appendix G).
Significance. If the empirical claims hold, this is a significant contribution: the first large-scale audit of political reframing in VLMs, with a design that separates refusal from framing and demonstrates a cross-generational form shift. The within-model Chinese-language gate (OR 3.67, p<10^-78, holding in all nine models under two independent judges and three seeds) is extremely robust and is the strongest single result. The paper also makes a strong reproducibility commitment, releasing all labels, rationales, and verbatim quotes, and it reports cross-seed stability and a selectivity design that helps separate governance-shaped effects from generic capability differences. However, the flagship refusal-to-reframing finding rests on the D4 label, whose human inter-rater agreement is Gwet AC1=0.39 with judge recall of only 44-46%, and the paper's own robustness derivation assumes a shared judge sensitivity across compared cells. The central claim is therefore defensible but not yet established at the level asserted in the abstract and introduction.
major comments (3)
- [§5.6 and Appendix G (Eq. 4-5)] The central claim that newer Qwen models are 'censored differently, trading a behavior users can detect for one they cannot' assumes that the rise in measured D4 across generations reflects a rise in true state-aligned framing rather than a rise in judge sensitivity. The 200-trial validation sample is too small to estimate per-generation recall, and the Appendix G cancellation argument requires the same sensitivity s in the cells being compared. The paper's own Appendix J shows that the Opus judge's miss rate is population-dependent (6/6 vs. 17/35, Fisher p=0.027), and Appendix H.3 reports that among D4+ responses overt endorsement rises from 47% to 75% while median response length roughly quadruples; both are plausible mechanisms for increasing judge detectability. Without per-generation validation or an explicit sensitivity analysis that bounds the trend, the 4.8% to 33.0% monotonic increase is not identified as a true behavioral change.
- [§4.3 and Appendix D] The D4 outcome itself has only Gwet AC1=0.39 among the three human experts, and both LLM judges recall only 44-46% of human-identified positives. The authors interpret the low recall as making reported rates conservative lower bounds, but that interpretation is valid only for absolute prevalence under near-perfect specificity; it does not make between-group odds ratios conservative unless the false-negative rate is constant across the groups. Since D4 is the primary outcome for all four findings, the manuscript should report validation stratified by the factors that drive the main comparisons (at minimum model generation and prompt language), or provide a formal bias analysis that varies sensitivity parametrically and shows the reported effects survive.
- [Appendix G, Eq. (5)] The 'rare-outcome regime' approximation (1 - s*pi_j ≈ 1) is applied to cells with rates as high as 36.5% (comment-text) and 33.0% (Qwen3.5-9B). At these rates the approximation error is not negligible, and the statement that odds ratios are 'approximately unbiased' under judge attenuation should be replaced by the exact expression or by a numerical check across the observed rate range. This is a technical but load-bearing step in the robustness argument for all of the paper's odds-ratio effects.
minor comments (5)
- [§5.7 and Table A11] The statement that the median cross-seed SD of the state-aligned rate is ≤2.4 percentage points across the 72 cells does not reconcile with Table A11, where GLM-4.6V-Flash has a per-model median SD of 2.80 pp; please clarify which summary statistic is being reported.
- [Appendix P] Using Claude Opus 4.7 as both the primary judge and the 'uncensored reference VLM' creates an appearance of circularity even though the appendix explicitly says the reference is not a gold standard; an independent non-China reference model would strengthen this illustrative comparison.
- [§5.6, footnote 3] Describing the exact rank-trend test p=0.083 as 'marginally non-significant' overstates the evidence; with only four generations this p-value is far from conventional significance, and the paper's descriptive interpretation of the form shift is the appropriate framing.
- [§4.3] Calling D4's Gwet AC1=0.39 'substantial' is inconsistent with standard interpretations of this coefficient; 'moderate' or 'low' would be more accurate and would better signal the interpretive difficulty of the construct.
- [Appendix O] The deep-stripped response-length metric depends on a fairly involved cleaning protocol; Table A17 reassuringly shows the fourfold length growth holds under all three accountings, but this should be stated in the main text where the length growth is first reported.
Circularity Check
No significant circularity; the audit is anchored to external human validation and a pre-registered rubric.
full rationale
The paper's central measurements are anchored to external evidence rather than to its own outputs or definitions. The D4 state-aligned framing construct is operationalized through a pre-registered rubric with per-entry expected-facts annotations, and the LLM-judge labels are validated against a three-expert human majority on a 200-trial sample; the reported low recall (Opus 0.44, GPT-5.5 0.46) and the Gwet AC1=0.39 on D4 are acknowledged limitations, not circular moves. The generational refusal-to-reframing trend is a descriptive summary of the independently labeled corpus, not a fitted parameter renamed as a prediction. Equation (1) is a logistic regression fitted to the same data, and the claimed language/origin independence is tested with an interaction term rather than imposed. Appendix G gives an explicit attenuation argument showing that odds ratios approximately cancel the judge sensitivity factor; it assumes a shared sensitivity across compared cells, which is a falsifiable correctness assumption that the authors themselves flag as violated in the Opus origin comparison (Appendix J), not a circular derivation. The only self-referential element is the use of Claude Opus 4.7 as both the primary judge and the illustrative 'uncensored reference' model in Appendix P; that appendix is qualitative, is not load-bearing for any quantitative claim, and is supplemented by human validation. No load-bearing step reduces to its own inputs by construction, and no central claim depends on a self-citation chain.
Assumptions & free parameters
free parameters (1)
- Inference sampling defaults (temperature, top_p, top_k) =
0.7, 0.8, 20
assumptions (5)
- domain assumption The expected-facts annotations on the 200 benchmark images correctly represent the documented historical and political record for each depicted subject.
- domain assumption LLM-judge labels for state-aligned framing on the full 21,708-trial corpus are sufficiently reliable despite a 200-trial validation with human AC1=0.39 and judge recall of 44-46%.
- domain assumption The two non-China models (Pixtral-12B and Llama-3.2-11B-Vision) are adequate representatives of non-China-origin VLMs for the origin comparison.
- domain assumption The Chinese and English prompts are semantically matched translations, so the language effect is not an artifact of prompt wording differences.
- standard math The additive logistic model (Eq. 1) and cluster-robust standard errors by image entry give valid inference for the reported odds ratios.
Cite this review
Pith. "Pith review of How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment." pith.science (2026). https://pith.science/paper/NNSWE7LM
@misc{pith2026260811816,
author = {Pith},
title = {Pith review of: How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/NNSWE7LM}},
note = {Machine review of arXiv:2608.11816}
}
read the original abstract
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Pravesh Agrawal et al. 2024. Pixtral 12B.arXiv preprint arXiv:2410.07073(2024)
arXiv 2024
-
[2]
Mohamed Ahmed, Jeffrey Knockel, and Rachel Greenstadt. 2025. An Analysis of Chinese Censorship Bias in LLMs.Proceedings on Privacy Enhancing Technologies 2025, 4 (2025), 112–129. doi:10.56553/popets-2025-0122
-
[3]
Shuai Bai et al . 2025. Qwen2.5-VL Technical Report.arXiv preprint arXiv:2502.13923(2025)
arXiv 2025
-
[4]
Yuntao Bai et al. 2022. Constitutional AI: Harmlessness from AI Feedback.arXiv preprint arXiv:2212.08073(2022)
arXiv 2022
-
[5]
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey.arXiv preprint arXiv:2404.18930(2024)
arXiv 2024
-
[6]
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623
2021
-
[7]
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021. Multimodal datasets: misogyny, pornography, and malignant stereotypes.arXiv preprint arXiv:2110.01963(2021)
arXiv 2021
-
[8]
Rishi Bommasani, Percy Liang, and Tony Lee. 2023. Holistic evaluation of lan- guage models.Annals of the New York Academy of Sciences1525, 1 (2023), 140–146
work page 2023
Show all 62 references
-
[9]
Lawrence D Brown, T Tony Cai, and Anirban DasGupta. 2001. Interval estimation for a binomial proportion.Statistical science16, 2 (2001), 101–133
2001
-
[10]
Stephen Casper et al. 2023. Open Problems and Fundamental Limitations of Re- inforcement Learning from Human Feedback.Transactions on Machine Learning Research (TMLR)(2023)
2023
-
[11]
Zhe Chen et al . 2024. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[12]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). 15607–15631
2023
-
[13]
China Digital Times. 2026. China Digital Times. https://chinadigitaltimes.net. Independent archive of content censored from the Chinese internet. Accessed 2026-08-11
2026
-
[14]
Cyberspace Administration of China. 2023. Interim Measures for the Management of Generative Artificial Intelligence Services. https://www.cac.gov.cn/2023-07/ 13/c_1690898327029107.htm
2023
-
[15]
Cyberspace Administration of China. 2025. Clear and Bright: Rectifying the Abuse of AI Technology (Special Campaign). https://www.cac.gov.cn/2025- 04/30/c_1747719097461951.htm
2025
-
[16]
Robert M. Entman. 1993. Framing: Toward Clarification of a Fractured Paradigm. Journal of Communication43, 4 (1993), 51–58
1993
-
[17]
Robert M Entman. 2007. Framing bias: Media in the distribution of power.Journal of communication57, 1 (2007), 163–173
2007
-
[18]
Aaron Grattafiori et al . 2024. The Llama 3 Herd of Models.arXiv preprint arXiv:2407.21783(2024)
2024 arXiv
-
[19]
Daya Guo et al. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.arXiv preprint arXiv:2501.12948(2025)
2025 arXiv
-
[20]
Kilem L. Gwet. 2008. Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement.Brit. J. Math. Statist. Psych.61, 1 (2008), 29–48
2008
-
[21]
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The polit- ical ideology of conversational AI: Converging evidence on ChatGPT’s pro- environmental, left-libertarian orientation.arXiv preprint arXiv:2301.01768 (2023)
2023 arXiv
-
[22]
Jen-tse Huang, Chang Chen, Shiyang Lai, Wenxuan Wang, Michelle R Kaufman, and Mark Dredze. 2026. Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation.arXiv preprint arXiv:2601.06600 (2026)
2026 arXiv
-
[23]
PeiHsuan Huang, ZihWei Lin, Simon Imbot, WenCheng Fu, and Ethan Tu. 2025. Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high.arXiv preprint arXiv:2506.01814(2025)
2025 arXiv
-
[24]
Taylor, and Stefano Soatto
Prannay Kaul, Zhizhong Li, Hao Yang, Yonatan Dukler, Ashwin Swaminathan, Christopher J. Taylor, and Stefano Soatto. 2024. THRONE: An Object-based Hal- lucination Benchmark for the Free-form Generations of Large Vision-Language Models. InProceedings of the IEEE/CVF Conference o...
2024 arXiv
-
[25]
Gary King, Jennifer Pan, and Margaret E Roberts. 2013. How censorship in China allows government criticism but silences collective expression.American political science Review107, 2 (2013), 326–343
2013
-
[26]
Gary King, Jennifer Pan, and Margaret E Roberts. 2014. Reverse-engineering censorship in China: Randomized experimentation and participant observation. Science345, 6199 (2014), 1251722
2014
-
[27]
Gary King, Jennifer Pan, and Margaret E Roberts. 2017. How the Chinese govern- ment fabricates social media posts for strategic distraction, not engaged argument. American political science review111, 3 (2017), 484–501
2017
-
[28]
Ju-Chun Ko. 2026. Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study.arXiv preprint arXiv:2602.06371(2026)
2026 arXiv
-
[29]
David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Penny- cook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonat...
2018
-
[30]
Lee and Katrina A
John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appropriate Reliance.Human Factors46, 1 (2004), 50–80
2004
-
[31]
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. 2023. Evaluating object hallucination in large vision-language models. InProceedings of the 2023 conference on empirical methods in natural language processing. 292–305
2023
-
[32]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruc- tion tuning.Advances in neural information processing systems36, 34892–34916
2023
-
[33]
Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024. Safety of multimodal large language models on images and texts.arXiv preprint arXiv:2402.00357
2024 arXiv
-
[34]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-EVAL: NLG evaluation using GPT-4 with better human alignment. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2511–2522. doi:10.18653/v1/2023...
2023 doi
-
[35]
Nestor Maslej et al. 2026. The AI Index 2026 Annual Report. AI Index Steering Committee, Institute for Human-Centered AI, Stanford University, Stanford, CA. https://hai.stanford.edu/ai-index/2026-ai-index-report
2026
-
[36]
Meta AI. 2024. Llama 3.2 Model Card (Llama-3.2-11B-Vision). https://github.com/meta-llama/llama-models/blob/main/models/llama3_ 2/MODEL_CARD_VISION.md. Released September 25, 2024
2024
-
[37]
Ali Naseh, Harsh Chaudhari, Jaechul Roh, Mingshi Wu, Alina Oprea, and Amir Houmansadr. 2025. R1dacted: Investigating Local Censorship in DeepSeek’s R1 Language Model.arXiv preprint arXiv:2505.12625(2025)
2025 arXiv
-
[38]
OpenAI. 2023. GPT-4 Technical Report.arXiv preprint arXiv:2303.08774(2023)
2023 arXiv
-
[39]
Long Ouyang et al. 2022. Training Language Models to Follow Instructions with Human Feedback. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 35. 27730–27744
2022
-
[40]
Jennifer Pan and Xu Xu. 2026. Political censorship in large language models originating from China.PNAS nexus5, 2 (2026), pgag013
2026
-
[41]
Vaidehi Patil, Yi-Lin Sung, Peter Hase, Jie Peng, Tianlong Chen, and Mohit Bansal
-
[42]
Gordon Pennycook and David G. Rand. 2021. The Psychology of Fake News. Trends in Cognitive Sciences25, 5 (2021), 388–402
2021
-
[43]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[44]
2018.Censored: distraction and diversion inside China’s Great Firewall
Margaret Roberts. 2018.Censored: distraction and diversion inside China’s Great Firewall. Princeton University Press
2018
-
[45]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect?. In International conference on machine learning. PMLR, 29971–30004
2023
-
[46]
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design. InInterna- tional Conference on Learning Representations (ICLR). OpenReview.net, Vienna, Austria, 24 pages
2024
-
[47]
Vera Liao, and Ziang Xiao
Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI)
2024
-
[48]
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural...
2019
-
[49]
Skitka, Kathleen L
Linda J. Skitka, Kathleen L. Mosier, and Mark Burdick. 1999. Does Automation Bias Decision-Making?International Journal of Human-Computer Studies51, 5 (1999), 991–1006
1999
-
[50]
Zhi Rui Tam, Yung-Yu Shih, Yen-Wei Lee, Ya-Ting Pai, Wen Yu Chang, and Yun- Nung Chen. 2026. VisTW: Benchmarking vision-language models for Taiwanese Mandarin in Taiwan. InFindings of the Association for Computational Linguistics: ACL 2026. 36711–36756. doi:10.18653/v1/2026.fi...
2026 doi
-
[51]
An Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Vy Tuong Dang, Anh Totti Nguyen, and Daeyoung Kim. 2025. Vision language models are biased. arXiv preprint arXiv:2505.23941(2025)
2025 arXiv
-
[52]
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. 2023. Decod- ingTrust: A Comprehensive Assessment of Trustworthiness in{GPT} Models. Neural Information Processing Systems Datasets; Benc...
2023
-
[53]
Peng Wang, Shuai Bai, et al . 2024. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution.arXiv preprint arXiv:2409.12191(2024)
2024 arXiv
-
[54]
Edwin B Wilson. 1927. Probable inference, the law of succession, and statistical inference.J. Amer. Statist. Assoc.22, 158 (1927), 209–212
1927
-
[55]
Yuan Yao et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone.arXiv preprint arXiv:2408.01800(2024)
2024 arXiv
-
[56]
Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao. 2026. SafeBench: A safety evaluation framework for multimodal large language models.International Journal of Computer Vision 134, 1 (2026). doi:10.1007/s11263-025-02613-...
2026 doi
-
[57]
Hengxiang Zhang, Hongfu Gao, Qiang Hu, Guanhua Chen, Lili Yang, Bingyi Jing, Hongxin Wei, Bing Wang, Haifeng Bai, and Lei Yang. 2024. Chinesesafe: A chinese benchmark for evaluating safety in large language models.arXiv preprint arXiv:2410.18491(2024)
2024 arXiv
-
[58]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36, 46595–46623
2023
-
[59]
Parker, and Munmun De Choudhury
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker, and Munmun De Choudhury. 2023. Synthetic Lies: Understanding AI-Generated Misinfor- mation and Evaluating Algorithmic and Human Solutions. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI)
2023
-
[60]
vocational training
Jinguo Zhu et al. 2025. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.arXiv preprint arXiv:2504.10479 (2025). 12 A Data Table A1 lists the exact publicly released, post-trained instruction-tuned checkpoints used in the audit. Ta...
2025 arXiv
-
[997]
Cross-seed stability
as thedescribebaseline, under identical sampling ( temperature=0.7, top_p=0.8, top_k=20, max_tokens=1024; reasoning disabled for the two reasoning-capable checkpoints), giving200× 9× 2× 3 = 10,800appendix-only trials. These trials are paired to the maindescribe baseline for se...
-
[2025]
Unlearning sensitive information in multimodal LLMs: Benchmark and attack-defense evaluation.arXiv preprint arXiv:2505.01456(2025)
2025 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.