Pith. sign in

REVIEW 3 major objections 4 minor 150 references

Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Writers who meet the AI in Spanish first use its English drafting feature roughly 64% less, a violation of the rational-choice axiom that choices in one language should not depend on experience in another.

desk verdict The headline choice-independence claim is not supported by the reported analysis; the second experiment is the more solid contribution. read the letter →

arxiv 2502.09532 v1 pith:V7JJVIVM submitted 2025-02-13 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords Human-AIinteractionChoiceindependenceMultilingualLLMsUserrelianceCo-writingPersuasivewritingCharitablegivingAlgorithmaversion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports two experiments asking whether people use a multilingual LLM writing assistant independently across languages. It finds they do not: writers who first used the assistant in Spanish used its AI text-generation feature roughly 64% less when writing in English afterward, even though the underlying model was unchanged. The result is offered as evidence that choice independence, a standard assumption of rational choice theory, is violated in an applied co-writing setting, extending earlier lab-only demonstrations. A second experiment on charitable giving finds that the usage gap did not change how persuasive the ads were, but donors' beliefs about whether an ad was written by AI did affect giving, especially among Spanish-speaking women.

What carries the argument

The central mechanism is the order-of-exposure design built into a customized version of the co-writing tool ABScribe, which offers an AI Drafter feature that generates new text from a user prompt. Writers in two bilingual conditions wrote charity ads in English and Spanish in opposite orders, and the count of AI Drafter uses in the second-language task served as the revealed-utility measure. The paper supplements usage counts with a weighted-average embedding similarity between AI-generated segments and the final ad, and with benchmark checks showing the underlying LLaMA 3.1 model follows instructions worse in Spanish than English. The independence axiom of expected utility theory supplies the normative baseline: if writers evaluated languages independently, usage in the second task should not depend on which language came first.

What would settle it

Run the same bilingual writing study with an additional English-then-English condition. If writers in the second English session also cut AI Drafter usage by roughly 60%, the observed gap is an order or fatigue effect rather than language transfer; if English-second usage stays near English-first levels, the choice-independence interpretation is supported.

Watch

Extended reading notes

Core claim

The paper's central claim is Result 1: prior exposure to the Spanish version of the LLM assistant reduces subsequent utilization of the English version, a large drop in AI Drafter usage from 28 total uses in the English-first condition to 10 in the Spanish-first condition. The authors interpret this as a violation of choice independence, arguing that users extrapolate the assistant's lower Spanish quality to its English performance and adjust reliance accordingly. The reverse pattern is also reported: participants who met the English assistant first used the Spanish assistant more. The downstream experiment adds two further findings: utilization differences do not change aggregate donation outcomes, and donors cannot reliably tell human from AI ads but condition their giving on what they believe, with Spanish-speaking female participants showing the strongest negative reaction to perceived AI authorship.

Load-bearing premise

The claim that the 64% drop in English AI Drafter use was caused by prior Spanish exposure assumes that no order, fatigue, or learning effect explains lower usage in the second writing task; the design has no English-then-English control showing usage would stay high without Spanish first.

Editorial extensions

If this is right

  • If the finding replicates, developers of multilingual writing tools should expect usage in one language to depend on users' experience in another, so cross-linguistic quality gaps can suppress adoption even where the model is strong.
  • Persuasiveness results imply that utilization drops of this size need not degrade output quality in charity-ad writing, because donations did not differ across treatment conditions.
  • Because most donors could not distinguish human from LLM ads yet responded to perceived AI authorship, disclosure or labeling decisions can directly affect charitable giving.
  • In the authors' sample, the belief effect is concentrated in Spanish-speaking women, who donated 22% less and were nearly four times more likely to donate nothing when they believed an AI wrote the ad; other groups showed weaker, non-significant versions of the same pattern.
  • Rolling out multilingual assistants without monitoring post-deployment language-specific quality could create second-order inequality, since lower-resource-language users may generalize poor experiences and under-use the tool in higher-resource languages as well.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the observed drop is causal, a plausible mechanism is that users form a holistic quality impression of the assistant rather than separate per-language beliefs; the paper does not directly measure such beliefs, so this remains an inference.
  • A natural testable extension would vary the linguistic and resource distance between first and second language, for example comparing Spanish with a typologically distant low-resource language, to see whether generalization grows with perceived similarity.
  • The donor-belief results suggest that a simple intervention, such as announcing human involvement in drafting or not disclosing AI use, could move donation totals in specific demographics; this follows from the paper's findings but was not tested as a treatment.
  • The design could be extended to repeated interactions over time to check whether the utilization gap decays with learning or persists, which the authors flag as outside their one-session scope.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports two preregistered experiments on multilingual human-AI co-writing. In Experiment 1, bilingual writers (n=16 per condition) use an LLM-based writing assistant (ABScribe) to produce charity advertisements for the WWF in English and Spanish, with the language order manipulated (ENG_ESP vs. ESP_ENG) plus no-LLM controls. The main reported finding (Result 1) is that writers exposed to the Spanish LLM first use the AI Drafter feature less in the subsequent English task (a drop of roughly 64%, from 28 to 10 total uses), which the authors interpret as a violation of choice independence. A secondary analysis suggests the reverse effect for Spanish when English is experienced first. Experiment 2 (n=720) evaluates the persuasiveness of the resulting ads in a donation task, finding no effects of the utilization differences on donations, but reporting that participants cannot reliably identify AI-generated ads and that Spanish-speaking female participants who believe they read an AI-generated ad donate significantly less. The paper claims to be the first to extend choice-independence violations from abstract lab tasks to an applied multilingual co-writing setting.

Significance. If the causal interpretation of Result 1 holds, the paper makes a meaningful contribution by showing that users' experiences with an LLM in one language spill over to their reliance on the same model in another language, with practical implications for the global deployment of multilingual assistants. The work has several concrete strengths: both experiments are preregistered; the code and data are made openly available; the language-performance assumptions are checked with external benchmarks (Multi-IF, PAWS-X, persuasion detection); and the downstream persuasion experiment with 720 participants is a valuable attempt to connect utilization changes to real-world outcomes. The paper is also explicitly an empirical extension with new data rather than a circular argument. However, the central causal claim rests on a between-group comparison that confounds language order with task position, and the reported analyses do not include the interaction test that would isolate a choice-independence violation.

major comments (3)
  1. [§5.1, Result 1 and Figure 4] The two t-tests used to support Result 1 are separate between-group comparisons on n=16 per group (ENG_1 vs. ENG_2: t=2.2, p=0.04; ESP_1 vs. ESP_2: t=2.58, p=0.017). These tests are not corrected for the multiple feature and language comparisons examined in the same section, and they do not exploit the within-subject pairing that the crossover design provides. The reported significance of the simple effects does not establish a differential effect of order by language; the absence of an interaction test is a particular problem because the observed pattern (both English and Spanish counts lower in Group B than in Group A) is exactly what would be expected from a stable participant-level difference in AI usage. A proper mixed-model analysis of the raw pair-level usage counts is required before the utilization drop can be attributed to prior exposure to Spanish.
  2. [§4.1.3 / §5.1, weighted average similarity] The weighted average similarity measure is used as a second piece of evidence for Result 1, but it is not validated against ground-truth text provenance. The paper reports that ENG_1 advertisements are 14.7% more similar to AI-generated text than ENG_2, using three embedding models, but it does not show that this metric discriminates AI-written from human-written text in this domain, nor does it report calibration or reliability across the two languages. Without such validation, the similarity difference may reflect topic, style, or length differences between first- and second-position writing tasks rather than differential reliance on AI-generated content. The authors should either validate the measure (e.g., against known human-only and LLM-only texts) or present it as a descriptive, non-causal auxiliary result.
  3. [§6.2, Caveats and Limitations] The limitations section acknowledges the small number of writers and the difficulty of recruiting bilingual participants, but it does not explicitly acknowledge the design-level confound between language order and task position, nor the fact that no same-language-order control condition was included. Given that these issues directly qualify the paper's headline result, the limitation should be stated transparently and the appropriate interaction analysis should be reported in the main text rather than only in future-work remarks. The current presentation of Result 1 as 'consistent with violations of choice independence' is stronger than the design and reported statistics support.
minor comments (4)
  1. [§5.2, Result 2] The sentence 'Average donations do not differ between treatments, providing any evidence for a detrimental effect of choice independence violations on persuasion' is contradictory as written; it should read 'providing no evidence' or 'we find no evidence'. The subsequent text correctly reports the absence of an effect, so this is a wording error that should be corrected.
  2. [§6, RQ2 and RQ3 discussion] The paper uses the phrases 'moderate evidence', 'minor evidence', and 'strong evidence' without defining an evidentiary threshold or reporting effect sizes. I recommend stating the chosen standard (e.g., based on p-values or Bayes factors) or replacing these qualitative labels with the actual estimates and confidence intervals.
  3. [References and appendix] Some placeholders remain unresolved: the regression results in §5.2 are referenced as 'Tables ?? and ??', and Table 5 is said to be 'available on the next page'. In addition, references [90]/[91] and [129]/[130] appear to be duplicated with different citation keys. These should be cleaned up before publication.
  4. [Figure 9] The label 'Average time spent outside' is unclear; I assume it means time spent outside the writing platform (e.g., in other browser tabs), but this should be defined in the figure caption or text, and it is unclear how this was measured.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the load-bearing results are direct empirical measurements of AI-drafter usage compared across language-order conditions, with external benchmarks and no fitted parameter renamed as a prediction.

full rationale

The paper's central claims are empirical observations of logged interaction counts, not quantities derived from fitted parameters or from the cited prior work. The key comparison (ENG_1 vs ENG_2: 28 vs 10 AI-drafter uses, t=2.2, p=0.04) is a direct measurement of user behavior; no parameter is fit to the outcome and then relabeled as a prediction. The benchmarks used to document English-Spanish performance gaps (Multi-IF, PAWS-X, and a persuasion-detection dataset) are external instruments with fixed scoring procedures, and Experiment 2's donation game is a separate outcome measure collected independently of Experiment 1's usage data. The authors cite their own prior work, Erlei et al. [30], to motivate and interpret the choice-independence construct, but the current experiment uses newly collected data and the conclusion does not reduce to that citation; the self-citation is contextual framing rather than load-bearing evidence. The absence of a same-language-order control and the lack of a reported Language×Order interaction are genuine threats to the causal interpretation, but they concern validity, not circularity: the observed usage drop is not defined to equal the treatment condition. No circular step can be exhibited from the paper's equations or derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No fitted parameters are used to define the outcomes; experimental design choices (temperature 0.8, incentive amounts, recipe prompts) are hand-set but not tuned to produce the results. The central claim rests on three domain assumptions, listed above.

assumptions (3)
  • domain assumption Utilization differences between first- and second-language tasks are caused by the language of prior exposure, not by task order, fatigue, or learning effects.
    The design has no same-language-order control, so the causal attribution in Result 1 requires this assumption.
  • domain assumption Weighted average similarity between AI-generated segments and final text is a valid measure of reliance on LLM output.
    Used in Section 5.1 to support the reliance claim; no validation against ground-truth authorship.
  • domain assumption Benchmark differences between English and Spanish Llama 3.1 (Multi-IF, PAWS-X, persuasion detection) correspond to a perceptible quality gap in the writing task.
    The manipulation relies on participants experiencing worse Spanish output; this is inferred from benchmarks, not directly measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages." pith.science (2026). https://pith.science/paper/V7JJVIVM

@misc{pith2026250209532,
  author       = {Pith},
  title        = {Pith review of: Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V7JJVIVM}},
  note         = {Machine review of arXiv:2502.09532}
}
read the original abstract

Recent advances in generative AI have precipitated a proliferation of novel writing assistants. These systems typically rely on multilingual large language models (LLMs), providing globalized workers the ability to revise or create diverse forms of content in different languages. However, there is substantial evidence indicating that the performance of multilingual LLMs varies between languages. Users who employ writing assistance for multiple languages are therefore susceptible to disparate output quality. Importantly, recent research has shown that people tend to generalize algorithmic errors across independent tasks, violating the behavioral axiom of choice independence. In this paper, we analyze whether user utilization of novel writing assistants in a charity advertisement writing task is affected by the AI's performance in a second language. Furthermore, we quantify the extent to which these patterns translate into the persuasiveness of generated charity advertisements, as well as the role of peoples' beliefs about LLM utilization in their donation choices. Our results provide evidence that writers who engage with an LLM-based writing assistant violate choice independence, as prior exposure to a Spanish LLM reduces subsequent utilization of an English LLM. While these patterns do not affect the aggregate persuasiveness of the generated advertisements, people's beliefs about the source of an advertisement (human versus AI) do. In particular, Spanish-speaking female participants who believed that they read an AI-generated advertisement strongly adjusted their donation behavior downwards. Furthermore, people are generally not able to adequately differentiate between human-generated and LLM-generated ads. Our work has important implications for the design, development, integration, and adoption of multilingual LLMs as assistive agents -- particularly in writing tasks.

Figures

Figures reproduced from arXiv: 2502.09532 by the authors.

Figure 1
Figure 1. The ABScribe writing interface used in the exper [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Experiment Workflow for LLM-Assisted Writing [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 4
Figure 4. Effect of Initial Language Exposure on AI Drafter Usage by Task Group. Left: Total usage count of the AI drafter feature shows a significant "gap" between task groups based on initial language exposure. The group exposed to English first (ENG_1, followed by ESP_2) shows substantially higher usage compared to the group exposed to Spanish first (ESP_1, followed by ENG_2), as indicated by the significant differ￾ences m… view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: Weighted Average Similarity Percentage Across Task Groups and Models. The similarity scores vary across task groups depending on the initial language exposure but pattern remains consistent across embedding models. ENG_1 and ESP_2, which involve starting with English, …
Figure 6
Figure 6. Figure 6: Left: Average donations across treatments. Right: [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: The figure shows the differences in average dona￾tion amounts (left) and donation shares (right) across actual ad sources (Control, Human, Human-AI, and AI). Bars are further categorized by the perceived source of the ad text (AI or Human). Error bars represent the con…
Figure 8
Figure 8. Figure 8: The figure shows the differences in average dona￾tion amounts (left) and donation shares (right) based on par￾ticipant demographics (sex) and language groups (Spanish￾speaking and English-speaking). Bars are further catego￾rized by the perceived source of the persuasiv…
Figure 9
Figure 9. Figure 9: Left: Average word count across treatments, indicat￾ing variations in output length depending on treatment and condition. Right: Average time spent outside the platform across treatments. Error bars represent 95% confidence inter￾vals. 8.3 Stage 1 - Writing Task Metric…
Figure 10
Figure 10. Figure 10: Differences in stated utility and feature usage percentages across treatment groups for English (Left) and Spanish (Right) tasks. Bars represent feature usage percent￾ages, while lines indicate normalized stated preferences as reported in questionnaire responses. This…
Figure 14
Figure 14. Figure 14: Left: Distribution of donation amounts across age groups, showing the variability in donation behavior as a function of age. Right: Comparison of age distributions be￾tween donors and non-donors, highlighting differences in median age and interquartile ranges for both…
Figure 12
Figure 12. Figure 12: Average scores for Behavioral Intention (left), Emo￾tional Appeal (middle), and Information Awareness (right) across various experimental conditions. Error bars represent 95% confidence intervals. 8.6 Stage 2: Donation Behaviour by Demographic Factors Asian Black Mixe…
Figure 13
Figure 13. Figure 13: Left: Average donation amount across different ethnic groups. Right: Share of donors (percentage of partic￾ipants who donated) by ethnicity, highlighting variations in donation behavior and likelihood among demographic groups. Error bars representing 95% confidence in…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

150 extracted references · 39 canonical work pages

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Mortensen, Noah A

    Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. 2023. Do All Languages Cost the Same? Tok- enization in the Era of Commercial Language Models. arXiv:2305.13707 [cs.CL] https://arxiv.org/abs/2305.13707

  3. [3]

    Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023. MEGA: Multilingual evaluation of generative AI. arXiv [cs.CL] (March 2023)

  4. [4]

    Reut Apel, Ido Erev, Roi Reichart, and Moshe Tennenholtz. 2020. Predicting Decisions in Language Based Persuasion Games. arXiv [cs.AI] (Dec. 2020)

  5. [5]

    Luis Arango, Stephen Pragasam Singaraju, and Outi Niininen. 2023. Consumer responses to AI-generated charitable giving ads. Journal of Advertising 52, 4 (2023), 486–503

  6. [6]

    Chaitanya Arora, Utkarsh Venaik, Pavit Singh, Sahil Goyal, Jatin Tyagi, Shyama Goel, Ujjwal Singhal, and Dhruv Kumar. 2024. Analyzing LLM Usage in an Advanced Computing Class in India. arXiv preprint arXiv:2404.04603 (2024)

  7. [7]

    Payal Arora. 2024. From pessimism to promise: Lessons from the Global South on designing inclusive tech. MIT Press

  8. [8]

    Richard D Ashmore, Kay Deaux, and Tracy McLaughlin-Volpe. 2004. An or- ganizing framework for collective identity: articulation and significance of multidimensionality. Psychological bulletin 130, 1 (2004), 80

Show all 150 references
  1. [9]

    Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.0...

  2. [10]

    Lucas Bellaiche, Rohin Shahi, Martin Harry Turpin, Anya Ragnhildstveit, Shawn Sprockett, Nathaniel Barr, Alexander Christensen, and Paul Seli. 2023. Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cognitive Research: Principles and Imp...

  3. [11]

    Alec Brandon, Christopher M Clapp, John A List, Robert D Metcalfe, and Michael Price. 2022. The Human Perils of Scaling Smart Technologies: Evidence from Field Experiments. Technical Report. National Bureau of Economic Research

  4. [12]

    Simon Martin Breum, Daniel Vædele Egdal, Victor Gram Mortensen, Anders Gio- vanni Møller, and Luca Maria Aiello. 2024. The persuasive power of Large Language Models. Proceedings of the International AAAI Conference on Web and Social Media 18 (May 2024), 152–163

  5. [13]

    Simon Martin Breum, Daniel Vædele Egdal, Victor Gram Mortensen, Anders Gio- vanni Møller, and Luca Maria Aiello. 2024. The persuasive power of large lan- guage models. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 18. 152–163

  6. [14]

    Jasper David Brüns and Martin Meißner. 2024. Do you create your content your- self? Using generative artificial intelligence for social media content creation diminishes perceived brand authenticity. Journal of Retailing and Consumer Services 79 (2024), 103790

  7. [15]

    Yaqi Chen, Haizhong Wang, Sally Rao Hill, and Binglian Li. 2024. Consumer attitudes toward AI-generated ads: Appeal types, self-efficacy and AI’s social role. Journal of Business Research 185 (2024), 114867

  8. [16]

    Jungsil Choi and Hyun Young Park. 2020. How Donor’s Regulatory Focus Changes the Effectiveness of a Sadness-Evoking Charity Appeal. Social Science Research Network (Sept. 2020)

  9. [17]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1–113

  10. [18]

    Marina Chugunova and Daniela Sele. 2022. We and It: An interdisciplinary review of the experimental evidence on how humans interact with machines. Journal of Behavioral and Experimental Economics 99 (2022), 101897

  11. [19]

    Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021. All that’s’ human’is not gold: Evaluating human evaluation of generated text. arXiv preprint arXiv:2107.00061 (2021)

  12. [20]

    Javier Conde, Miguel González, Nina Melero, Raquel Ferrando, Gonzalo Martínez, Elena Merino-Gómez, José Alberto Hernández, and Pedro Reviriego

  13. [21]

    Emma Dafouz-Milne. 2008. The pragmatic role of textual and interpersonal metadiscourse markers in the construction and attainment of persuasion: A cross-linguistic study of newspaper discourse. J. Pragmat. 40, 1 (Jan. 2008), 95–113

  14. [22]

    Paramveer S Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel P Robert. 2024. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. arXiv [cs.HC] (Feb. 2024)

  15. [23]

    Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2015. Algorithm aversion: people erroneously avoid algorithms after seeing them err. Journal of experimental psychology: General 144, 1 (2015), 114

  16. [24]

    Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2018. Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Management science 64, 3 (2018), 1155–1170

  17. [25]

    Anil R Doshi and Oliver Hauser. 2023. Generative artificial intelligence enhances creativity. A vailable at SSRN (2023)

  18. [26]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  19. [27]

    Esin Durmus, Liane Lovitt, Alex Tamkin, Stuart Ritchie, Jack Clark, and Deep Ganguli. 2024. Measuring the Persuasiveness of Language Models

  20. [28]

    Alexander Erlei, Richeek Das, Lukas Meub, Avishek Anand, and Ujwal Gadiraju

  21. [29]

    Alexander Erlei, Franck Nekdem, Lukas Meub, Avishek Anand, and Ujwal Gadiraju. 2020. Impact of algorithmic decision making on human behavior: Evidence from ultimatum bargaining. In Proceedings of the AAAI conference on human computation and crowdsourcing, Vol. 8. 43–52

  22. [30]

    Alexander Erlei, Abhinav Sharma, and Ujwal Gadiraju. 2024. Understanding Choice Independence and Error Types in Human-AI Collaboration. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems . 1–19

  23. [31]

    Kawin Ethayarajh and Dan Jurafsky. 2022. The Authenticity Gap in Human Evaluation. arXiv [cs.CL] (May 2022)

  24. [32]

    Meta Fundamental AI Research Diplomacy Team (FAIR)†, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al. 2022. Human-level play in the game of Diplomacy by combining language models with strateg...

  25. [33]

    Xiaojun Fan, Nianqi Deng, Yi Qian, and Xuebing Dong. 2020. Factors affecting the effectiveness of cause-related marketing: A meta-analysis.Journal of Business Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages CHI ’25, April 2...

  26. [34]

    Carla Ferraro, Vlad Demsar, Sean Sands, Mariluz Restrepo, and Colin Campbell

  27. [35]

    Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024. Open LLM Leaderboard v2. https://huggingface.co/spaces/ open-llm-leaderboard/open_llm_leaderboard

  28. [36]

    Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra, Yudai Yamazaki, Yasutaka Nishimura, Sina J Semnani, Kazushi Ikeda, Weiyan Shi, and Monica S Lam. 2024. Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval. arXiv preprint arXiv:2407.03585 (2024)

  29. [37]

    Business Horizons (2024)

    The paradoxes of generative AI-enabled customer service: A guide for managers. Business Horizons (2024)

  30. [38]

    Ella Glikson and Anita Williams Woolley. 2020. Human trust in artificial intel- ligence: Review of empirical research. Academy of Management Annals 14, 2 (2020), 627–660

  31. [39]

    Laura Globig, Rachel Xu, Steve Rathje, and Jay J Van Bavel. 2024. Perceived (Mis) alignment in generative Artificial Intelligence Varies Across Cultures. (2024)

  32. [40]

    Xiao Ge, Chunchen Xu, Daigo Misaki, Hazel Rose Markus, and Jeanne L Tsai

  33. [41]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems

    How Culture Shapes What People Want From AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15

  34. [42]

    Simone Grassini and Mika Koivisto. 2024. Understanding how personality traits, experiences, and attitudes shape negative bias toward AI-generated artworks. Scientific Reports 14, 1 (2024), 4113

  35. [43]

    Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, Ujwal Gadiraju, Aditya Vashistha, Vivek Seshadri, et al. 2024. Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology. In The 2024 ACM Conferen...

  36. [44]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2023. Can AI Write Persuasive Propaganda?

  37. [45]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2023. Can ai write persuasive propaganda? SocArXiv. April 8 (2023)

  38. [46]

    Gaole He, Gianluca Demartini, and Ujwal Gadiraju. 2025. Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant. In Proceedings of the CHI Conference on Human Factors in Computing Systems

  39. [47]

    Gaole He, Lucie Kuiper, and Ujwal Gadiraju. 2023. Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23, Vol. 144). ACM, 1–18. https://doi.o...

  40. [48]

    Jakub Harasta, Tereza Novotná, and Jaromir Savelka. 2024. It cannot be right if it was written by AI: On lawyers’ preferences of documents perceived as authored by an LLM vs a human. arXiv [cs.HC] (July 2024)

  41. [49]

    Arid Hasan, Prerona Tarannum, Krishno Dey, Imran Razzak, and Usman Naseem

    Md. Arid Hasan, Prerona Tarannum, Krishno Dey, Imran Razzak, and Usman Naseem. 2024. Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings. arXiv:2408.02237 [cs.CL] https://arxiv.org/abs/2408.02237

  42. [50]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Under- standing. Proceedings of the International Conference on Learning Representations (ICLR) (2021)

  43. [51]

    Steffen Herbold, Annette Hautli-Janisz, Ute Heuer, Zlata Kikteva, and Alexander Trautsch. 2023. AI, write an essay for me: A large-scale comparison of human- written versus ChatGPT-generated essays. arXiv preprint arXiv:2304.14276 (2023)

  44. [52]

    Yun He, Di Jin, Chaoqi Wang, Chloe Bi, Karishma Mandyam, Hejia Zhang, Chen Zhu, Ning Li, Tengyu Xu, Hongjiang Lv, et al. 2024. Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following. arXiv preprint arXiv:2410.15553 (2024)

  45. [53]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI With Shared Human Values. Proceedings of the International Conference on Learning Representations (ICLR) (2021)

  46. [54]

    Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, and Satoshi Nakamura. 2014. Reinforcement learning of cooperative persuasive dialogue policies using framing. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Pa...

  47. [55]

    Charles A Holt. 1986. Preference reversals and the independence axiom. The American Economic Review 76, 3 (1986), 508–515

  48. [56]

    Steffen Herbold, Annette Hautli-Janisz, Ute Heuer, Zlata Kikteva, and Alexander Trautsch. 2023. A large-scale comparison of human-written versus ChatGPT- generated essays. Scientific reports 13, 1 (2023), 18617

  49. [57]

    Sally Hibbert, Andrew Smith, Andrea Davies, and Fiona Ireland. 2007. Guilt appeals: Persuasion knowledge and charitable giving. Psychology and Marketing 24, 8 (Aug. 2007), 723–742

  50. [58]

    Muhammad Abid Jamil, Muhammad Arif, Normi Sham Awang Abubakar, and Akhlaq Ahmad. 2016. Software testing techniques: A literature review. In 2016 6th international conference on information and communication technology for the Muslim world (ICT4M) . IEEE, 177–182

  51. [59]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)

  52. [60]

    Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023. Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting. arXiv preprint arXiv:2305.07004 (2023)

  53. [61]

    Martin Huschens, Martin Briesch, Dominik Sobania, and Franz Rothlauf. 2023. Do You Trust ChatGPT?–Perceived Credibility of Human and AI-Generated Content. arXiv preprint arXiv:2309.02524 (2023)

  54. [62]

    Karsten Jonsen, Jacqueline Fendt, and Sébastien Point. 2018. Convincing quali- tative research: What constitutes persuasive writing? Organizational Research Methods 21, 1 (2018), 30–67

  55. [63]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choud- hury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. arXiv preprint arXiv:2004.09095 (2020)

  56. [64]

    Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is ChatGPT a good translator? A preliminary study. arXiv preprint arXiv:2301.08745 1, 10 (2023)

  57. [65]

    Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar. 2024. Better to ask in English: Cross-lingual evaluation of large language models for healthcare queries. In Proceedings of the ACM on Web Conference 2024. 2627–2638

  58. [66]

    Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T Hancock. 2023. Working with AI to persuade: Examining a large language model’s ability to generate pro-vaccination messages. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023), 1–29

  59. [67]

    Tae Soo Kim, Yoonjoo Lee, Minsuk Chang, and Juho Kim. 2023. Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23, Article ...

  60. [68]

    Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and Applications of Large Language Models. (July 2023). https://doi.org/10.48550/arXiv.2307.10169

  61. [69]

    Daniel Kahneman and Amos Tversky. 2013. Prospect theory: An analysis of decision under risk. In Handbook of the Fundamentals of Financial Decision Making. WORLD SCIENTIFIC, 99–127

  62. [70]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. (2023)

  63. [71]

    Mina Lee, Percy Liang, and Qian Yang. 2022. Coauthor: Designing a human- ai collaborative writing dataset for exploring language model capabilities. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–19

  64. [72]

    Nils Köbis and Luca D Mossink. 2021. Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry. Computers in human behavior 114 (2021), 106553

  65. [73]

    Olanrewaju Israel Lawal, Olubayo Adekanmbi, and Anthony Soronnadi. 2024. Contextual Evaluation of LLM’s Performance on Primary Education Science Learning Contents in the Yoruba Language. In 5th Workshop on African Natural Language Processing

  66. [74]

    Sue Lim and Ralf Schmälzle. 2024. The effect of source disclosure on evaluation of AI-generated messages. Computers in Human Behavior: Artificial Humans 2, 1 (2024), 100058

  67. [75]

    John A List. 2022. The voltage effect: How to make good ideas great and great ideas scale. Crown Currency

  68. [76]

    Mina Lee, Percy Liang, and Qian Yang. 2022. CoAuthor: Designing a human-AI collaborative writing dataset for exploring language model capabilities. In CHI Conference on Human Factors in Computing Systems . ACM, New York, NY, USA

  69. [77]

    Xianming Li and Jing Li. 2023. AnglE-optimized Text Embeddings.arXiv preprint arXiv:2309.12871 (2023)

  70. [78]

    Stephen MacNeil, Andrew Tran, Juho Leinonen, Paul Denny, Joanne Kim, Arto Hellas, Seth Bernstein, and Sami Sarsa. 2022. Automatically generating cs learning materials with large language models. arXiv preprint arXiv:2212.05113 (2022)

  71. [79]

    S C Matz, J D Teeny, S S Vaid, H Peters, G M Harari, and M Cerf. 2024. The potential of generative AI for personalized persuasion at scale. Sci. Rep. 14, 1 (Feb. 2024), 4692. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Biswas et al

  72. [80]

    Zhuoran Lu, Sheshera Mysore, Tara Safavi, Jennifer Neville, Longqi Yang, and Mengting Wan. 2024. Corporate Communication Companion (CCC): An LLM- empowered Writing Assistant for Workplace Social Media. arXiv preprint arXiv:2405.04656 (2024)

  73. [81]

    Xiaoyu Luo, Daping Liu, Fan Dang, and Hanjiang Luo. 2024. Integration of LLMs and the Physical World: Research and Application. In Proceedings of the ACM Turing A ward Celebration Conference-China 2024. 1–5

  74. [82]

    Kobe Millet, Florian Buehler, Guanzhong Du, and Michail D Kokkoris. 2023. Defending humankind: Anthropocentric bias in the appreciation of AI art. Computers in Human Behavior 143 (2023), 107707

  75. [83]

    Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–34

  76. [84]

    Meta. 2024. Meta AI Is Now Multilingual, More Creative and Smarter. https://about.fb.com/news/2024/07/meta-ai-is-now-multilingual-more- creative-and-smarter/ Accessed: 2024-11-15

  77. [85]

    Selina Meyer, David Elsweiler, Bernd Ludwig, Marcos Fernandez-Pichel, and David E Losada. 2022. Do we still need human assessors? prompt-based gpt-3 user simulation in conversational ai. In Proceedings of the 4th Conference on Conversational User Interfaces. 1–6

  78. [86]

    Mahsan Nourani, Donald R Honeycutt, Jeremy E Block, Chiradeep Roy, Tahrima Rahman, Eric D Ragan, and Vibhav Gogate. 2020. Investigating the importance of first impressions and explainable ai with interactive video analysis. In Extended Abstracts of the 2020 CHI Conference on H...

  79. [87]

    Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the pro- ductivity effects of generative artificial intelligence. Science 381, 6654 (2023), 187–192

  80. [88]

    P Karen Murphy. 2001. What makes a text persuasive? Comparing students’ and experts’ conceptions of persuasiveness. Int. J. Educ. Res. 35, 7 (Jan. 2001), 675–698

  81. [89]

    Devon Myers, Rami Mohawesh, Venkata Ishwarya Chellaboina, Anantha Lak- shmi Sathvik, Praveen Venkatesh, Yi-Hui Ho, Hanna Henshaw, Muna Al- hawawreh, David Berdik, and Yaser Jararweh. 2024. Foundation and large language models: fundamentals, challenges, opportunities, and socia...

  82. [91]

    Teemu Pöyhönen, Mika Hämäläinen, and Khalid Alnajjar. 2022. Multilingual Persuasion Detection: Video Games as an Invaluable Data Source for NLP.ArXiv abs/2207.04453 (2022). https://api.semanticscholar.org/CorpusID:250426445

  83. [92]

    Morris, Brandon Duderstadt, and Andriy Mulyar

    Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar

  84. [93]

    arXiv:2402.01613 [cs.CL]

    Nomic Embed: Training a Reproducible Long Context Text Embedder. arXiv:2402.01613 [cs.CL]

  85. [94]

    Saumya Pareek, Eduardo Velloso, and Jorge Goncalves. 2024. Trust Development and Repair in AI-Assisted Decision-Making during Complementary Expertise. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 546– 561

  86. [95]

    Mohi Reza, Nathan Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan “michael” Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, and Joseph Jay Williams. 2023. ABScribe: Rapid Exploration & Organization of Mul- tiple Writing Variations in Human-AI Co-Writing...

  87. [96]

    Mike Rose and Michael Anthony Rose. 2009. Writer’s block: The cognitive dimension. SIU Press

  88. [97]

    Irene Rae. 2024. The Effects of Perceived AI Use On Content Perceptions. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14

  89. [98]

    Rambachan and Others

    L. Rambachan and Others. 2024. Large language models don’t behave like peo- ple, even though we may expect them to. https://computing.mit.edu/news/large- language-models-dont-behave-like-people-even-though-we-may-expect- them-to/. Accessed: 2024-11-21

  90. [99]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084 [cs.CL] https://arxiv.org/abs/ 1908.10084

  91. [100]

    Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2024. On the conversational persuasiveness of large language models: A randomized controlled trial. arXiv preprint arXiv:2403.14380 (2024)

  92. [101]

    Christina Schamp, Mark Heitmann, Tammo HA Bijmolt, and Robin Katzen- stein. 2023. The effectiveness of cause-related marketing: A meta-analysis on consumer responses. Journal of Marketing Research 60, 1 (2023), 189–215

  93. [102]

    Alexander K Saeri, Peter Slattery, Joannie Lee, Thomas Houlden, Neil Farr, Romy L Gelber, Jake Stone, Lee Huuskes, Shane Timmons, Kai Windle, et al

  94. [103]

    Sivan Schwartz, Avi Yaeli, and Segev Shlomov. 2023. Enhancing trust in LLM- based AI automation agents: New considerations and future challenges. arXiv preprint arXiv:2308.05391 (2023)

  95. [104]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2023. A Missing Piece in the Puzzle: Considering the Role of Task Complexity in Human-AI Decision Making. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization (Limassol, Cyprus) (UMAP ’23). As...

  96. [105]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  97. [106]

    Zhengyu Shen and Liyin Jin. 2024. Bargaining with algorithms: How consumers respond to offers proposed by algorithms versus humans. Journal of Retailing (2024)

  98. [107]

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. Language Models are Multilingual Chain-of-Thought Reasoners. arXiv:2210.03057 [cs.CL] https://arxiv.o...

  99. [108]

    Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger

  100. [109]

    In Proceedings of the 28th International Conference on Intelligent User Interfaces

    Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces. 410–422

  101. [110]

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhu- patiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295 (2024)

  102. [111]

    Daniel B Shank, Courtney Stefanik, Cassidy Stuhlsatz, Kaelyn Kacirek, and Amy M Belfi. 2023. AI composer bias: Listeners like music less when they think it was composed by an AI. Journal of Experimental Psychology: Applied 29, 3 (2023), 676

  103. [112]

    Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024. The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts. arXiv [cs.CL] (Jan. 2024). https://arxiv.org/abs/2401.13136

  104. [113]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)

  105. [114]

    Yaacov Trope and Nira Liberman. 2010. Construal-level theory of psychological distance. Psychological review 117, 2 (2010), 440

  106. [115]

    Minkyu Shin and Jin Kim. 2024. Large language models can enhance persuasion through linguistic feature alignment. SSRN Electronic Journal (Feb. 2024). https: //doi.org/10.2139/ssrn.4725351

  107. [116]

    Ben Sumner. 2024. More New Languages Supported in Microsoft 365 Copi- lot. https://techcommunity.microsoft.com/blog/microsoft365copilotblog/more- new-languages-supported-in-microsoft-365-copilot/4274332 Accessed: 2024- 11-15

  108. [117]

    John von Neumann, Oskar Morgenstern, and Ariel Rubinstein. 1944. Theory of Games and Economic Behavior (60th Anniversary Commemorative Edition) . Princeton University Press. http://www.jstor.org/stable/j.ctt1r2gkx

  109. [118]

    Lukas Teufelberger, Xintong Liu, Zhipeng Li, Max Moebus, and Christian Holz

  110. [119]

    arXiv preprint arXiv:2407.21593 (2024)

    LLM-for-X: Application-agnostic Integration of Large Language Models to Support Personal Writing Workflows. arXiv preprint arXiv:2407.21593 (2024)

  111. [120]

    Suzanne Tolmeijer, Ujwal Gadiraju, Ramya Ghantasala, Akshit Gupta, and Abra- ham Bernstein. 2021. Second chance for a first impression? Trust development in intelligent system interaction. In Proceedings of the 29th ACM Conference on user modeling, adaptation and personalizati...

  112. [121]

    Xiaoyi Wang and Xingyi Qiu. 2024. The positive effect of artificial intelligence technology transparency on digital endorsers: Based on the theory of mind perception. Journal of Retailing and Consumer Services 78 (2024), 103777

  113. [122]

    Azmine Toushik Wasi, Mst Rafia Islam, and Raima Islam. 2024. LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning.arXiv [cs.HC] (March 2024)

  114. [123]

    Srinivasan (Cheenu) Venkatachary. 2024. AI Overviews in Search are coming to more places around the world. https://blog.google/products/search/ai- overviews-search-october-2024/ Accessed: 2024-11-15

  115. [124]

    Jan G Voelkel, Robb Willer, et al . 2023. Artificial intelligence can persuade humans on political issues. (2023)

  116. [125]

    Walter Wymer and Hellen Gross. 2023. Charity advertising: A literature review and research agenda. Journal of Philanthropy and Marketing 28, 4 (Nov. 2023)

  117. [126]

    Alicia von Schenk, Victor Klockmann, and Nils Köbis. 2023. Social preferences toward humans and machines: a systematic experiment on the role of machine payoffs. Perspectives on Psychological Science (2023), 17456916231194949

  118. [127]

    Oana Vuculescu, Franziska Günzel-Jensen, Lars Frederiksen, and Carsten Bergenholtz. 2024. Leveling Up or Leveling Down? The Impact of Large Lan- guage Models on Student Performance in Higher Education. The Impact of Large Language Models on Student Performance in Higher Educat...

  119. [128]

    Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li. 2023. Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning. arXiv preprint arXiv:2310.03731 (2023)

  120. [129]

    Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. arXiv preprint arXiv:1908.11828 (2019)

  121. [130]

    Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A Cross- lingual Adversarial Dataset for Paraphrase Identification. In Proc. of EMNLP

  122. [131]

    Irene Weber. 2024. Large Language Models as Software Components: A Taxon- omy for LLM-Integrated Applications. arXiv preprint arXiv:2406.10300 (2024). Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages CHI ’25, April 26-May 1,...

  123. [132]

    Joel Wester, Sander De Jong, Henning Pohl, and Niels Van Berkel. 2024. Ex- ploring People’s Perceptions of LLM-generated Advice. Computers in Human Behavior: Artificial Humans (2024), 100072

  124. [133]

    Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi

  125. [134]

    Jie Xu and Guanxiong Huang. 2020. The relative effectiveness of gain-framed and loss-framed messages in charity advertising: Meta-analytic evidence and implications. International Journal of Nonprofit and Voluntary Sector Marketing 25, 4 (2020), e1675

  126. [135]

    Rongwu Xu, Brian S Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2023. The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation. arXiv [cs.CL] (Dec. 2023)

  127. [136]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al . 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024)

  128. [137]

    Yunhao Zhang and Renée Gosline. 2023. Human favoritism, not AI aversion: Peo- ple’s perceptions (and bias) toward generative AI, human experts, and human– GAI collaboration in persuasive content generation. Judgment and Decision Making 18 (2023), e41

  129. [138]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)

  130. [139]

    Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023. Low-resource languages jailbreak gpt-4. arXiv preprint arXiv:2310.02446 (2023)

  131. [140]

    Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: story writing with large language models. InProceedings of the 27th International Conference on Intelligent User Interfaces . 841–852

  132. [141]

    Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. The design and im- plementation of xiaoice, an empathetic social chatbot.Computational Linguistics 46, 1 (2020), 53–93

  133. [142]

    arXiv [cs.CL] (Jan

    How Johnny can persuade LLMs to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing LLMs. arXiv [cs.CL] (Jan. 2024)

  134. [143]

    Wenxuan Zhang, Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2024. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models. Adv. Neural Inf. Process. Syst. 36 (2024)

  135. [144]

    Xing Zhang, Xinyue Wang, Durong Wang, Quan Xiao, and Zhaohua Deng. 2024. How the linguistic style of medical crowdfunding charitable appeal influences individuals’ donations. Technol. Forecast. Soc. Change 203 (June 2024), 123394

  136. [145]

    Xiao Zhang, Ruoyu Xiang, Chenhan Yuan, Duanyu Feng, Weiguang Han, Ale- jandro Lopez-Lira, Xiao-Yang Liu, Meikang Qiu, Sophia Ananiadou, Min Peng, et al. 2024. Dólares or dollars? unraveling the bilingual prowess of financial llms between spanish and english. InProceedings of t...

  137. [148]

    Zoie Zhao, Sophie Song, Bridget Duah, Jamie Macbeth, Scott Carter, Monica P Van, Nayeli Suseth Bravo, Matthew Klenk, Kate Sick, and Alexandre LS Filipow- icz. 2023. More human than human: LLM-generated narratives outperform human-LLM interleaved narratives. In Proceedings of t...

  138. [149]

    Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-Following Evaluation for Large Language Models. arXiv:2311.07911 [cs.CL] https://arxiv.org/abs/2311. 07911

  139. [151]

    Zhao Zou, Omar Mubin, Fady Alnajjar, and Luqman Ali. 2024. A pilot study of measuring emotional response and perception of LLM-generated questionnaire and human-generated questionnaires. Scientific reports 14, 1 (2024), 2781. 8 Appendix 8.1 AI modifier prompts • Positive Narra...

  140. [2022]

    In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems

    For what it’s worth: Humans overwrite their economic self-interest to avoid bargaining with AI systems. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–18

  141. [2023]

    VOLUNTAS: International Journal of Voluntary and Nonprofit Organizations 34, 3 (2023), 626–642

    What works to increase charitable donations? A meta-review with meta- meta-analysis. VOLUNTAS: International Journal of Voluntary and Nonprofit Organizations 34, 3 (2023), 626–642

  142. [2024]

    arXiv:2403.15491 [cs.CL] https://arxiv.org/abs/2403.15491

    Open Source Conversational LLMs do not know most Spanish words. arXiv:2403.15491 [cs.CL] https://arxiv.org/abs/2403.15491

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.