Pith. sign in

REVIEW 5 major objections 5 minor 68 references

Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that adding an African American voice to spoken chatbots improves how AAE-speaking users rate them, while written AAE chatbots underperform their standard-English counterparts.

desk verdict A genuinely useful controlled comparison of AAE intensity in chatbots, but the spoken benefit is mostly the voice, not the dialect text—the abstract overstates what the design can separate. read the letter →

arxiv 2501.03441 v2 pith:G373KUVQ submitted 2025-01-07 cs.CL

classification cs.CL
keywords AfricanAmericanEnglishdialectpersonalizationchatbotevaluationtext-to-speechspokendialoguesystemslinguisticsimilaritylargelanguagemodelshuman-computerinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish whether chatbots that use African American English (AAE) actually serve AAE-speaking users better than standard-English chatbots do. The authors build text chatbots by translating responses into AAE at three intensity levels with three large language models, and spoken chatbots by voicing those responses with an African American accent using text-to-speech. When AAE-speaking university students evaluated pre-recorded dialogues, the text AAE chatbots underperformed the standard-English baseline, but spoken chatbots with an African American voice and mild AAE scored higher on warmth, similarity to self, and engagement preference. The upshot, if correct, is that dialect personalization works in chatbots only when the voice carries it, and heavy dialect expression backfires.

What carries the argument

The central mechanism is the text-versus-speech contrast built from two separable components: an SAE-to-AAE translation function $E(I, D_a, D_b)$ implemented by prompting an LLM at three intensity levels (Low, Medium, High), and a text-to-speech stage using the F5 model conditioned on a reference clip of an African American speaker from the Corpus of Regional African American Language. Dialect expression is deliberately separated from response generation so that AAE changes only surface style, not content. Evaluation is static: 100 ten-turn SODA dialogues across five chatbot domains are translated, voiced, and rated by AAE-speaking judges on fifteen metrics covering dialect expression, speech quality, user alignment, and engagement. The comparison of the same content across text and speech, with dialect intensity varied, is what isolates modality as the decisive factor.

What would settle it

Run the same text-versus-speech comparison as a live interaction study in which AAE-speaking participants converse freely with the chatbots rather than watch recordings; if the spoken chatbot's advantage over text on warmth or engagement disappears or reverses, the modality-contrast claim is refuted. A more targeted check would compare the AA-accented SAE spoken chatbot against the SA-accented SAE baseline while controlling whether raters know the voice's intended ethnicity, testing whether the effect depends on explicit recognition rather than felt rapport.

Watch

Extended reading notes

Core claim

The central claim is that linguistic personalization to African American English is modality-dependent. In text, every AAE chatbot—across three LLM families and three AAE intensity levels—scored at or below the Standard American English baseline on trustworthiness, role appropriateness, and engagement preference with AAE-speaking evaluators. In speech, the same dialect content paired with an African American accent improved key outcomes: the best configuration, an AA voice speaking standard English or low-intensity AAE, outperformed the standard-voice baseline on warmth, similarity to self, communication ease, role appropriateness, and engagement preference. The effect is non-monotonic: High AAE, which mostly means heavy phonetic rewriting, lowered inoffensiveness and naturalness and fell back toward or below baseline. The paper concludes that the voice, not the dialect text, is what carries personalization benefit, and that technology limits—especially text-to-speech trained mostly on standard English—still constrain authentic AA speech generation.

Load-bearing premise

The central claim rests on the assumption that AAE-speaking university students' ratings of pre-recorded, third-party dialogues predict how the same chatbots would be experienced in live, interactive use.

Editorial extensions

If this is right

  • Spoken chatbots with an African American voice and low-intensity AAE outperform a standard-voice, standard-English baseline on warmth, similarity to self, and engagement preference for AAE-speaking evaluators.
  • Text-based AAE chatbots do not outperform the SAE baseline on any measured characteristic; higher AAE intensity makes them worse on trustworthiness, role appropriateness, and engagement.
  • Heavy AAE expression, dominated by phonetic changes, is perceived as less inoffensive, less natural, and less clear, so extreme dialect levels are counterproductive.
  • The best spoken configuration uses an AA accent with SAE or Low/Medium AAE text, indicating the accent carries most of the personalization benefit in voice interactions.
  • Modality determines whether dialect personalization helps: the same AAE content that fails in text succeeds in speech.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the modality contrast holds in live use, designers should prioritize voice and accent personalization over dialect-heavy text generation; mild AAE plus an AA voice is a safer default than heavy phonetic rewriting.
  • The paper's annotator comments suggest LLMs default to a young male AAE persona; a testable extension is whether controlling persona explicitly changes the text-chatbot results, since persona mismatch rather than dialect per se could drive underperformance.
  • The non-monotonic effect of dialect intensity predicts that an adaptive system mirroring each user's own AAE level turn by turn would outperform any fixed level; that is a direct, testable consequence the paper leaves implicit.
  • Because the TTS model is trained mostly on standard English, the clarity drop at High AAE may be a technology ceiling rather than a user preference; as accented TTS improves, the optimal dialect level could shift upward.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper develops text-based and spoken chatbots that vary in African American English (AAE) dialect intensity and, for spoken chatbots, use an African American (AA) accent. The authors evaluate these chatbots with AAE-speaking participants on metrics such as comprehension, warmth, trustworthiness, similarity to self, and engagement preference, comparing them to standard English baselines. The main reported finding is a modality contrast: text-based AAE chatbots generally underperform their SAE counterparts, while spoken chatbots with an AA accent and AAE elements outperform the SAE baseline on warmth, similarity to self, and engagement preference, with the best spoken condition pairing an AA accent with SAE text.

Significance. If the result holds, the paper would provide useful, actionable evidence for dialect personalization in conversational AI, showing that voice-based personalization can benefit AAE-speaking users even when text-based dialect rendering does not. The study is notable for systematically varying AAE intensity across three LLM families, using multi-turn dialogues from five application domains, and for centering evaluation on self-identified AAE speakers rather than on generic crowdworkers. The authors also release their code and data, which strengthens reproducibility. However, the inferential base is thin: the evaluation uses small, partially non-overlapping rater groups, no significance tests, and a single voice exemplar per accent condition, so the central modality-contrast claim is not yet established at the level the abstract states.

major comments (5)
  1. [§4.4.2, Fig. 4] The spoken benefit is not attributable to AAE elements because the best spoken condition pairs an AA accent with SAE dialect features. The abstract claims that spoken chatbots 'benefit from an African American voice and AAE elements,' but the paper does not report a direct pairwise contrast between the AA-accent/SAE condition (S) and the AA-accent/Low-AAE condition (L) on Warmth, Similarity to Self, or Engagement Preference. Without that contrast, the observed gains over the SAE baseline could be driven entirely by the AA voice.
  2. [§4.4, Figs. 3 and 4] The evaluation uses only 12 evaluators for text chatbots and 8 for spoken chatbots, and no significance tests or effect sizes are reported. Many of the confidence intervals in Figures 3 and 4 overlap substantially, so the qualitative claims that Low AAE 'enhances' warmth, similarity, and engagement preference, or that High AAE 'largely fails,' are not backed by inferential evidence. The authors should report pairwise tests or at minimum bootstrap confidence intervals for the key comparisons.
  3. [§7, Evaluator Differences] The text and spoken evaluations were conducted by different, only partially overlapping groups of evaluators, as the authors acknowledge. The manual verification that findings were 'largely consistent' is not quantified and does not address whether the text-versus-spoken contrast could be explained by rater variation. This is load-bearing because the central claim is a modality contrast, not just a within-modality effect.
  4. [§3.2, Appendix B] The spoken condition uses exactly one AA voice exemplar (ATL_se0_ag2_f_02_1) and one SA voice exemplar, so 'African American voice' is confounded with speaker identity, pitch, and prosody. The authors should either use multiple voice exemplars or explicitly frame the result as specific to the chosen voice; otherwise the claimed voice-based benefit may not generalize beyond a single speaker.
  5. [Abstract and §4.4.2] The abstract's wording overstates the findings. It says spoken chatbots benefit from 'an African American voice and AAE elements,' but Section 4.4.2 states that the most effective spoken configuration pairs an AA accent with SAE dialect features. The abstract and conclusion should be revised to say that the spoken benefit is associated with the AA voice, while the marginal contribution of AAE text elements in spoken output remains unclear.
minor comments (5)
  1. [Title] The title contains a typo: 'V oice' should be 'Voice'.
  2. [Throughout] There are inconsistent spellings of 'AAVE' as 'AA VE' and 'AAVE,' and inconsistent spacing in terms such as 'T ext' and 'V oice.' A copyedit pass would improve readability.
  3. [§4.3] The feature-tagging accuracy is reported as 91% for Claude and 86% for GPT-4o, but there is no inter-annotator agreement or error analysis for the gold test set itself. Since the test set is used to validate the tagging approach, reporting agreement would strengthen the claim that the automatic tagger is reliable.
  4. [Table 7] The 'Engagement Preference' metric asks whether the evaluator would prefer the AAE chatbot instead of the 'Original Chatbot,' but 'Original Chatbot' is not defined in the table. For spoken chatbots it is unclear whether the baseline is the SAE-accented or the SAE-dialect/SA-accent version.
  5. [§4.4.1] The claim that 'none of the studied models achieved strong representation of the grounding persona' is based on a single 'Text Persona Adherence' item, and the annotator comments about 'young male' AAE are anecdotal. Reporting representative examples and a coding scheme for those comments would make this observation more credible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims rest on external human evaluations and literature-validated tagging, not on fitted parameters or self-citation.

full rationale

This paper does not present a mathematical derivation, so there is no equation-level reduction to its own inputs. The central claims about text-based versus spoken AAE chatbots are supported by human evaluations conducted by AAE-speaking judges, which are external to the paper's own modeling choices. The AAE feature tagger is validated on a test set built from labeled examples in existing AAE literature, giving independent support for the feature measurements. The spoken chatbot comparisons use a fixed set of voices and pre-recorded dialogues, and the authors explicitly acknowledge the static-evaluation setup as a limitation; these are methodological threats to generalizability, not circular reasoning. There is no fitted parameter that is later renamed as a prediction, no load-bearing self-citation chain, and no uniqueness argument imported from prior work by the same authors. The paper's conclusions are therefore not equivalent to its inputs by construction, and no specific circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No mathematical derivation or fitted model parameters are present. The study's claims rest on design choices and evaluator assumptions listed above, plus the external validity of the static evaluation setup.

assumptions (5)
  • domain assumption The Low, Medium, and High translation prompts produce valid, ordered levels of AAE intensity.
    Section 3.1 defines the levels through prompt wording; Section 4.3 checks them with LLM tagging, but the ordering is not independently benchmarked against human judgments.
  • domain assumption Self-reported AAE background identifies people who are realistic end-users of an AAE-speaking chatbot.
    Recruitment in Section 4.4 uses a self-report form with questions about upbringing, years of use, frequency, and contexts; no external language proficiency measure is used.
  • domain assumption Static, third-party evaluation of recorded dialogues approximates live human-chatbot interaction.
    The paper acknowledges this in Section 7 and argues it is standard practice, but the central conclusions about user preference depend on the proxy.
  • domain assumption Claude-Sonnet-3.5's feature tagger, at 91% accuracy on a 90-text test set, accurately measures AAE features in generated chatbot responses.
    Section 4.3 uses the tagger to support claims about phonetic dominance in High AAE; the validation set is small and only covers 136 feature labels.
  • domain assumption The F5 TTS system with a single CORAAL reference clip produces an African American accented voice adequate for the study.
    Section 3.2 uses one clip from ATL_se0_ag2_f_02_1; the perceived persona is measured but the voice quality itself is not compared to other accents.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots." pith.science (2026). https://pith.science/paper/G373KUVQ

@misc{pith2026250103441,
  author       = {Pith},
  title        = {Pith review of: Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G373KUVQ}},
  note         = {Machine review of arXiv:2501.03441}
}
read the original abstract

As chatbots become integral to daily life, personalizing systems is key for fostering trust, engagement, and inclusivity. This study examines how linguistic similarity affects chatbot performance, focusing on integrating African American English (AAE) into virtual agents to better serve the African American community. We develop text-based and spoken chatbots using large language models and text-to-speech technology, then evaluate them with AAE speakers against standard English chatbots. Our results show that while text-based AAE chatbots often underperform, spoken chatbots benefit from an African American voice and AAE elements, improving performance and preference. These findings underscore the complexities of linguistic personalization and the dynamics between text and speech modalities, highlighting technological limitations that affect chatbots' AA speech generation and pointing to promising future research directions.

Figures

Figures reproduced from arXiv: 2501.03441 by the authors.

Figure 1
Figure 1. Overview of the dialect translation and voice generation approaches taken for Text and Spoken Chatbots. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Comparison of average per-turn rates of AAE [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Evaluation results of the 9 AA Text Chatbots (L: Low AAE, M: Medium AAE, H: High AAE). Error bars [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Evaluation results of the 4 AA Spoken Chatbots (Dialect level - S: SAE, L: Low AAE, M: Medium AAE, [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Evaluation interface for Text Chatbots [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Evaluation interface for Spoken Chatbots. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 37 canonical work pages

  1. [1]

    Eleni Adamopoulou and Lefteris Moussiades. 2020. https://doi.org/10.1016/j.mlwa.2020.100006 Chatbots: History , technology, and applications . Machine Learning with Applications, 2:100006

  2. [2]

    Vibhav Agarwal, Pooja Rao, and Dinesh Babu Jayagopi. 2021. https://doi.org/10.18653/v1/2021.nlp4convai-1.26 Towards code-mixed H inglish dialogue generation . In Proceedings of the 3rd Workshop on Natural Language Processing for Conversational AI, pages 271--280, Online. Association for Computational Linguistics

  3. [3]

    Abdulla Alsharhan, Mostafa Al-Emran, and Khaled Shaalan. 2024. https://doi.org/10.1109/TEM.2023.3298360 Chatbot Adoption : A Multiperspective Systematic Review and Future Research Agenda . IEEE Transactions on Engineering Management, 71:10232--10244

  4. [4]

    AI Anthropic. 2024. Claude 3.5 sonnet model card addendum. Claude-3.5 Model Card, 3

  5. [5]

    Gaurav Arora, Srujana Merugu, and Vivek Sembium. 2023. https://doi.org/10.18653/v1/2023.findings-acl.506 C o M ix: Guide transformers to code-mix using POS structure and phonetics . In Findings of the Association for Computational Linguistics: ACL 2023, pages 7985--8002, Toronto, Canada. Association for Computational Linguistics

  6. [6]

    Debasmita Bhattacharya, Eleanor Lin, Run Chen, and Julia Hirschberg. 2024. https://doi.org/10.21437/Interspeech.2024-1224 Switching tongues, sharing hearts: Identifying the relationship between empathy and code-switching in speech . In Interspeech 2024, pages 492--496

  7. [7]

    Su Lin Blodgett, Johnny Wei, and Brendan O ' Connor. 2018. https://doi.org/10.18653/v1/P18-1131 T witter U niversal D ependency parsing for A frican- A merican and mainstream A merican E nglish . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1415--1425, Melbourne, Australia. Assoc...

  8. [8]

    Jacqueline Brixey and David Traum. 2024. Why should a dialogue system speak more than one language? In Proceedings of the International Workshop on Spoken Dialogue Systems (IWSDS), Sapporo, Japan

Show all 68 references
  1. [9]

    Gayatri Ramamoorthy Brown. 2017. Pronoun marking in African American English-speaking children with and without specific language impairment. Louisiana State University and Agricultural & Mechanical College

  2. [10]

    Guendalina Caldarini, Sardar Jaf, and Kenneth McGarry. 2022. https://doi.org/10.3390/info13010041 A Literature Survey of Recent Advances in Chatbots . Information, 13(1):41

  3. [11]

    Ana Paula Chaves and Marco Aurelio Gerosa. 2021. https://doi.org/10.1080/10447318.2020.1841438 How Should My Chatbot Interact ? A Survey on Social Characteristics in Human – Chatbot Interaction Design . International Journal of Human–Computer Interaction, 37(8):729--758

  4. [12]

    Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, and Xie Chen. 2024. F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885

  5. [13]

    Choi, Minha Lee, and Sangsu Lee

    Yunjae J. Choi, Minha Lee, and Sangsu Lee. 2023. https://doi.org/10.1145/3544548.3581445 Toward a multilingual conversational agent: Challenges and expectations of code-mixing multilingual users . In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems,...

  6. [14]

    Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen McKeown. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.421 Evaluation of African American Language Bias in Natural Language Generation . In Proceedings of the 2023 Conference on Emp...

  7. [15]

    Los Angeles Unified School District. 2016. https://www.lausd.org/cms/lib/CA01000043/Centricity/domain/576/instruction/sel/6.3.16

  8. [16]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  9. [17]

    Matja z Ezgeta. 2012. Internal grammatical conditioning in african-american vernacular english. Maribor International Review, 5(1):9--26

  10. [18]

    Alessio Falai. 2022. https://amslaurea.unibo.it/id/eprint/25805/ Conditioning text-to-speech synthesis on dialect accent: a case study . Master's thesis, University of Bologna

  11. [19]

    Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024. https://doi.org/10.48550/arXiv.2406.08818 Linguistic Bias in ChatGPT : Language Models Reinforce Dialect Discrimination . arXiv preprint

  12. [20]

    Howard Fogel and Linnea C Ehri. 2006. Teaching african american english forms to standard american english-speaking teachers: Effects on acquisition, attitudes, and responses to student use. Journal of Teacher Education, 57(5):464--480

  13. [21]

    Eric Graves, Shreyas Aswar, Rujuta Desai, Srilekha Nampelli, Sunandan Chakraborty, and Ted Hall. 2024. Aave corpus generation and low-resource dialect machine translation. In Proceedings of the 7th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies, pages 50--59

  14. [22]

    Lisa Green. 2013. http://apics-online.info/contributions/14 African american english structure dataset . In Susanne Maria Michaelis, Philippe Maurer, Martin Haspelmath, and Magnus Huber, editors, Atlas of Pidgin and Creole Language Structures Online. Max Planck Institute for E...

  15. [23]

    Lisa J. Green. 2002. African American English: A Linguistic Introduction. Cambridge University Press

  16. [24]

    Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.473 Investigating A frican- A merican V ernacular E nglish in transformer-based text generation . In Proceedings of th...

  17. [25]

    it’s kind of like code-switching

    Christina N Harrington, Radhika Garg, Amanda Woodward, and Dimitri Williams. 2022. “it’s kind of like code-switching”: Black older adults’ experiences with a voice assistant for health information seeking. In Proceedings of the 2022 CHI Conference on Human Factors in Computing...

  18. [26]

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. https://doi.org/10.5281/zenodo.1212303 spaCy: Industrial-strength Natural Language Processing in Python

  19. [27]

    Qiushi Huang, Xubo Liu, Tom Ko, Bo Wu, Wenwu Wang, Yu Zhang, and Lilian Tang. 2024. https://doi.org/10.18653/v1/2024.findings-acl.959 Selective prompting tuning for personalized conversations with LLM s . In Findings of the Association for Computational Linguistics: ACL 2024, ...

  20. [28]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  21. [29]

    Zhijing Jin, Nils Heil, Jiarui Liu, Shehzaad Dhuliawala, Yahang Qi, Bernhard Sch \"o lkopf, Rada Mihalcea, and Mrinmaya Sachan. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.717 Implicit personalization in language models: A systematic study . In Findings of the Associ...

  22. [30]

    Allison Jones and Georgia Zellou. 2024. https://doi.org/10.3389/fcomp.2024.1436341 Voice accentedness, but not gender, affects social responses to a computer tutor . Frontiers in Computer Science, 6

  23. [31]

    Anjali Kantharuban, Jeremiah Milbauer, Emma Strubell, and Graham Neubig. 2024. Stereotype or personalization? user identity biases chatbot recommendations. arXiv preprint arXiv:2410.05613

  24. [32]

    Tyler Kendall and Charlie Farrington. 2023. https://doi.org/10.7264/1ad5-6t35 The corpus of regional african american language. version 2023.06

  25. [33]

    Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, et al. 2023. Soda: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Me...

  26. [34]

    Rickford, Dan Jurafsky, and Sharad Goel

    Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R. Rickford, Dan Jurafsky, and Sharad Goel. 2020. https://doi.org/10.1073/pnas.1915768117 Racial disparities in automated speech recognition . Proceedings of the National Ac...

  27. [35]

    Bernd Kortmann, Kerstin Lunkenheimer, and Katharina Ehret, editors. 2020. https://ewave-atlas.org/ eWAVE 3.0: The Electronic World Atlas of Varieties of English

  28. [36]

    Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, et al. 2024. Voicebox: Text-guided multilingual universal speech generation at scale. In Advances in Neural Information Processing Systems, volume 36

  29. [37]

    Yuting Liao and Jiangen He. 2020. Racial mirroring effects on human-agent interaction in psychotherapeutic conversations. In Proceedings of the 25th international conference on intelligent user interfaces, pages 430--442

  30. [38]

    Zhengyuan Liu, Stella Xin Yin, and Nancy Chen. 2024. https://doi.org/10.18653/v1/2024.sigdial-1.43 Optimizing code-switching in conversational tutoring systems: A pedagogical framework and evaluation . In Proceedings of the 25th Annual Meeting of the Special Interest Group on ...

  31. [39]

    Bei Luo, Raymond Y. K. Lau, Chunping Li, and Yain-Whar Si. 2022. https://doi.org/10.1002/widm.1434 A critical review of state-of-the-art chatbot designs and applications . WIREs Data Mining and Knowledge Discovery, 12(1):e1434

  32. [40]

    Andre Martin and Khalia Jenkins. 2024. Speaking your language: The psychological impact of dialect integration in artificial intelligence systems. Current Opinion in Psychology, page 101840

  33. [41]

    Quim Motger, Xavier Franch, and Jordi Marco. 2022. https://doi.org/10.1145/3527450 Software- Based Dialogue Systems : Survey , Taxonomy , and Challenges . ACM Comput. Surv., 55(5):91:1--91:42

  34. [42]

    Vladimir Nechaev and Sergey Kosyakov. 2024. https://arxiv.org/abs/2405.13162 Non-autoregressive real-time accent conversion model with voice cloning . Preprint, arXiv:2405.13162

  35. [43]

    David Obremski, Paula Friedrich, Nora Haak, Philipp Schaper, and Birgit Lugrin. 2022. https://doi.org/10.3389/frobt.2022.983955 The impact of mixed-cultural speech on the stereotypical perception of a virtual robot . Frontiers in Robotics and AI, 9

  36. [44]

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. https://doi.org/10.1109/ICASSP.2015.7178964 Librispeech: An asr corpus based on public domain audio books . In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), page...

  37. [45]

    Gain Park, Jiyun Chung, and Seyoung Lee. 2024. Human vs. machine-like representation in chatbot mental health counseling: the serial mediation of psychological distance and trust on compliance intention. Current Psychology, 43(5):4352--4363

  38. [46]

    Askarbek Pazylbekov, Daryn Kalym, Anuar Otynshin, and Anara Sandygulova. 2019. https://doi.org/10.1109/HRI.2019.8673232 Similarity attraction for robot's dialect in language learning using social robots . In 2019 14th ACM/IEEE International Conference on Human-Robot Interactio...

  39. [47]

    PBS. 2005. https://www.pbs.org/speak/education/curriculum/high/aae/ African american english

  40. [48]

    William Peebles and Saining Xie. 2023. https://doi.org/10.1109/ICCV.2023.00423 Scalable diffusion models with transformers . In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195--4205. IEEE

  41. [49]

    Arianna Peoples. 2023. https://www.sjsu.edu/writingcenter/docs/handouts/AAVE-Dismantling

  42. [50]

    Piercy, Gretchen Montgomery-Vestecka, and Sun Kyong Lee

    Cameron W. Piercy, Gretchen Montgomery-Vestecka, and Sun Kyong Lee. 2025. https://doi.org/10.1016/j.ijhcs.2024.103407 Gender and accent stereotypes in communication with an intelligent virtual assistant . International Journal of Human-Computer Studies, 195:103407

  43. [51]

    Claudio Santos Pinhanez, Raul Fernandez, Marcelo Carpinette Grave, Julio Nogima, and Ron Hoory. 2024. https://doi.org/10.1145/3640543.3645165 Creating an african american-sounding tts: Guidelines, technical challenges, and surprising evaluations . In Proceedings of the 29th In...

  44. [52]

    Rifki Afina Putri, Faiz Ghifari Haznitrama, Dea Adhista, and Alice Oh. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1145 Can LLM generate culturally relevant commonsense QA data? case study in I ndonesian and S undanese . In Proceedings of the 2024 Conference on Empirical...

  45. [53]

    Amon Rapp, Lorenzo Curti, and Arianna Boldi. 2021. https://doi.org/10.1016/j.ijhcs.2021.102630 The human side of human-chatbot interaction: A systematic literature review of ten years of research on text-based chatbots . International Journal of Human-Computer Studies, 151:102630

  46. [54]

    Vinotha Ravichandran, Hepsiba D, L. D. Vijay Anand, and Deepak John Reji. 2024. https://arxiv.org/abs/2401.11771 Empowering communication: Speech technology for indian and western accents through ai-powered speech synthesis . Preprint, arXiv:2401.11771

  47. [55]

    Rickford

    John R. Rickford. 1999. African American Vernacular English: Features, Evolution, Educational Implications. Blackwell Publishers, Malden, MA

  48. [56]

    Jack Sidnell. 2002. Outline of aave grammar. Retrieved April, 24:2011

  49. [57]

    Jack Sidnell. 2012. African american vernacular english (ebonics). Language Varieties

  50. [58]

    Richard L Street, Kimberly J O’Malley, Lisa A Cooper, and Paul Haidet. 2008. Understanding concordance in patient-physician relationships: personal and ethnic dimensions of shared identity. The Annals of Family Medicine, 6(3):198--205

  51. [59]

    Junko Takeshita, Shiyu Wang, Alison W Loren, Nandita Mitra, Justine Shults, Daniel B Shin, and Deirdre L Sawinski. 2020. Association of racial/ethnic and gender concordance between patients and physicians with patient experience ratings. JAMA network open, 3(11):e2024583--e2024583

  52. [60]

    Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. 2023. Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue systems. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9677--9705

  53. [61]

    Walt Wolfram. 2004. The grammar of urban african american vernacular english. Handbook of varieties of English, 2:111--32

  54. [62]

    Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. 2023. https://doi.org/10.1109/CVPR.2023.00163 Convnext v2: Co-designing and scaling convnets with masked autoencoders . In Proceedings of the IEEE/CVF Conference on Computer Vis...

  55. [63]

    Nathan I Wood. 2019. Departing from doctor-speak: a perspective on code-switching in the medical setting. Journal of general internal medicine, 34(3):464--466

  56. [64]

    Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. https://doi.org/10.18653/v1/P18-1205 Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational ...

  57. [65]

    Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022. https://doi.org/10.18653/v1/2022.acl-long.258 VALUE : U nderstanding dialect disparity in NLU . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1:...

  58. [66]

    Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. https://doi.org/10.18653/v1/2023.acl-long.44 Multi- VALUE : A framework for cross-dialectal E nglish NLP . In Proceedings of the 61st Annual Meeting of the Association for Computational ...

  59. [67]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  60. [68]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.