{"id":"42589cd2-53ca-437b-bfb4-b70c457e9143","arxiv_id":"2505.05660","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An experiment with 985 participants found that LLM agents using AAE or Queer slang did not increase reliance or trust over a standard English agent, and AAE speakers significantly preferred the standard English agent.","lead":"This study asked hundreds of Black American and LGBTQ+ participants whether they would rely on and trust an AI assistant more when it replied in their own sociolect, African American English or Queer slang, versus standard American English. Across 985 screened participants, the standard English assistant was generally trusted and preferred, while Queer slang users felt a stronger social presence from the slang-speaking assistant.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sociolect stimuli are not matched on perceived confidence/competence; observed reliance and preference gaps may reflect translation artifacts rather than sociolect use per se.","rationale":"The paper is a serious, well-documented mixed-methods study with large samples, within-subjects design, counterbalancing, a sociolect screener, and qualitative coding that broadly aligns with the quantitative results. The reader's conditional verdict is appropriate: the headline findings are plausible but rest on an unverified matched-cues assumption. The stimulus construction in Section 4.3 explicitly aimed to hold warmth and confidence constant across agents, but the verification pipeline only validated sociolect authenticity, not the matched communicative dimensions. Because the behavioral outcomes (reliance, trust, preference) are the paper's central contribution, a confound in the stimuli would directly threaten the claim that sociolect usage per se drives the effects rather than the particular translated phrases. The proposed rating study is a direct, feasible check: if the sociolect stimuli are rated as less confident or less competent, the causal interpretation is undermined; if they are matched, the concern is resolved. No data or code are shared, which makes independent re-analysis harder, but the concern is about experimental validity rather than internal inconsistency. I therefore recommend keeping the CONDITIONAL verdict: accept the paper's contribution as a carefully executed study whose central claim should be re-verified with matched sociolect stimuli and public materials.","tokens_in":43578,"tokens_out":2283,"duration_ms":25165,"concrete_test":"Run a preregistered rating study with a fresh Prolific panel (for example, 100 AAE speakers and 100 Queer slang speakers who pass the same screeners) who rate the final 10 AAE and 10 Queer slang stimuli and their matched SAE counterparts on perceived confidence, warmth, competence, naturalness, and formality using 7-point scales, without any video task. If the mean confidence or competence rating of the sociolect phrases differs from the matched SAE phrases by more than one scale point, the reliance and preference differences in Tables 14-16 cannot be attributed solely to sociolect use; the authors would need to re-run the study with sociolect stimuli matched on these ratings, or include the ratings as covariates in the analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the only systematic difference between the agents be the sociolect. Section 4.3 states the design goal was to create responses with 'a sense of warmth and medium confidence,' and the SAE suggestions were built from 14 warmth phrases and 14 confidence expressions taken from Zhou et al. The AAE and Queer slang stimuli were then produced by GPT-4 translation and screened only for sociolect authenticity: at least 4 of 5 verifiers had to label the phrase as the intended sociolect (Appendices J and N). No check was performed on whether the translations preserved the intended warmth, confidence, competence, or naturalness. The phrase lists in Appendices K and O show substantial divergence: AAE translations include 'Fa sho, kinda certain it's...' and 'Ima lean on it's...' for SAE 'Of course, I'm fairly certain it's...' and 'I would lean it's...'; Queer translations include 'Yasss queen! I would lean it's...' and 'Work it, diva! I would say it's...' for 'Yes! I would lean it's...' and 'Well done, I would say it's...'. These sociolect versions plausibly read as less confident, less competent, or more exaggerated, independent of sociolect recognition. The manipulation check (Appendix G) only asked whether the agent sounded like it used the sociolect, not whether the sociolect and SAE agents were matched on confidence, warmth, formality, or naturalness. Consequently, the lower reliance, trust, and satisfaction observed for the AAELM, and the weaker effects for the QSLM, could be driven by properties of the specific translated phrases rather than by sociolect usage per se. This is the most load-bearing assumption in the paper's causal interpretation of the behavioral differences.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two within-subjects user studies (498 AAE speakers and 487 Queer slang speakers; analyses on manipulation-check passers are n=399 and n=406) in which participants answered difficult video-based factual questions with suggestions from an LLM agent speaking either standard American English (SAE) or the participant's self-identified sociolect. The authors measure behavioral reliance (the proportion of trials on which the participant accepts the LLM's suggestion), perceptions (trust, satisfaction, frustration, social presence), pairwise agent preference with free-text rationales, and associations between these variables. They find that AAE speakers relied more on, trusted, preferred, and were more satisfied with the SAE agent; Queer slang speakers showed only a marginal reliance difference, no significant preference, but significantly greater social presence with the Queer slang agent and lower frustration with the SAE agent. The paper interprets these results as evidence that sociolect personalization does not straightforwardly improve reliance or perception and that reliance and social presence can diverge.","tokens_in":43914,"tokens_out":9590,"duration_ms":103838,"significance":"If the causal attribution holds, this is a timely and valuable contribution to FAccT and HCI: it provides large-sample behavioral evidence that personalizing an LLM to a user's minoritized sociolect can backfire, and that a sociolect-speaking agent can increase social presence without increasing reliance or trust. The study's strengths include its behavioral outcome measure rather than self-report alone, within-subjects design with counterbalancing, community-targeted recruitment, detailed appendices with the full stimulus sets, and mixed-methods qualitative coding that connects open-ended comments to the quantitative patterns. The main risk is that the stimuli are not verified to isolate sociolect identity from phrase-level differences in confidence, warmth, and naturalness; this must be addressed before the central claim is fully supported.","major_comments":[{"comment":"The sociolect stimuli are not validated to match their SAE counterparts on confidence, warmth, naturalness, or formality. Section 4.3 states the design goal was 'to create LLM responses that elicited a sense of warmth and medium confidence,' and the SAE items use confidence expressions with reliability percentages between 30% and 70% (Appendix H). The AAE and Queer slang phrases were generated by GPT-4 and screened only for sociolect authenticity (at least 4 of 5 verifiers; Appendices J/N), and the manipulation check (Appendix G) only asked whether the agent sounded like it used the sociolect. The final phrase lists (Tables 5 and 9) show systematic shifts in epistemic markers and warmth phrases (e.g., SAE 'I would lean it’s...' vs. AAE 'Ima lean on it’s...'; SAE 'Of course, I’m fairly certain it’s...' vs. AAE 'Fa sho, kinda certain it’s...'; SAE 'Yes! I would lean it’s...' vs. Queer 'Yasss queen! I would lean it’s...'; SAE 'Well done, I would say it’s...' vs. Queer 'Work it, diva! I would say it’s...'). These changes plausibly alter perceived confidence, competence, and naturalness independently of sociolect identity. Because the reliance measure is behavioral and answer content is withheld, the lower reliance, trust, and satisfaction for the AAELM, and the weaker effects for the QSLM, could reflect the particular translated phrases rather than sociolect usage per se. Please add a norming study in which the target populations rate both sets of phrases on confidence, warmth, naturalness, and formality, or otherwise control for these dimensions (e.g., multiple translation sets or mixed-effects models with perceived confidence as a covariate).","section":"4.3; Appendices I–K, M–O"},{"comment":"The abstract's summary statement overstates the Queer slang reliance result. It says 'both AAE and Queer slang speakers relied more on the SAE agent, and had more positive perceptions of the SAE agent,' but §5.1 and Table 14 report the Queer slang reliance difference as marginal (p = 0.085, d = 0.08) and the text explicitly states it 'did not reach statistical significance.' Likewise, for Queer slang speakers, trust and satisfaction differences are not significant (§5.2, Table 15), so 'more positive perceptions' is not supported on those measures. Please revise the abstract and any summary claims to state that AAE speakers significantly relied on, trusted, and preferred the SAE agent, while Queer slang speakers showed a marginal reliance effect and mixed perceptions (with significantly greater social presence for the Queer slang agent and lower frustration with the SAE agent).","section":"Abstract; §5.1, Table 14"},{"comment":"The Discussion generalizes from a single set of ten GPT-4-generated phrases per sociolect to conclusions about 'sociolect usage by LLMs.' Given the qualitative evidence that some participants found the AAELM's language unnatural, forced, or mocking (§5.4), the observed effects may be specific to this particular instantiation of AAE rather than to AAE as a sociolect. The paper already acknowledges in Section 7 that some phrases may sound artificial, but the main takeaway in Section 6 is phrased broadly ('personalization and anthropomorphic design of agents in this scenario can hinder...'). Please either temper the general claim to the specific stimuli used, or include multiple independently generated phrase sets per sociolect to support a stimulus-general conclusion.","section":"§6 Discussion; §4.3"}],"minor_comments":[{"comment":"The phrase 'due to the its ties with Queer individuals' contains a typo; it should be 'due to its ties.'","section":"§2.1"},{"comment":"The sentence 'we performed paired t-test comparisons between the relevant variables reported for SAELM and AAELM' should read 'SAELM and QSLM' when describing the Queer slang analysis.","section":"§5.2"},{"comment":"Appendix K says 'we achieved a diverse selection of 20 sociolect aligned phrases,' while Section 4.3 and Table 5 list 10; please make the counts consistent.","section":"Appendix K"},{"comment":"The paper repeatedly uses 'a LLM' (e.g., in the abstract); it should be 'an LLM.'","section":"Abstract and throughout"},{"comment":"Please consider adding a data availability statement and reporting confidence intervals for the reported effect sizes, since the effects are small (d ≈ 0.08–0.44) and confidence intervals would aid interpretation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong fit for FAccT and the empirical design is ambitious. My main technical concern is the stimulus confound: the sociolect phrases are not matched to their SAE counterparts on confidence and warmth, so the causal claim is not yet fully supported. I believe a norming study or reanalysis can address this, and the abstract's overstatement of the Queer slang reliance result must be corrected. I do not see a need for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is genuinely interesting: AAE speakers relied on and trusted the standard-American-English agent more than the AAE-speaking one, while Queer slang speakers felt more social presence from the Queer-slang agent without reliably preferring it. The paper is a substantial empirical study—nearly a thousand participants across two groups, within-subjects design, behavioral reliance measured by actual accept/reject decisions, and a mixed-methods follow-up. That is enough to make it worth a serious referee.\n\nThe design is generally careful. The sociolect stimuli were generated by GPT-4 and filtered by crowd panels for authenticity; the manipulation check confirmed participants perceived the intended sociolect. The qualitative coding is consistent with the quantitative patterns. I believe the core finding—that self-reported affinity can diverge from behavioral reliance—survives a careful reading.\n\nThe soft spots are real, though. The abstract says Queer slang speakers 'relied more' on the SAE agent, but that result is p=0.085, not significant; the text does report it as marginal, but the abstract overstates it. More substantively, the sociolect translations were never checked for perceived confidence, warmth, naturalness, or competence. The phrase lists in the appendix show translations like 'Fa sho, kinda certain it's...' and 'Yasss queen! I would lean it's...' for SAE counterparts; these may read as less confident or more exaggerated regardless of sociolect recognition. Since the manipulation check only asked about sociolect identity, the observed trust and reliance gaps could be due to the specific translated phrases rather than sociolect use per se. The authors acknowledge some limitations about naturalness in their Limitation section, but they don't address the matched-cues confound head-on. That's the main thing a revision would need to fix—ideally by validating that the sociolect and SAE versions are matched on warmth, confidence, and competence, or at least by explicitly arguing why the pattern of results (e.g., Queer slang social presence increasing while trust doesn't move) is hard to explain purely by phrase-level artifacts.\n\nNo data or code are shared, which limits reproducibility, but the appendix is detailed enough to reconstruct the procedure.\n\nWho this is for: HCI and FAccT researchers working on LLM personalization, dialect bias, and trust in AI. It deserves a serious referee slot; with the matched-cues issue addressed, it could be a solid publication. I'd accept it with major revision. Would I cite it? Probably yes, for the behavioral-reliance divergence finding, with a caveat about the confound.","headline":"Solid large-sample study on sociolect use in LLMs, with a real finding that behavioral reliance can diverge from self-reported affinity, but the causal claim is weakened by unmatched sociolect stimuli.","tokens_in":44447,"tokens_out":2136,"would_cite":true,"duration_ms":23757,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sociolect-speaking AI can feel warmer yet earn less trust.","keywords":["sociolect adaptation","African American English","Queer slang","LLM personalization","user reliance","trust","social presence","cultural appropriation"],"falsifier":"Run the same video-question study with two matched sets of sociolect stimuli, one set as in this paper and a second set generated by a different method or written by native speakers and pre-tested to be equal to the standard English set in warmth and confidence on the same scales. If the second set produces no reliance or trust difference between the agents, the claim that sociolect use per se changes reliance is falsified; if the second set reproduces the pattern, the phrase-generation route is exonerated.","tokens_in":43413,"feed_emoji":"🗣️","tokens_out":5593,"duration_ms":57956,"temperature":0.7,"pith_summary":"The paper asks whether an LLM agent that speaks a user's own sociolect, such as African American English or Queer slang, makes that user rely on it more and perceive it better, as personalization promises. In controlled question-answering studies with 498 AAE speakers and 487 Queer slang speakers, both groups accepted more suggestions from the standard American English agent. AAE speakers also trusted, preferred, and felt less frustrated with the standard agent. Queer slang speakers, however, felt significantly more social presence from the Queer slang agent even though they relied on it no more than the standard one. The paper concludes that sociolect adaptation can create warmth without creating reliance, and that behavioral measures can diverge from what self-reported preference and personalization would predict.","feed_headline":"Speaking users' dialect can make an AI feel warmer but less trustworthy","feed_subtitle":"AAE and Queer slang speakers relied more on the standard-English AI, while Queer users felt more connection from slang.","key_machinery":"The central mechanism is the templated LLM suggestion: 14 warmth phrases crossed with 14 confidence expressions carrying 30–70% reliability, generating 196 standard English suggestions, then translated into AAE via in-context learning from AAE corpora or into Queer slang via persona prompting, with each translation vetted by crowdsourced verification panels requiring at least 4 of 5 raters. Participants answered deliberately difficult questions about short videos and, for each question, chose between “Use LLM’s Response” and “I’ll figure it out myself”; that binary choice is the behavioral reliance measure. The within-subjects design exposes each participant to both a sociolect agent and a standard English agent, and a manipulation check ensures the participant perceived the sociolect as intended before their quantitative responses are counted. This mechanism isolates language style as the only difference while holding content, confidence expressions, and question difficulty fixed.","core_discovery":"On the paper's own terms, the central discovery is that minoritized anthropomorphic cues do change user behavior and perception, but not in the direction of simple personalization. AAE participants relied more on the standard English agent than on the AAE-speaking one ($d = 0.12$, $p < 0.05$), trusted it more, were more satisfied with it, were less frustrated by it, and preferred it outright. Queer slang participants showed no significant preference or trust difference between agents, relied on the Queer slang agent only marginally less ($p = 0.09$), felt less frustration with the standard agent, yet reported significantly greater social presence from the Queer slang agent ($d = -0.23$, $p < 0.05$). Across both groups, reliance correlated weakly but significantly with trust, satisfaction, and, for the sociolect agents, social presence, and participants relied more on whichever agent they explicitly preferred. The paper reads this as evidence that LLM adaptation to minoritized language has emotional and social effects that are partly decoupled from trust and reliance.","pith_inferences":["I would predict that in a casual, social, or entertainment task rather than a factual question-answering task, the social-presence benefit of sociolect could translate into higher reliance, reversing the direction seen here; the paper's own context-dependence findings support testing this.","A testable extension: if the sociolect suggestions were written by native speakers rather than machine-translated, the trust gap might shrink or flip, because many negative comments targeted perceived inauthenticity rather than the sociolect itself.","The divergence between AAE and Queer slang outcomes may partly reflect different acquisition paths, with AAE often a first language tied to family and community while Queer slang is often adopted later, so identity ownership and perceived encroachment likely moderate the effect; the paper gestures at this but does not test it directly.","I would infer that LLM personalization to minoritized language should be opt-in and context-aware rather than automatic, given that both groups showed some positive affect toward the sociolect agent but preferred the standard agent for the task."],"forward_implications":["Designers cannot assume that matching a user's sociolect will increase reliance or trust; in this study it moved reliance in the opposite direction for AAE speakers and left it unchanged for Queer slang speakers.","Perceived social presence and behavioral reliance can diverge: Queer slang speakers felt more connected to the sociolect agent while relying on it no more than the standard agent, so engagement metrics should not be read as trust metrics.","Testing behavioral outcomes, such as whether users accept suggestions, alongside self-reported perceptions is necessary, since self-reported preference alone would have missed part of the AAE group's reliance pattern.","The direction and size of sociolect effects differ across sociolects, with AAE and Queer slang producing different trust, preference, and social-presence profiles, so findings for one minoritized language variety should not be generalized to another without testing.","If LLM providers adapt to a user's dialect, they should weigh cultural-appropriation concerns and context, because a substantial share of participants described the sociolect output as forced, unnatural, or disrespectful."],"supporting_citations":[{"why":"Supplies the behavioral reliance measure (accept the suggestion vs. figure it out yourself) and the confidence expressions with 30–70% reliability levels.","marker":"[121]"},{"why":"Supplies the warmth-prefix generation method and the finding that warmth increases reliance on medium-confidence responses, which this study builds on.","marker":"[122]"},{"why":"Contributes the human-AI team performance paradigm and decision framing that the reliance task adapts.","marker":"[9]"},{"why":"Provides the CORAAL corpus used as in-context examples for translating standard English suggestions into AAE.","marker":"[55]"},{"why":"Provides the AAVE/SAE tweet pairs used as additional in-context translation examples for the AAE agent.","marker":"[42]"},{"why":"Documents LLM difficulty with AAE and supplies SAE-to-AAE translation examples and the bias motivation for the study.","marker":"[27]"},{"why":"Provides the social presence scale adapted for the follow-up questionnaire.","marker":"[51]"},{"why":"Provides the trust items adapted for the follow-up questionnaire.","marker":"[35]"},{"why":"Provides the frustration measurement adapted from the NASA-TLX for the follow-up questionnaire.","marker":"[48]"}],"fun_headline_variants":["Dialect AI feels warmer but gets less user trust","Standard English AI wins reliance and trust over dialect AI","Sociolect AI: more presence, but standard wins on reliance","When AI speaks your dialect, users lean on standard instead"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the machine-translated sociolect phrases differ from their standard English counterparts only in sociolect, not in perceived warmth, confidence, naturalness, or formality; if those dimensions also shifted, the lower trust and reliance could come from the specific phrases rather than from sociolect use as such.","fun_headline_variants_meta":{"raw":{"variants":["Dialect AI feels warmer but gets less user trust","Standard English AI wins reliance and trust over dialect AI","Sociolect AI: more presence, but standard wins on reliance","When AI speaks your dialect, users lean on standard instead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000874,"raw_usage":{"total_tokens":3842,"prompt_tokens":1064,"completion_tokens":2778,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":2711}},"tokens_in":680,"tokens_out":2778,"duration_ms":32157,"temperature":1.0,"reasoning_tokens":2711,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:59:35.406784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same video-question study with two matched sets of sociolect stimuli, one set as in this paper and a second set generated by a different method or written by native speakers and pre-tested to be equal to the standard English set in warmth and confidence on the same scales. If the second set produces no reliance or trust difference between the agents, the claim that sociolect use per se changes reliance is falsified; if the second set reproduces the pattern, the phrase-generation route is exonerated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the social presence scale adapted for the follow-up questionnaire."}],"review_version":1}