{"id":"56b7021a-7a58-47a0-a684-73ded57343dd","arxiv_id":"2507.01719","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative study of 108 participants in Latin America proposes that culturally appropriate health chatbots must model relational, economic, and material contexts, presented as a Pluriversal CAI for Health framework.","lead":"Researchers ran eight participatory workshops with 108 people in Peru, Argentina, and Latin American migrant communities in the UK to explore how health chatbots should handle culture. They propose a framework, Pluriversal Conversational AI for Health, that says chatbots must account for family relationships, local logistics, economics, and politics, not just language and stereotypes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Core empirical support for the 'culture loses meaning' claim is missing: the paper asserts entanglement of culture with economic/logistical systems, but presents no within-narrative analysis showing these themes co-occur in participants' stories.","rationale":"Good faith reading: the paper is a carefully executed exploratory qualitative study. It uses appropriate participatory methods, reports positionality, follows reflexive thematic analysis, includes COREQ in an appendix, acknowledges limitations in Section 7, and links to related work. The framework is a useful synthesis. However, the central claim as stated is not only a sampling-generalization claim; it is a claim about the inadequacy of academic culture categories. The reader's weakest_assumption (sampling representativeness) is real but partly mitigated by the paper's own limitations and by the fact that the framework is presented as an initial iteration. A more load-bearing issue is construct validity: the empirical chain from workshops to 'culture loses meaning' has a missing link. Section 5's categories separate culture from constraints; Section 6's assertion of entanglement is not backed by narrative-level analysis. This is testable by re-coding. If co-occurrence is weak, the framework remains a plausible conceptual proposal but not an empirically grounded finding, and the paper's central contribution should be reframed accordingly. Thus I do not recommend rejection; conditional acceptance with a request for the co-occurrence analysis (or an explicit reframing of the claim as conceptual) is appropriate. Agreement with the reader is partial: I share concern about evidence sufficiency but locate it differently.","tokens_in":24841,"tokens_out":4000,"duration_ms":50798,"concrete_test":"Using the full code tree and transcripts (or the 682 coded excerpts), unitize the data by participant story/narrative rather than by code. For each unit, record whether at least one code from 'Constraints to healthcare access' (e.g., financial constraints, geographic/mobility constraints, lack of resources, crime/corruption) and at least one code from 'Socio-cultural ecosystem' (e.g., family interdependence, traditional practices, community dynamics) appear. Compute the observed co-occurrence rate and compare it to the rate expected under independence (e.g., chi-square or Fisher's exact test). If co-occurrence is not significantly above chance, the ground-level entanglement claim is not supported by the data as analyzed, and the framework's motivation would need to be reframed as an analytic or theoretical argument, weakening the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that at ground level culture is inextricably entangled with economics, politics, geography, and logistics, so CAI needs the proposed Pluriversal framework—is presented as a finding ('Our findings show...'), but Section 5.1 reports the thematic analysis in separate categories: 'Socio-cultural ecosystem' (family, tradition, spirituality, community) and 'Constraints to healthcare access' (finance, geography, resources, crime). The Discussion (Section 6.2.1) then asserts entanglement, illustrated by a hypothetical Cuban asthma example, not by participant data. No unit-of-analysis analysis (story/narrative level) is reported showing that material constraints and cultural themes co-occur in the same narratives. Because the framework also deliberately incorporates all elements of Liu et al.'s TCE and adds categories, it is a superset; without evidence that TCE coding fails to capture the data, the claim that academic boundaries 'lose meaning' is an interpretive assertion rather than an empirical result. This matters because the framework's added value and the paper's central contribution rest on this assertion, not on the sampling limitations, which Section 7 already acknowledges. If the claim is meant as conceptual argument, it should be labeled as such; as an empirical finding it is under-supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory qualitative study with 108 Spanish-speaking participants across eight participatory workshops in Peru, Argentina, and Latin American migrant communities in the UK. Using projective storytelling, consequence scanning, and a customized GPT-3.5-based chatbot, the authors identify themes around health conversations, barriers to healthcare access, and opportunities/risks of conversational AI. The central claim is that academic boundaries around 'culture' lose meaning at the ground level because cultural experience is entangled with economic, political, geographic, and logistical systems; the paper proposes a 'Pluriversal Conversational AI for Health' framework that encompasses three realms (ecosystems, relationships, individual) plus cross-cutting themes, extending Liu et al.'s Taxonomy of Cultural Elements (TCE).","tokens_in":25069,"tokens_out":6641,"duration_ms":69795,"significance":"If the findings and framework hold, this is a timely contribution to culturally-appropriate conversational AI for health, representing an understudied region and a bottom-up, participatory approach. Strengths include the use of COREQ reporting, explicit researcher positionality, multiple workshop sites and stakeholders, a detailed code tree in the appendix, and a thoughtful discussion linking the data to pluriversal design and conviviality. The paper also raises valuable open technical questions. However, the central empirical claim that culture is inextricably entangled with material systems is presented as a direct finding yet lacks systematic within-narrative evidence, and the framework is positioned as correcting existing taxonomies without a demonstrated coding comparison.","major_comments":[{"comment":"The paper's central claim, stated in the Abstract ('academic boundaries on notions of culture lose meaning at the ground level') and in §6.2.1 ('our data revealed that notions of culture are so entangled...'), is presented as an empirical finding. However, Section 5.1 reports the thematic analysis in separate categories: 'Socio-cultural ecosystem' and 'Constraints to healthcare access' are distinct categories, and no within-narrative or unit-of-analysis analysis is reported that would demonstrate co-occurrence of cultural and material/geographic/logistical themes within the same participant stories. The illustrative quotes in §5.1.6 (e.g., the Cardiff 'vacunadas' quote and the Huancayo discrimination quote) do show such co-occurrence, but the reader is not told how representative these are across the 682 excerpts. The Discussion illustrates the entanglement claim with a hypothetical Cuban asthma example rather than with participant data. To make this claim load-bearing, the authors should either re-analyze the transcripts at the story level and report the frequency and patterns of theme co-occurrence, or explicitly relabel the claim as a conceptual interpretation synthesized from the data rather than a direct empirical result.","section":"§5.1, §6.2.1"},{"comment":"The framework is described as incorporating 'all elements of the TCE' and adding categories; it is therefore a superset. The paper's added value, however, rests on the claim that a strict TCE-based analysis would be insufficient ('a strict and exclusive definition loses meaning on the ground'). This insufficiency is asserted in §6.2.1 and §6.2.2 but never demonstrated empirically. For example, the authors do not report a coding comparison in which TCE elements were applied to the data and shown to leave material constraints or relational dynamics as residual categories. Without such evidence, the framework is a plausible conceptual proposal, but the claim that it addresses a failure of existing taxonomies is not supported. The authors should either provide a systematic mapping (e.g., a table showing which TCE elements were insufficient and what data fell outside them) or revise the framing to present the framework as a domain-specific extension that is complementary to, rather than corrective of, the TCE.","section":"§6.2.2"}],"minor_comments":[{"comment":"The analysis is described as performed by one coder (Peters) with feedback from Da Re and Calvo; given that the paper follows reflexive thematic analysis, the single-coder approach is defensible, but the rationale should be stated explicitly, and the description of the coding rounds (62 initial codes, 682 excerpts, 107 final codes) could be expanded to clarify how the feedback from other researchers changed the code tree.","section":"§4.3.1"},{"comment":"The abstract states 'Our findings show...' without qualification, although Section 7 appropriately acknowledges that results cannot represent the rest of Latin America or even the two countries entirely; consider adding a qualifier such as 'in this sample' to the abstract and framing the implications as hypotheses to be tested in other settings.","section":"Abstract, §7"},{"comment":"Table 2 is difficult to parse because the Realm column values (Ecosystems, Relationships, Individual, Cross-cutting) are not visually distinguished from the theme names in the second column, and the caption uses 'Spheres' while the text uses 'Realms'; please reformat the table and align terminology, and ensure Figure 2 is legible in the final version.","section":"Table 2"},{"comment":"The citation of Frankfurt (1971) as an example of documented cultural biases in LLMs is puzzling, since the cited work is a philosophy paper on free will; please verify that this citation is correct and place it in the appropriate context.","section":"§2.1"},{"comment":"The quote in which a participant obtains a Cipro dosage after claiming to be a doctor demonstrates a breakdown of medical guardrails; the authors mention the female gendering of the assistant but do not comment on the safety vulnerability revealed here, which is directly relevant to the 'boundaries and limits' theme.","section":"§5.1.7"},{"comment":"In §6.2.2, 'Table 3' is referenced but the table shown is numbered Table 2; please renumber consistently.","section":"§6.2.2"},{"comment":"There are several typographical and spelling errors, including 're highlighted' in §6.2.2, 'Karusula' in §6.3, 'Hovey' in §6.2.1, and 'Caros Paz' in §5.1.5; a careful proofread is needed.","section":"Whole manuscript"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid qualitative contribution on an important and timely topic, and the authors' participatory, COREQ-compliant methods are commendable. The main concern is that the central claim about the entanglement of culture with material systems is presented as an empirical finding without the required within-narrative analysis, and the framework's positioning relative to the TCE would benefit from a systematic comparison. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid paper and worth engaging, more than the abstract lets on. The real contribution is the framework: extending Liu et al.'s Taxonomy of Cultural Elements into health-specific relational and material domains—family interdependence, mobility and logistics, power dynamics—grounded in new participatory data from Latin America. That is a real addition, not just a mash-up of existing categories. The workshops are described in unusual detail, the COREQ reporting is a genuine plus, and the links to prior LMIC health work (Jain, Karusala and Anderson, Carlo et al.) are apt. Credit where due: the limits section is honest, and the quotes do show that family, money, geography, and stigma co-occur in people's stories, not just in the analysts' categories.\n\nThe reader's stress-test concern lands, but only partially. The abstract says 'findings show' that academic boundaries on culture lose meaning, yet the thematic analysis reports 'Socio-cultural ecosystem' and 'Constraints to healthcare access' as separate categories, and the Discussion's entanglement argument leans on a hypothetical Cuban asthma example. That is a fair criticism: the claim is presented as an empirical result when it is more of an interpretive synthesis. However, the raw data are not silent on this. The Huancayo story about in-laws, gender roles, and home birth, and the Carlos Paz insurance/NGO story, weave cultural and material factors together at the narrative level. So the concern is not that the claim is unsupported—it is that the paper does not make the within-narrative entanglement explicit. A revision should either add a short analytic pass showing the co-occurrence, or label the claim as conceptual. This is fixable.\n\nOne more soft spot: single-coder analysis with team feedback is defensible under reflexive thematic analysis, so I would not demand inter-rater reliability. But with 107 codes and only illustrative quotes in the arXiv text, I could not verify how cleanly themes were separated or whether the framework's realms really map onto the code tree. A fuller codebook would help.\n\nWho is this for? HCI and health-AI people designing or evaluating culturally aware conversational systems, and NLP researchers who need qualitative grounding for cultural taxonomies. It deserves a serious referee. The framework stands as a useful mapping device, and the oversold claim is an adjustment, not a fundamental flaw. Send it to review, but ask for a tighter distinction between empirical finding and conceptual proposal.","headline":"A carefully reported exploratory study whose Pluriversal CAI for Health framework is a genuinely useful extension of the TCE; the 'culture loses meaning' claim is oversold as a finding but the entanglement it asserts is visible in the participant data.","tokens_in":25581,"tokens_out":1886,"would_cite":true,"duration_ms":24286,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Health chatbots in Latin America will need to model families and material realities, not just language and values, this participatory study argues.","keywords":["conversational AI","health chatbots","cultural appropriateness","Latin America","pluriversal design","participatory workshops","thematic analysis","majority world"],"falsifier":"A comparative user study in which Latin American users rate responses from a framework-informed health chatbot versus a conventional chatbot on identical health queries would settle it: if the framework-informed responses are not judged more culturally appropriate and trustworthy, the claim that family dynamics and material constraints are load-bearing for cultural fit fails.","tokens_in":24669,"feed_emoji":"🩺","tokens_out":8419,"duration_ms":87408,"temperature":0.7,"pith_summary":"This paper argues that making conversational AI for health truly appropriate for Latin America cannot be accomplished by adding more cultural data or fine-tuning models alone. Drawing on eight participatory workshops with 108 Spanish-speaking citizens and health and AI professionals in Peru, Argentina, and UK-based Latin American migrant communities, the authors identify sites where health chatbots are likely to misalign with local realities. Their central finding is that at ground level, culture is inseparable from economics, politics, geography, family structure, and local logistics, so a chatbot that ignores these factors will feel culturally wrong even if its language and values are tuned. To guide designers, they propose a Pluriversal Conversational AI for Health framework that maps individual, relational, and ecosystem influences and suggests that more relationality and tolerance, rather than just more data, may be needed.","feed_headline":"Health chatbots must see the family, not just the patient","feed_subtitle":"A participatory study with 108 Latin Americans concludes that health chatbots must model family and material constraints, not just more…","key_machinery":"The load-bearing device is the Pluriversal CAI for Health framework, a concentric-circle map of what a health chatbot must hold in view: the Individual (demographics, emotional state, medical history, knowledge and information needs), Relationships (family interdependence, kinship structure, responsibilities and reciprocity, community dynamics, friends and workplace), and Ecosystems (healthcare system, natural environment, mobility and infrastructure, economic and political systems, history and traditions), plus cross-cutting cultural elements adapted from an existing taxonomy and the added theme of power dynamics and discrimination. The framework is generated from qualitative workshop data—projective storytelling worksheets in which participants described third-person health scenarios, followed by role-play conversations with a persona-prompted GPT-3.5 Turbo chatbot—and analyzed through reflexive inductive thematic analysis. It does the work of converting scattered stories about health conversations into a checklist of where cultural misalignment can arise and where future CAI evaluation and benchmarking should look.","core_discovery":"The study's central claim is that culturally appropriate health CAI in the majority world requires a holistic framework because academic boundaries around 'culture' do not hold in lived health experience. Based on thematic analysis of workshop stories and discussions, the authors find that health conversations in Latin America are shaped by entities beyond the individual user: the family as a decision-making unit, community dynamics, traditional and alternative health practices, misinformation, financial constraints, crime, transport, regional disease, and food availability. They conclude that a chatbot that recommends an electronic air filter to someone without window glass and with regular power outages will be experienced as culturally incompetent, even though the failure is economic and material. The proposed Pluriversal CAI for Health framework therefore includes three realms—the Individual, Relationships, and Ecosystems—with cross-cutting themes of artefacts and technologies, concepts, norms and morals, values and beliefs, and power dynamics, and it incorporates plural notions of conviviality and relationality rather than treating cultural fit as a fixed set of trainable elements.","pith_inferences":["Editorial inference: the framework implies a practical data architecture in which chatbots maintain separate context slots for individual, relational, and ecosystem information, and can flag when relational or material context is missing rather than assuming a lone user.","Editorial inference: the relational emphasis suggests a testable design: in collectivist settings, interventions that also address the older patient's adult children or in-laws may outperform one-to-one personalized coaching, and this could be measured in a randomized trial.","Editorial inference: because all participants were Spanish-speaking and mostly reached through existing networks, the strongest test of pluriversality is replication with Indigenous-language communities and non-mestizo populations; until then, the framework's Latin America-wide reach is provisional.","Editorial inference: the paper's 'humility and tolerance' direction points to an alternative benchmark—measuring whether a chatbot that asks preference-eliciting questions and hedges assumptions is rated more appropriate in unfamiliar cultural contexts than one that tries to mirror a detected culture."],"forward_implications":["Health chatbots deployed in Latin America should treat the patient and their family as a single decision-making unit, accounting for how health decisions affect and depend on relatives.","Culturally appropriate CAI must incorporate local logistics and material conditions such as medicine availability, transport, costs, and food access, or its advice will read as foreign even if the language is perfect.","A credible chatbot could act as an on-ramp to formal care and an information bridge between appointments, but only if it is explicitly designed not to replace human care, which participants feared.","Because participants reported that patients may be more honest with a computer, CAI has a potential role in reducing stigma-related nondisclosure and guiding users toward appropriate services.","Multimodal, multimedia communication such as images, videos, stories, and WhatsApp-style delivery is likely necessary to serve mixed literacy levels and learning preferences."],"supporting_citations":[{"why":"Supplies the Taxonomy of Cultural Elements that the proposed framework adapts and extends.","marker":"(Liu et al., 2024)"},{"why":"Documents the data provenance gap, with South American organizations accounting for fewer than 0.2% of training tokens, motivating the need for bottom-up approaches.","marker":"(Longpre et al., 2025)"},{"why":"Provides the Ecuador findings on mobile health feasibility and fragmented healthcare infrastructure that the study uses to anchor CAI opportunities.","marker":"(Carlo et al., 2020)"},{"why":"Supplies the convivial-tools analysis and evidence of relational, plural health knowledges that the framework's relationality draws on.","marker":"(Karusala and Anderson, 2022)"},{"why":"Defines pluriversality, 'a world where many worlds fit,' which gives the framework its name and worldview.","marker":"(Escobar, 2018)"},{"why":"Introduces convivial tools, the notion the paper uses to argue CAI should support autonomy and mutual care rather than impose values.","marker":"(Illich, 1973)"},{"why":"Provides the case of Indian ASHA workers using WhatsApp across households, illustrating diffuse boundaries between patient, family, and community.","marker":"(Jain, 2023)"},{"why":"Scoping review identifying 'lack of adeptness with local contexts' as a common challenge for AI health interventions in low- and middle-income countries, framing the study agenda.","marker":"(Ciecierski-Holmes et al., 2022)"},{"why":"Provides the reflexive thematic analysis method used to derive the themes and framework from workshop transcripts.","marker":"(Braun and Clark, 2021)"}],"fun_headline_variants":["Health chatbots need to see the whole household, not just the patient","Chatbots for health can't ignore power outages and poverty","For Latin America, health chatbots must model family and context","Pluriversal AI: health chatbots must respect local realities","Health bots fail when they ignore home economics and power cuts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's generalizability rests on the assumption that eight workshops with 108 Spanish-speaking participants, recruited mostly through existing professional and community networks in Peru, Argentina, and UK migrant communities, surface themes stable enough to represent Latin American and majority-world health experiences.","fun_headline_variants_meta":{"raw":{"variants":["Health chatbots need to see the whole household, not just the patient","Chatbots for health can't ignore power outages and poverty","For Latin America, health chatbots must model family and context","Pluriversal AI: health chatbots must respect local realities","Health bots fail when they ignore home economics and power cuts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1270,"prompt_tokens":975,"completion_tokens":295,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":225}},"tokens_in":591,"tokens_out":295,"duration_ms":3633,"temperature":1.0,"reasoning_tokens":225,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:43:33.848467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A comparative user study in which Latin American users rate responses from a framework-informed health chatbot versus a conventional chatbot on identical health queries would settle it: if the framework-informed responses are not judged more culturally appropriate and trustworthy, the claim that family dynamics and material constraints are load-bearing for cultural fit fails.","supporting_citations":[],"review_version":1}