{"id":"3d728b6e-e494-4c8d-9343-51c5f0470474","arxiv_id":"2501.19211","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interview study finds four user mental models of how GenAI chatbot ecosystems process data, with third-party ecosystems (ChatGPT/Expedia) understood more simply and trusted more than first-party ecosystems (Gemini/Google).","lead":"This paper interviews 21 people about how they think generative AI chatbots such as ChatGPT and Gemini work behind the scenes. It finds that people hold more consistent and trusting mental models of third-party ecosystems like ChatGPT with Expedia than of first-party ones like Google Gemini.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comparative claim is confounded: ChatGPT's condition redirected users to Expedia while Gemini completed booking in-chat, so mental-model differences may reflect UI/task flow, not first-party vs third-party ecosystem.","rationale":"The reader's weakest assumption focused on the scripted-task measurement issue, but the more fundamental problem is the non-equivalence of the two experimental conditions: ChatGPT's plugin did not complete bookings and redirected users off-platform, while Gemini completed bookings in-chat. This structural difference directly determines what participants can observe and draw, so the claimed uniformity of the Agent model for third-party ecosystems may simply mirror the visible off-platform redirect, not a stable mental model. The paper itself acknowledges this asymmetry in Section 3.5 and P15's quote in Section 4.2.4, yet the abstract and discussion still frame the result as a property of first-party vs third-party ecosystems. The reader's rationale independently identified the non-equivalent conditions as issue (2), which is why my agreement is partial rather than full. I do not recommend REJECT because the descriptive taxonomy of four mental models remains a useful qualitative contribution and the comparative claim could be salvaged by re-framing it as an observation about interface-visible data flows or by running a matched study. However, the central claim as currently stated in the abstract must be softened or re-tested, so the CONDITIONAL verdict is appropriate.","tokens_in":22794,"tokens_out":4942,"duration_ms":49080,"concrete_test":"Conduct a controlled follow-up using two matched simulated chatbot interfaces that differ only in plugin ownership: one 'first-party' (e.g., Google Hotels icon) and one 'third-party' (e.g., Expedia icon), both showing the same in-chat successful booking confirmation with no off-platform redirect. If the distribution of mental models (via the same drawing task) remains significantly different (all Agent for third-party), the original claim is supported; if the distributions converge, the original result is an artifact of the unequal task flow.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central comparative claim—that users hold more consistent and simpler mental models toward third-party ecosystems and therefore trust them more—is not identifiable from this design. Section 3.5 states that ChatGPT did not support booking hotels directly from its interface even with the Expedia plugin, while Gemini did; the study stopped participants before submission and showed a Gemini confirmation screenshot. Consequently, the ChatGPT condition literally ended with the chatbot handing the user to Expedia's external site, making the 'Agent' model (User→Chatbot→Plugin) an accurate description of the visible interaction, not a mental model of the ecosystem. The first-party condition completed the booking inside the chat, forcing participants to guess about hidden Google infrastructure, which naturally produces multiple models (Key Player, Medium, Representation, Agent). P15's quote in §4.2.4 explicitly states 'the user finishes the last step on their own whereas Gemini directly books the rooms.' The difference in mental-model diversity and in trust/concern may therefore be an artifact of the off-platform redirect and the screenshot manipulation, rather than evidence about first-party vs third-party ecosystems per se. The numeric inconsistency (Table 3 shows 20/21 Gemini concerns, §4.4 says '14 of 15') further weakens the concern counts, but the design asymmetry is the deeper problem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a semi-structured interview study with 21 participants who completed hotel-booking tasks with Google Gemini and ChatGPT, drew diagrams of the data flow, and answered follow-up questions. The authors identify four mental models of GenAI chatbot ecosystems (Key Player, Medium, Representation, Agent) centered on the chatbot's perceived role. They further claim that participants held a more consistent and simpler mental model toward the third-party ecosystem (ChatGPT/Expedia) than toward the first-party ecosystem (Gemini/Google), and that this resulted in higher trust and fewer privacy concerns toward the third-party ecosystem. The paper closes with design and policy implications for transparency and entity disclosure in GenAI ecosystems.","tokens_in":22916,"tokens_out":4984,"duration_ms":48178,"significance":"The qualitative taxonomy of four mental models is a useful contribution to HCI and privacy research on GenAI, and the drawing-based elicitation with participant quotes and diagrams provides rich, interpretable data. The full-agreement coding process and the explicit attention to data-flow perceptions are strengths. If the comparative result were robust, it would challenge common assumptions that users trust first-party services more than third-party ones. However, the central first-party/third-party comparison is currently not identifiable from the design, because the two conditions differ not only in ecosystem type but also in visible task flow: ChatGPT redirected participants to Expedia, while Gemini completed the booking in-chat. The paper is therefore best viewed as an exploratory taxonomy with a plausible but unproven association between perceived data-flow complexity and trust; the causal wording in the abstract and conclusion needs substantial revision.","major_comments":[{"comment":"The central first-party vs third-party comparison is confounded with task flow. Section 3.5 states that ChatGPT did not support booking hotels directly from its interface even with the Expedia plugin, while Gemini could complete the booking in-chat; participants were stopped before submission and shown a confirmation screenshot from Gemini (Figure 1). P15's quote in §4.2.4 explicitly describes the user finishing the last step on Expedia, so the 'Agent' model for ChatGPT accurately reflects the visible interaction with an external site rather than a mental model of a third-party ecosystem. The abstract's causal claim that third-party ecosystems 'resulted in' higher trust thus cannot be identified from this design. I recommend reframing the result as a comparison of the two specific configurations (integrated first-party booking vs off-platform redirect to Expedia) or adding an analysis that separates perceived data-flow transparency from ecosystem type.","section":"§3.5, §4.2.4, §4.3"},{"comment":"The privacy-concern counts are internally inconsistent and the discrepancy is load-bearing. Table 3 lists concerns for 20 of 21 participants in the Bard/Gemini condition (only P15 has 'No'), but §4.4 says '14 of 15 participants consistently expressed concerns'; in contrast, §4.3's '16 out of 21' for ChatGPT matches Table 3 (5 with concerns). The text should report the exact counts from Table 3 and, if a subset was analyzed (e.g., only 15 codable transcripts), state that subset explicitly. As written, the reader cannot verify the 'fewer concerns' claim.","section":"§4.4, Table 3"},{"comment":"The measurement procedure may elicit interface descriptions rather than stable mental models. The drawing task immediately followed a single scripted booking attempt using a shared lab account with fake personal information, and the final booking step was never completed; instead a pre-made Gemini confirmation screenshot was shown. Section 3.5 acknowledges that company-brand reputation influenced the models and that participants were intentionally exposed to all entities. Because participants relied on the visible Expedia icon in the ChatGPT interface (as noted in §4.3), the claimed mental models may be artifacts of the immediately preceding visual and task experience. The paper should provide evidence that the models predate the task (e.g., pre-task elicitation or questions about prior use) or temper the claim that these are users' stable mental models of chatbot ecosystems.","section":"§3.3, §3.5"}],"minor_comments":[{"comment":"Table 1 lists P2's education as 'Mater'; this appears to be a typo for 'Master'.","section":"Table 1"},{"comment":"The introduction to the mental models uses 'Representations' (plural) as the model name, whereas the model is elsewhere called 'Representation'; please standardize the terminology.","section":"§4.2"},{"comment":"The phrase 'carry out more precious user education' should likely be 'more precise user education'.","section":"§5.1"},{"comment":"Please clarify in the text whether the Gemini confirmation screenshot was shown in both the Gemini and ChatGPT conditions; the current description suggests it was taken from a Gemini pre-study test, which may have influenced the ChatGPT-condition drawings.","section":"Figure 1 and §3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper fits IUI's scope, and the four-model taxonomy is a useful qualitative contribution. The main risk is that the headline first-party/third-party result is not identifiable from the current design; I would urge the editor to require a reframing of that claim rather than a surface revision. The numeric inconsistency in §4.4 should also be corrected before any revision is considered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the GenAI mental models paper. What you should know: the four-model taxonomy is a reasonable qualitative contribution, but the headline comparative claim—users trust third-party ecosystems more than first-party—is not supported by the design as it stands. The study conditions were not equivalent, and the numbers don't line up.\n\nWhat's actually new: the authors use the established drawing-elicitation method to study GenAI chatbot ecosystems, and the resulting taxonomy (Key Player, Medium, Representation, Agent) centered on the chatbot's role is a useful way to organize users' mental models. The observation that all 21 participants held the Agent model for ChatGPT while Gemini elicited all four is striking, and the paper is transparent about many limitations, including system instability and the US-only sample. The coding process is described in detail, and the decision to skip inter-rater reliability after full agreement is defensible per their cited norms.\n\nThe soft spots are real. First, the confound: as the paper itself states in Section 3.5, ChatGPT did not support booking hotels directly from its interface even with the Expedia plugin, while Gemini did. The ChatGPT condition ended with the chatbot handing the user off to Expedia's site, so the \"Agent\" model is an accurate description of the visible interaction, not necessarily an internal mental model of the ecosystem. Conversely, Gemini completed the booking in-chat, forcing participants to guess about hidden Google infrastructure, which naturally produces diverse models. P15's quote in §4.2.4 makes this explicit: \"the user finishes the last step on their own whereas Gemini directly books the rooms.\" So the mental-model difference may be an artifact of off-platform redirect versus in-chat completion, not first-party vs third-party per se. The paper should either re-frame the claim as a hypothesis or re-analyze with this asymmetry controlled.\n\nSecond, the numeric inconsistency: Table 3 shows 20/21 participants with privacy concerns for Gemini, but Section 4.4 says \"14 of 15 participants consistently expressed concerns.\" That's a direct contradiction. Also, the binary concern coding (yes/no) has no defined rule or reliability metric, so the \"16 out of 21\" trust difference for ChatGPT is not robustly established. These are fixable, but they need fixing.\n\nWho this is for: HCI and privacy researchers working on AI transparency and usable privacy. The taxonomy is worth having, and the design/policy implications about entity disclosure are reasonable. I'd bring it to a reading group as a case study in how task design can shape measured mental models. It deserves a serious referee—the method is sound and the topic is timely—but only after major revision: fix the numbers, define the concern coding, and either soften or properly qualify the comparative claim.\n\nRecommendation: send it to peer review with a request for major revision, not desk reject.","headline":"Solid taxonomy, but the headline comparative claim is confounded by unequal task conditions and inconsistent numbers.","tokens_in":23537,"tokens_out":3391,"would_cite":false,"duration_ms":29436,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users hold four distinct mental models of AI chatbot ecosystems, and trust third-party plugins more than first-party ones.","keywords":["mental models","generative AI chatbots","privacy and security","human-computer interaction","chatbot ecosystems","trust","plugins","first-party vs third-party"],"falsifier":"Run the identical hotel-booking task with the ChatGPT interface modified so that the Expedia branding and any third-party plugin icon are removed while the functionality stays the same; if participants then no longer converge on the Agent model and no longer show higher trust and fewer concerns than with Gemini, the paper's link between visible third-party entities, mental-model consistency, and trust is not supported. A complementary falsifier is a field study where users describe how they believe their own everyday chatbot conversations move data, outside any scripted booking task; if the four-model taxonomy and the first-party/third-party asymmetry do not appear there, the results are task-specific rather than general mental models.","tokens_in":22500,"feed_emoji":"🧠","tokens_out":8441,"duration_ms":70753,"temperature":0.7,"pith_summary":"Through 21 interviews in which people drew the data-flow diagrams of two real chatbot tasks, this paper tries to establish that users do not have a single picture of how generative-AI chatbot ecosystems work: they hold four distinct mental models, each defined by the role they assign to the chatbot (key player, medium, representation of its parent company, or independent agent). The paper further claims that the mental model is not stable across ecosystems: for Google's Gemini, a first-party ecosystem, participants split across all four models, while for ChatGPT with an Expedia plugin, a third-party ecosystem, all 21 participants used the same simple Agent model. That consistency is associated with higher trust and fewer privacy concerns toward the third-party system, opposite to the usual assumption that users trust first-party services more. The paper argues this happens because the third-party interface makes the data flow visible and familiar, whereas the integrated first-party ecosystem is opaque. If correct, the finding means users' privacy decisions about AI assistants are shaped by how clearly they can see who touches their data, not by brand familiarity alone.","feed_headline":"Third-party AI plugins earn more trust than first-party ones","feed_subtitle":"Visible third-party names like Expedia make ChatGPT's data flow feel simpler and safer than Gemini's.","key_machinery":"The central object is the mental model, defined as the user's internal picture of which entities exist in the chatbot ecosystem and how personal data flows between them. The paper elicits these models with a drawing exercise: after a scripted hotel-booking interaction with each chatbot, participants drew diagrams of the entities and data flows, and the researchers coded the diagrams together with the interview transcripts. The distinguishing mechanism is the 'role of the chatbot' as the axis of classification, producing four named models (Key Player, Medium, Representation, Agent). That role variable is what carries the argument: the paper ties differences in trust and privacy concern to which role the participant assigned, and ties the assignment itself to whether the ecosystem is first-party or third-party.","core_discovery":"On its own terms, the paper offers a taxonomy and an association. The taxonomy: four mental models, centered on the chatbot's role in the ecosystem—Key Player (the chatbot collects personal information and passes it to the parent company to complete the task), Medium (the chatbot is a passive conduit to the parent company, which then works with plugins), Representation (the chatbot and the parent company are the same entity, so trust in the company transfers directly), and Agent (the chatbot forwards data directly to the plugin, with no parent company in the data flow). The association: of the 21 participants, those who interacted with Gemini as the first-party ecosystem showed all four models and all but one reported privacy concerns; when the same participants interacted with ChatGPT as the third-party ecosystem, all of them used the Agent model and 16 of 21 reported no concerns. The paper explicitly reads this as evidence that a simpler, more consistent mental model accompanies higher trust and lower concern, and that the third-party ecosystem's visible, familiar plugin (Expedia) gives users a clearer sense of where their data goes than the highly integrated first-party ecosystem (Gemini and Google Hotels).","pith_inferences":["A testable extension, not run by the paper, is to replace the Expedia plugin with an unfamiliar brand: if the trust advantage disappears, then the effect is driven by plugin reputation, not by third-party ownership per se.","An implicit consequence is that users may be disclosing more to third-party plugin ecosystems than their own understanding would support if the parent company were more visible, an asymmetry the paper does not discuss.","Because the drawing task immediately followed the booking interaction, an unstated possibility is that the 'Agent' model for ChatGPT reflects the momentary salience of the Expedia icon rather than a stable mental model users would carry into other tasks.","The taxonomy could be extended to other emerging GenAI ecosystems, such as AI assistants embedded in operating systems, where the first-party/third-party boundary is again likely to shape perceived data flow."],"forward_implications":["If users' mental models guide their privacy behavior, then a first-party ecosystem like Gemini that leaves users confused about data flow will need transparency features—such as visible entity badges or data-flow diagrams—to reach the same level of trust as a clearly intermediated third-party ecosystem.","For policy, the paper suggests regulators should require disclosure of the entities involved in a chatbot ecosystem, not just the types of data collected, because perceived entity boundaries directly affect users' concern.","For design, visible third-party branding (like the Expedia icon in ChatGPT) can act as a mental-model cue that simplifies the perceived data flow; designers of first-party ecosystems may need to create equivalent transparency cues deliberately.","For researchers, the four-model taxonomy gives a shared vocabulary for studying user understanding of AI ecosystems, analogous to prior folk-model taxonomies in security and privacy, and could be extended to other GenAI products.","The paper notes these mental models are incomplete or inaccurate relative to the true data flow, so a consequence is that users are making trust and disclosure decisions on the basis of simplified, sometimes wrong, pictures of the system."],"supporting_citations":[{"why":"Supplies the folk-models method of classifying users' security-related mental models that this study adapts.","marker":"[63]"},{"why":"Provides the drawing-elicitation technique and folk-model analysis used to surface data-flow beliefs.","marker":"[68]"},{"why":"Establishes users' mental models of the Internet and their data-flow implications, which the study extends.","marker":"[29]"},{"why":"Documents users' mental models and privacy concerns with LLM-based conversational agents, the direct empirical baseline.","marker":"[71]"},{"why":"Shows that transparency in AI systems is linked to trust, supporting the paper's interpretation of visible third-party entities.","marker":"[20]"},{"why":"Introduces the drawing exercise as a way to elicit mental models of data flow, used in the study's protocol.","marker":"[34]"}],"fun_headline_variants":["Users trust visible AI plugins over seamless integrated suites","Clear data flow boosts trust in AI chatbots, study finds","ChatGPT's simpler model wins user trust over Gemini","Mental models reveal why third-party AI tools feel safer","Four mental models explain AI chatbot privacy perceptions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a mental model measured in one scripted hotel-booking task, using a shared lab account, fake personal information, fixed prompts, and a screenshot of a successful booking, reflects the mental models people actually apply in their own daily chatbot use, rather than describing only the interface they were just shown.","fun_headline_variants_meta":{"raw":{"variants":["Users trust visible AI plugins over seamless integrated suites","Clear data flow boosts trust in AI chatbots, study finds","ChatGPT's simpler model wins user trust over Gemini","Mental models reveal why third-party AI tools feel safer","Four mental models explain AI chatbot privacy perceptions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000567,"raw_usage":{"total_tokens":2677,"prompt_tokens":925,"completion_tokens":1752,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":1678}},"tokens_in":541,"tokens_out":1752,"duration_ms":11881,"temperature":1.0,"reasoning_tokens":1678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:54:30.284696+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical hotel-booking task with the ChatGPT interface modified so that the Expedia branding and any third-party plugin icon are removed while the functionality stays the same; if participants then no longer converge on the Agent model and no longer show higher trust and fewer concerns than with Gemini, the paper's link between visible third-party entities, mental-model consistency, and trust is not supported. A complementary falsifier is a field study where users describe how they believe their own everyday chatbot conversations move data, outside any scripted booking task; if the four-model taxonomy and the first-party/third-party asymmetry do not appear there, the results are task-specific rather than general mental models.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the folk-models method of classifying users' security-related mental models that this study adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the drawing-elicitation technique and folk-model analysis used to surface data-flow beliefs."},{"cited_title":"My} data just goes {Everywhere:","cited_arxiv_id":null,"evidence_quote":"Establishes users' mental models of the Internet and their data-flow implications, which the study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that transparency in AI systems is linked to trust, supporting the paper's interpretation of visible third-party entities."},{"cited_title":"When I am on Wi-Fi, I am fearless","cited_arxiv_id":null,"evidence_quote":"Introduces the drawing exercise as a way to elicit mental models of data flow, used in the study's protocol."}],"review_version":1}