Pith. sign in

REVIEW 4 major objections 6 minor 5 cited by

Agentic Workflows for Conversational Human-AI Interaction Design

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Structured agentic workflows with humans in the loop can help users and designers handle ambiguity and transience in conversational human-AI interaction.

desk verdict A solid, honestly narrated RtD study with a useful three-stage workflow and a real evidence gap on the designer-facing proxy user; worth peer review after fixes. read the letter →

arxiv 2501.18002 v1 pith:LYDLPNCH submitted 2025-01-29 cs.HC

classification cs.HC
keywords human-AIinteractionAIagentsconversationalinterfacesresearch-through-designagenticworkflowsambiguitytransiencepromptrecommendation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the two central difficulties of conversational human-AI interaction—ambiguous user goals and brief, transient conversations—can be addressed by a structured agentic workflow with humans in the loop. It reports a research-through-design study in which a chat-based probe was built and iterated with ten users over four design cycles. The probe divides each interaction into three stages: contextualization, goal formulation, and prompt articulation, with AI agents recommending goals and prompts at each step. In the final version, designers can also converse with an Interactive Proxy User agent built from parsed app logs and survey responses, which stands in for a real user during early design testing. If the claim holds, designers get a concrete way to reduce user trial-and-error and to surface user needs from conversations that are otherwise too short-lived to study.

What carries the argument

The load-bearing object is a three-stage workflow populated by named agents: Contextual Persona agents (role-played personas retrieved by embedding a user's visual context and matching it against a large persona dataset), a Proxy User agent (role-played from pre-survey persona data), a Goal Refinement agent (merges and decomposes selected goals), and, in the final version, an Interactive Proxy User for designers that retrieves relevant app-log memories through cosine-similarity search. These agents do not act on the environment; they recommend goals and prompts, share cognitive and motor effort, and make the system's capabilities visible. The workflow carries the argument by turning ambiguity into explicit choices and by turning transient log data into something a designer can converse with.

What would settle it

Record interviews with real participants about what they were trying to accomplish in a logged conversation, then ask the Interactive Proxy User the same questions and compare the answers; if the proxy's answers fit the logs but contradict the participant's own account more often than not, the workflow's designer-facing value is refuted.

Watch

Extended reading notes

Core claim

The central claim is that agentic workflows with humans-in-the-loop can help designers and users navigate ambiguity and transience in conversational human-AI interaction. Concretely, a workflow that first gathers context, then helps users formulate and refine goals, and only then articulates prompts lets users clarify intention before committing to trial-and-error prompting; a designer-facing Proxy User agent that role-plays actual participants from parsed app logs lets designers interrogate transient interactions as though the users were still present. The paper grounds this in four iterations of a probe: the first suggested prompts from visual context alone, the second decoupled goal formulation from prompt articulation and added personalized Proxy Users, the third introduced user control through goal merging and decomposition, and the fourth added the designer-facing Interactive Proxy User with retrieval-augmented memory.

Load-bearing premise

The designer-facing contribution rests on the assumption that a Proxy User role-playing a participant from app logs and survey data can stand in for that real user well enough to surface genuine needs, and the paper's own authors were the only ones who tested this capability.

Editorial extensions

If this is right

  • Users who encounter a goal-formulation stage before prompt articulation can reflect and explore before trial-and-error prompting, and several reported that the separation helped them refine their intentions.
  • Presenting recommendations grouped by context type, with the generating persona described alongside, helps users judge relevance and doubles as an explainability feature.
  • Too much context can amplify ambiguity, so users should be able to filter and attach or detach specific context dimensions rather than sharing everything by default.
  • Designers can treat a conversational probe as a needfinding machine, where fulfilled conversations reveal reusable AI functionalities and failed ones reveal gaps worth deeper research.
  • An Interactive Proxy User grounded in app logs can act as a conversational search engine over usage data, complementing surveys and static log analysis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the goal-formulation stage is doing the real work, the same pattern could be lifted into any large-language-model interface: a pre-chat goal-refinement step with merging and decomposition might lower prompt-rewriting effort outside the paper's probe.
  • The Interactive Proxy User is most plausibly read as a log-summarization tool rather than a faithful simulation of a user, so its answers should be treated as queries over documented behavior, not as independent evidence about user psychology.
  • A direct A/B test is within reach: count user edits and the number of turns needed to reach a satisfactory answer with and without the structured workflow, with the paper's claims predicting fewer edits under the workflow.
  • The observation that proxies with richer logs were most useful suggests a data-threshold effect, meaning the designer-facing approach pays off mainly once longitudinal usage data has accumulated, which tempers its value for brand-new applications.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper presents a research-through-design (RtD) study of agentic workflows for conversational human-AI interaction (CHAI). The authors developed a chat-based web app as a design probe over four iterations; the final workflow comprises three stages—contextualization, goal formulation, and prompt articulation—in which agents such as Contextual Personas, Proxy Users, and Goal Refinement agents assist users. A fourth iteration adds a designer-facing Interactive Proxy User, an agent that role-plays a specific participant using parsed app logs and survey data, intended to let designers explore user needs conversationally. The paper reports an annotated portfolio of the probe versions, a reflexive thematic analysis of user feedback, and a categorization of user prompts into six types. It concludes that these agentic workflows with humans-in-the-loop can help both users and designers navigate ambiguity and transience in CHAI.

Significance. If the claims are supported, the paper contributes a concrete, transferable workflow (contextualization, goal formulation, prompt articulation) and a novel designer-facing proxy-user mechanism, both of which could reduce trial-and-error prompting and help designers access user needs in transient interactions. The paper has clear strengths: the design rationale and iteration-by-iteration reflections are detailed, the reflexive thematic analysis process is explicitly described, the annotated portfolio is well-structured, and the authors are candid about limitations such as the small language model and the exploratory nature of RtD. However, the evidence base is thin (10 self-selected participants, no baseline comparison, no external validation of the designer-facing component), and there are several reporting inconsistencies that must be corrected. The study is a worthwhile design exploration, but the central claim extends beyond what the current evidence supports.

major comments (4)
  1. [Section 4.5.3, Section 8] The claim that the workflows help designers rests entirely on the authors' self-testing of the Interactive Proxy User. The text reports: 'As authors and designers, we tested the designer-facing interface and found it functioned effectively as a conversational search engine' (Section 4.5.3). There is no evidence from external designers, no comparison with static log or survey analysis, and no check that the proxy's role-played statements match the real user's actual perspective. Because this designer-facing half is essential to the central claim ('help designers and users to navigate ambiguity and transience'), the conclusion is not yet supported. Please provide an independent evaluation or explicitly reframe the contribution as a design exploration rather than a validated designer-facing outcome.
  2. [Abstract, Section 1, Section 3.2] The participant count is internally inconsistent: the abstract and Section 3.2 state 10 participants, while Section 1 says 'iteratively designed with 16 users over four design cycles.' Section 3.2 also reports percentages (e.g., 63.6% with Gemini, 63.6% using AI multiple times daily) that are impossible for a denominator of 10. This inconsistency undermines confidence in the reported evidence base and must be resolved with a clear statement of the actual participant count and the correct denominators for all reported percentages.
  3. [Section 4.5.3] The usefulness of the designer-facing Interactive Proxy User is explicitly conditioned on rich log data: 'Proxy Users role-playing participants with the most extensive data collected were the most useful.' This suggests the method may be weakest precisely in the transient, low-data settings that the paper aims to address. The discussion should address this tension or present data on how the proxy performs for participants with sparse usage logs.
  4. [Section 1, contribution (3)] The paper promises code, prompts, and supporting materials, but the link is a placeholder ('https://github.upon.publication'). Since the final probe artifact is listed as a contribution, the repository (or a supplement) should be made available for the contribution to be assessable; if this is not possible, the claim should be revised.
minor comments (6)
  1. [Section 3.4] The citation '(Figure ??)' is unresolved; the stacked bar plots of Likert responses are referenced but no such figure is included or numbered.
  2. [Section 5.2] 'as seem in Figure 7' should read 'as seen in Figure 7'.
  3. [Section 4.3.2] The term 'super prompt' is introduced without definition or explanation.
  4. [Section 4.4.3] 'surprising relevant' should read 'surprisingly relevant'.
  5. [Section 2.2] 'role-paly prompting' is a typo for 'role-play prompting'.
  6. [Section 3.3] The phrase 'customized to each version of the probe to according to the main goals' contains a grammatical error and should be revised.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the RtD findings are interpretive accounts of user feedback and usage logs, not reductions to fitted parameters or self-cited theorems.

full rationale

This paper does not contain a derivation chain in which an output is equivalent to its inputs by construction. The central claim—that agentic workflows with humans-in-the-loop can help users and designers navigate ambiguity and transience—is supported by an annotated portfolio, reflexive thematic analysis of participant feedback, affinity diagramming of 47 logged conversations, and iterative design reflection. There are no fitted parameters later renamed as predictions, no equations whose outputs are defined in terms of the claims, and no uniqueness theorem imported from the authors' prior work. The closest self-referential element is Section 4.5.3, where the authors state, 'As authors and designers, we tested the designer-facing interface and found it functioned effectively as a conversational search engine'; this is a self-assessment of the Interactive Proxy User rather than independent validation, so it is an evidentiary limitation (and the paper's Section 7 acknowledges the exploratory scope of the study), but it is not circular: the claim about the proxy is not an input to the analysis that produced it. The manuscript also contains reporting inconsistencies, including 10 participants in the abstract and methodology versus '16 users' in the introduction, a missing figure reference, and a placeholder repository URL; these are correctness and clarity concerns, not circularity. No self-citation is load-bearing: the one author-overlapping citation ([16]) concerns social simulations and is not used to justify the present workflow claims. Overall, the findings are qualitative interpretations of data collected from the probe, making the derivation chain self-contained rather than circular.

Assumptions & free parameters 1 free parameters · 4 assumptions · 4 invented entities

The paper introduces several agent types and one hand-picked retrieval count, but these are design artifacts rather than physical entities. The main assumptions are about the centrality of ambiguity and transience, the validity of role-played personas, the sufficiency of a 10-person qualitative sample, and the adequacy of the persona dataset. No quantitative fitted parameters are used to derive the central claim.

free parameters (1)
  • Number of persona descriptions retrieved per context = 3
    Chosen by the designers in version 1 and kept in later versions; no empirical tuning or sensitivity analysis is reported.
assumptions (4)
  • domain assumption Ambiguity and transience are the two core challenges of CHAI design.
    Frames the entire RtD inquiry in Section 1. If these are not the core issues, the workflow may address the wrong targets.
  • domain assumption Role-play prompting with LLMs produces personas that are useful proxies for user behavior.
    Invoked in Sections 2.2 and 4.3. The workflow's value depends on agents producing relevant, trustworthy recommendations and proxies.
  • domain assumption Self-reported feedback and usage logs from 10 participants are sufficient to derive transferable design themes.
    Thematic analysis in Section 5.1 generalizes from a small, self-selected sample to broader CHAI design implications.
  • domain assumption The persona hub dataset and semantic embeddings provide contextually relevant personas.
    Used in versions 1 through 4 for retrieval (Section 4.2.2). The paper itself notes in Limitations that the dataset 'was not optimized for generating recommendations.'
invented entities (4)
  • Contextual Persona agent
    purpose: Generates prompt and goal recommendations based on visual context and persona descriptions.
    Introduced in version 1; implemented inside the probe. Effectiveness is only shown through user anecdotes and the authors' observations, with no external validation.
  • Proxy User agent
    purpose: Role-plays the user's persona from pre-survey data to generate personalized recommendations.
    Introduced in version 2; no external validation beyond in-study feedback and usage.
  • Goal Refinement agent
    purpose: Merges and decomposes selected goals into actionable sub-goals.
    Introduced in version 3; effectiveness is reported through user comments, not through measurable task outcomes.
  • Interactive Proxy User (designer-facing)
    purpose: Enables designers to converse with a role-played user grounded in app usage logs via RAG.
    Introduced in version 4; only the authors tested it and judged it effective, so there is no independent evidence of its reliability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Agentic Workflows for Conversational Human-AI Interaction Design." pith.science (2026). https://pith.science/paper/LYDLPNCH

@misc{pith2026250118002,
  author       = {Pith},
  title        = {Pith review of: Agentic Workflows for Conversational Human-AI Interaction Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LYDLPNCH}},
  note         = {Machine review of arXiv:2501.18002}
}
read the original abstract

Conversational human-AI interaction (CHAI) have recently driven mainstream adoption of AI. However, CHAI poses two key challenges for designers and researchers: users frequently have ambiguous goals and an incomplete understanding of AI functionalities, and the interactions are brief and transient, limiting opportunities for sustained engagement with users. AI agents can help address these challenges by suggesting contextually relevant prompts, by standing in for users during early design testing, and by helping users better articulate their goals. Guided by research-through-design, we explored agentic AI workflows through the development and testing of a probe over four iterations with 10 users. We present our findings through an annotated portfolio of design artifacts, and through thematic analysis of user experiences, offering solutions to the problems of ambiguity and transient in CHAI. Furthermore, we examine the limitations and possibilities of these AI agent workflows, suggesting that similar collaborative approaches between humans and AI could benefit other areas of design.

Figures

Figures reproduced from arXiv: 2501.18002 by the authors.

Figure 1
Figure 1. Conversational human-AI interactions can be understood as dialogue exchanges between user goals and AI function [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. User activity across iterations of the design probe. A total of 47 conversations were documented across 18 unique [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. v1: Our main hypothesis was that contextual persona’s recommendations would resonate with users due to the shared [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: v2: In response to users’ desire for more personalized recommendations, version 2 introduced recommendations from [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: v3: the main goal of the probe became to support goal formulation. We introduced additional context while keeping [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: v4: The designer-facing interface enables selecting a specific user to create a Proxy User agent for conversational [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Action prompts delegated the resolution of a problem to the virtual agent, e.g., “I need a larger, more detailed image to help me [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 7
Figure 7. Figure 7: Users interacted with the probe using prompts we categorized in action, classification, context analysis, search, [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let's Get You Hired: A Job Seeker's Perspective on Multi-Agent Recruitment Systems for Explaining Hiring Decisions

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A multi-agent LLM chatbot for job seekers was perceived by 20 interviewed participants as more actionable, trustworthy, and fair than their recalled experiences with traditional hiring methods.

  2. CoGen3D: An Agentic Human-AI Co-Design Pipeline for 3D Asset Generation for Virtual Reality

    cs.HC 2026-07 conditional novelty 5.0 of 10

    A staged conversational co-design pipeline lets non-experts produce VR-ready 3D assets that increase scene engagement and modulate affect, with a clear 2D-to-3D quality drop and no authorship leniency.

  3. BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web

    cs.MA 2025-08 unverdicted novelty 4.0 of 10

    BetaWeb promises a blockchain-enabled trustworthy agentic web, but the submitted manuscript body is a different mining-robot paper, leaving the proposal without supporting evidence.

  4. Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

    cs.AI 2025-07 reject novelty 4.0 of 10

    Hide-and-Shill is a multi-agent reinforcement learning detector for DeFi shilling that combines GRPO, LLM features, social graphs, and price reactions, but its reported F1 of 0.90 is not supported by reproducible or i...

  5. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

    cs.SE 2025-05 conditional novelty 3.0 of 10

    A qualitative taxonomy positions vibe coding and agentic coding as complementary paradigms rather than rivals in AI-assisted software development.

Reference graph

Works this paper leans on

79 extracted references · 38 canonical work pages · cited by 5 Pith papers

  1. [1]

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219 (2024)

  2. [2]

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauff- mann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905 (2024)

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  4. [4]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–13

  5. [5]

    Tyler Angert, Miroslav Ivan Suzara, Jenny Han, Christopher Lawrence Pondoc, and Hariharan Subramonyam. 2023. Spellburst: A Node-based Interface for Exploratory Creative Coding with Natural Language Prompts. UIST (2023)

  6. [6]

    Riku Arakawa, Jill Fain Lehman, and Mayank Goel. 2024. Prism-q&a: Step-aware voice assistant on a smartwatch enabled by multimodal procedure tracking and large language models. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–26

  7. [7]

    Jeffrey Bardzell, Shaowen Bardzell, Peter Dalsgaard, Shad Gross, and Kim Halskov

  8. [8]

    Kirsten Boehner, Janet Vertesi, Phoebe Sengers, and Paul Dourish. 2007. How HCI interprets the probes. In Proceedings of the SIGCHI conference on Human factors in computing systems . 1077–1086

Show all 79 references
  1. [9]

    John Bowers. 2012. The logic of annotated portfolios: communicating the value of’research through design’. In Proceedings of the designing interactive systems conference. 68–77

  2. [10]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101

  3. [11]

    Virginia Braun and Victoria Clarke. 2024. A range of ways of approaching (re- flexive) TA. Retrieved February 5, 2024 from https://www.thematicanalysis.net/ understanding-ta/

  4. [12]

    Susan E Brennan. 1990. Conversation as direct manipulation: An iconoclastic view. The art of human-computer interface design (1990), 393–404

  5. [13]

    Chia-Yuan Chang, Zhimeng Jiang, Vineeth Rakesh, Menghai Pan, Chin- Chia Michael Yeh, Guanchu Wang, Mingzhi Hu, Zhichao Xu, Yan Zheng, Mahash- weta Das, et al. 2024. MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation. arXiv preprint arXiv:2501.00332 (2024)

  6. [14]

    Victoria Clarke and Virginia Braun. 2013. Successful qualitative research: A practical guide for beginners. Successful qualitative research (2013), 1–400

  7. [15]

    Francesco Colace, Massimo De Santo, Marco Lombardi, Francesco Pascale, Anto- nio Pietrosanto, Saverio Lemma, et al. 2018. Chatbot for e-learning: A case of study. International Journal of Mechanical Engineering and Robotics Research 7, 5 (2018), 528–533

  8. [16]

    Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Srihas Rao, Arthur Caetano, Misha Sra, et al. 2024. Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory. arXiv preprint arXiv:2406.14373 (2024)

  9. [17]

    Ernest Davis. 2023. Benchmarks for automated commonsense reasoning: A survey. Comput. Surveys 56, 4 (2023), 1–41

  10. [18]

    Allan de Barcelos Silva, Marcio Miguel Gomes, Cristiano André da Costa, Rodrigo da Rosa Righi, Jorge Luis Victoria Barbosa, Gustavo Pessin, Geert De Doncker, and Gustavo Federizzi. 2020. Intelligent personal assistants: A systematic literature review. Expert Systems with Appli...

  11. [19]

    Elizabeth Diller, Ricardo Scofidio, and Diana Murphy. 2002. Blur: the making of nothing. (No Title) (2002)

  12. [20]

    Mustafa Doga Dogan, Eric J Gonzalez, Karan Ahuja, Ruofei Du, Andrea Colaço, Johnny Lee, Mar Gonzalez-Franco, and David Kim. 2024. Augmented Object Intelligence with XR-Objects. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–15

  13. [21]

    Tanja Döring, Axel Sylvester, and Albrecht Schmidt. 2013. A design space for ephemeral user interfaces. In Proceedings of the 7th International Conference on Tangible, Embedded and Embodied Interaction . 75–82

  14. [22]

    Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, et al. 2024. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Mult...

  15. [23]

    KJ Kevin Feng, Maxwell James Coppock, and David W McDonald. 2023. How Do UX Practitioners Communicate AI as a Design Material? Artifacts, Conceptions, and Propositions. In Proceedings of the 2023 ACM Designing Interactive Systems Conference. 2263–2280. https://doi.org/10.1145/...

  16. [24]

    Joel E Fischer, Stuart Reeves, Martin Porcheron, and Rein Ove Sikveland. 2019. Progressivity for voice interface design. In Proceedings of the 1st international conference on conversational user interfaces . 1–8

  17. [25]

    Asbjørn Følstad and Petter Bae Brandtzæg. 2017. Chatbots and the new world of HCI. interactions 24, 4 (2017), 38–42

  18. [26]

    William Gaver. 2012. What should we expect from research through design?. In Proceedings of the SIGCHI conference on human factors in computing systems . 937–946

  19. [27]

    William Gaver, Peter Gall Krogh, Andy Boucher, and David Chatting. 2022. Emergence as a feature of practice-based design research. In Proceedings of the 2022 ACM designing interactive systems conference . 517–526

  20. [28]

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024. Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094 (2024)

  21. [29]

    Colin M Gray, Erik Stolterman, and Martin A Siegel. 2014. Reprioritizing the relationship between HCI research and practice: bubble-up and trickle-down effects. In Proceedings of the 2014 conference on Designing interactive systems . 725–734

  22. [30]

    Siddharth Gupta, Deep Borkar, Chevelyn De Mello, and Saurabh Patil. 2015. An e-commerce website based chatbot. International Journal of Computer Science and Information Technologies 6, 2 (2015), 1483–1485

  23. [31]

    Nuria Haristiani. 2019. Artificial Intelligence (AI) chatbot as language learning medium: An inquiry. In Journal of Physics: Conference Series , Vol. 1387. IOP Publishing, 012020

  24. [32]

    Amine Ben Hassouna, Hana Chaari, and Ines Belhaj. 2024. Llm-agent-umf: Llm-based agent unified modeling framework for seamless integration of multi active/passive core-agents. arXiv preprint arXiv:2409.11393 (2024)

  25. [33]

    Jeffrey Heer. 2019. Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences 116, 6 (2019), 1844–1850

  26. [34]

    John Horner and Michael E Atwood. 2006. Effective design rationale: under- standing the barriers. Rationale management in software engineering (2006), 73–90

  27. [35]

    Anders Humlum and Emilie Vestergaard. 2024. The Adoption of ChatGPT. University of Chicago, Becker Friedman Institute for Economics Working Paper 2024-50 (2024)

  28. [36]

    Edwin L Hutchins, James D Hollan, and Donald A Norman. 1985. Direct manip- ulation interfaces. Human–computer interaction 1, 4 (1985), 311–338. Caetano et al

  29. [37]

    Florian Johannsen, Susanne Leist, Daniel Konadl, and Michael Basche. 2018. Comparison of Commercial Chatbot solutions for Supporting Customer Interac- tion. In European Conference on Information Systems. https://api.semanticscholar. org/CorpusID:56145368

  30. [38]

    Ashutosh Joshi, Sheikh Muhammad Sarwar, Samarth Varshney, Sreyashi Nag, Shrivats Agrawal, and Juhi Naik. 2024. REAPER: Reasoning based retrieval planning for complex RAG systems. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . ...

  31. [39]

    Philipp Kirschthaler, Martin Porcheron, and Joel E Fischer. 2020. What can i say? effects of discoverability in vuis on task performance and user experience. In Proceedings of the 2nd Conference on Conversational User Interfaces . 1–9

  32. [40]

    Rafal Kocielnik, Raina Langevin, James S George, Shota Akenaga, Amelia Wang, Darwin P Jones, Alexander Argyle, Callan Fockele, Layla Anderson, Dennis T Hsieh, et al. 2021. Can I talk to you about your social needs? Understanding preference for conversational user interface in ...

  33. [41]

    Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong. 2024. Better Zero-Shot Reasoning with Role-Play Prompting. In Proceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Lin...

  34. [42]

    What it wants me to say

    Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D Gordon. 2023. “What it wants me to say”: Bridging the abstraction gap between end-user programmers and code- generating large language models. In Proceedings of the 20...

  35. [43]

    Jiaojiao Ma, Pengcheng Wang, Benqian Li, Tian Wang, Xiang Shan Pang, and Dake Wang. 2024. Exploring User Adoption of ChatGPT: A Technology Accep- tance Model Perspective. International Journal of Human–Computer Interaction (2024), 1–15

  36. [44]

    Nikolas Martelaro and Wendy Ju. 2017. The needfinding machine. In Proceedings of the Companion of the 2017 ACM/IEEE International Conference on Human-Robot Interaction. 355–356

  37. [45]

    Damien Masson, Sylvain Malacria, Géry Casiez, and Daniel Vogel. 2024. Direct- gpt: A direct manipulation interface to interact with large language models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–16

  38. [46]

    D. Norman. 2013. The Design of Everyday Things: Revised and Expanded Edition . Basic Books. https://books.google.com/books?id=nVQPAAAAQBAJ

  39. [47]

    Donald A Norman. 2010. Natural user interfaces are not natural. interactions 17, 3 (2010), 6–10. https://doi.org/10.1145/1744161.1744163

  40. [48]

    Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Techno...

  41. [49]

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109 (2024)

  42. [50]

    Dr Marta Perez Garcia, Sarita Saffon Lopez, and Hector Donis. 2018. Everybody is talking about Virtual Assistants, but how are people really using them?. In Proceedings of the 32nd International BCS Human Computer Interaction Conference. BCS Learning & Development

  43. [51]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics. https://arxiv.org/abs/1908.10084

  44. [52]

    Chirag Shah and Ryen W White. 2024. Agents Are Not Enough. arXiv preprint arXiv:2412.16241 (2024)

  45. [53]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature 623, 7987 (2023), 493–498

  46. [54]

    Joongi Shin, Michael A Hedderich, Bartłomiej Jakub Rey, Andrés Lucero, and Antti Oulasvirta. 2024. Understanding Human-AI Workflows for Generating Personas. In Proceedings of the 2024 ACM Designing Interactive Systems Conference. 757–781

  47. [55]

    Marita Skjuve, Asbjørn Følstad, and Petter Bae Brandtzaeg. 2023. The user experience of ChatGPT: findings from a questionnaire study of early users. In Proceedings of the 5th international conference on conversational user interfaces . 1–10

  48. [56]

    Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati. 2024. Chain of thoughtlessness: An analysis of cot in planning. arXiv preprint arXiv:2405.04776 (2024)

  49. [57]

    Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2020. Energy and policy considerations for modern deep learning research. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 13693–13696

  50. [58]

    Hari Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–19

  51. [59]

    Rustem Turtayev, Artem Petrov, Dmitrii Volkov, and Denis Volk. 2024. Hacking CTFs with Plain Agents. arXiv preprint arXiv:2412.02776 (2024)

  52. [60]

    Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. When com- binations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour (2024), 1–11

  53. [61]

    Aditya Nrusimha Vaidyam, Hannah Wisniewski, John David Halamka, Matcheri S Kashavan, and John Blake Torous. 2019. Chatbots and conversational agents in mental health: a review of the psychiatric landscape. The Canadian Journal of Psychiatry 64, 7 (2019), 456–464

  54. [62]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reason- ing in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  55. [63]

    Luoxuan Weng, Xingbo Wang, Junyu Lu, Yingchaojie Feng, Yihan Liu, and Wei Chen. 2024. InsightLens: Discovering and Exploring Insights from Conversational Contexts in Large-Language-Model-Powered Data Analysis. arXiv preprint arXiv:2404.01644 (2024)

  56. [64]

    Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, et al . 2024. Oasis: Open agents social interaction simulations on one million agents. arXiv preprint arXiv:2411.11581 (2024)

  57. [65]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR)

  58. [66]

    Nur Yildirim, Alex Kass, Teresa Tung, Connor Upton, Donnacha Costello, Robert Giusti, Sinem Lacin, Sara Lovic, James M O’Neill, Rudi O’Reilly Meehan, et al

  59. [67]

    Nur Yildirim, Changhoon Oh, Deniz Sayar, Kayla Brand, Supritha Challa, Violet Turri, Nina Crosby Walton, Anna Elise Wong, Jodi Forlizzi, James McCann, et al. 2023. Creating design resources to scaffold the ideation of AI concepts. In Proceedings of the 2023 ACM Designing Inter...

  60. [68]

    Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fer- nanda Viégas. 2023. Investigating how practitioners use human-ai guidelines: A case study on the people+ ai guidebook. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–13

  61. [69]

    JD Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang

  62. [70]

    Nima Zargham, Leon Reicherts, Michael Bonfert, Sarah Theres Voelkel, Johannes Schoening, Rainer Malaka, and Yvonne Rogers. 2022. Understanding circum- stances for desirable proactive behaviour of voice assistants: The proactivity dilemma. In Proceedings of the 4th conference o...

  63. [71]

    Saber Zerhoudi and Michael Granitzer. 2024. PersonaRAG: Enhancing Retrieval- Augmented Generation Systems with User-Centric Agents. CoRR abs/2407.09394 (2024). https://doi.org/10.48550/ARXIV.2407.09394 arXiv:2407.09394

  64. [72]

    Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Haolin Chen, et al . 2024. xlam: A family of large action models to empower ai agent systems. arXiv preprint arXiv:2409.03215 (2024)

  65. [73]

    Lei Zhang, Jin Pan, Jacob Gettig, Steve Oney, and Anhong Guo. 2024. VRCopilot: Authoring 3D Layouts with Generative AI Models in VR. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburgh, PA, USA) (UIST ’24). Association for Compu...

  66. [74]

    Qingxiao Zheng, Yiliu Tang, Yiren Liu, Weizi Liu, and Yun Huang. 2022. UX research on conversational human-AI interaction: A literature review of the ACM digital library. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–24

  67. [75]

    John Zimmerman, Jodi Forlizzi, and Shelley Evenson. 2007. Research through design as a method for interaction design research in HCI. In Proceedings of the SIGCHI conference on Human factors in computing systems . 493–502

  68. [76]

    John Zimmerman, Erik Stolterman, and Jodi Forlizzi. 2010. An analysis and critique of Research through Design: towards a formalization of a research approach. In proceedings of the 8th ACM conference on designing interactive systems. 310–319

  69. [2016]

    In Proceedings of the 2016 ACM Conference on Designing Interactive Systems

    Documenting the research through design process. In Proceedings of the 2016 ACM Conference on Designing Interactive Systems . 96–107

  70. [2022]

    In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems

    How experienced designers of enterprise applications engage AI as a design material. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–13. https://dl.acm.org/doi/pdf/10.1145/3491102.3517491

  71. [2023]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–21

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.