Pith. sign in

REVIEW 4 major objections 5 minor 71 references

Exploring the Innovation Opportunities for Pre-trained Models

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Pre-trained models win by understanding content, not generating it, according to an analysis of 85 applications.

desk verdict A useful qualitative map of what HCI researchers build with pre-trained models, but the headline 'understand content dominates' is partly an artifact of the corpus and counting. read the letter →

arxiv 2505.15790 v1 pith:4CO2L6YL submitted 2025-05-21 cs.HC cs.AI

classification cs.HCcs.AI
keywords pre-trainedmodelsgenerativeAIlargelanguageinnovationhuman-computerinteractiondesignpatternsartifactanalysiscontentunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to give innovators a map of where pre-trained models are actually creating value, using 85 HCI research applications as a stand-in for commercially successful products. Its central claim is that most of the value lies in content understanding, not generation: 69.4% (204 of 294) of the capabilities the authors extracted are 'understand content' actions such as summarizing, answering, refining, identifying, and ranking. A corollary finding is that none of the 85 applications required excellent model performance to create user value; moderate to good performance was enough. If the paper is right, product teams should look first for opportunities built on understanding existing content at moderate quality, rather than chasing generation features.

What carries the argument

The carrying mechanism is a two-part analytical scaffold. The first part is artifact analysis of 85 applications, using a capability grammar of '[action verb] + [output form or structure] + [input data]' to turn application descriptions into 294 countable capabilities, then into 33 clusters, 13 actions, and three high-level themes; this grammar is what allows the paper to claim that content understanding dominates. The second part is a task-expertise/model-performance matrix, rating each application on how hard the task is for a person (expert, typical adult, less than typical adult) and the minimum model performance needed to create user value (moderate, good, excellent); the matrix is what lets the paper claim that excellent performance is never required. The paper's third output, seven interaction design patterns spanning the applications, documents the interface solutions that repeatedly address user problems such as vague desires, difficulty expressing intent, and blank-page paralysis.

What would settle it

Count the capability themes and required performance levels of features in shipping, revenue-generating pre-trained model products: if more than half of the features in a comparable corpus of commercial products are generation capabilities requiring excellent performance, the paper's conclusion that understanding at moderate performance dominates would be overturned.

Watch

Extended reading notes

Core claim

Analyzing 85 applications drawn from the CHI and DIS literature, the paper finds that the applications researchers build with pre-trained models mostly use those models to understand content rather than to generate it. Using a capability grammar of action verb plus output form plus input data, the authors extracted 294 specific capabilities, clustered them into 33 clusters and 13 actions, and grouped the actions into three themes: generate new content, transform content, and understand content. Understand content accounts for 69.4% of the capabilities (204 of 294), while generate new content accounts for 22.4% and transform content for 8.2%; over the three years of the corpus, the main capability of applications shifted from generation toward understanding. When the applications were placed on a task-expertise/model-performance matrix, none of them required excellent model performance to create user value, and only three had task expertise below that of a typical adult. The paper presents these findings as a first resource for innovators, showing where pre-trained models can succeed and a set of seven emerging interaction design patterns that support those successes.

Load-bearing premise

The load-bearing premise is that applications published by HCI researchers at CHI and DIS are a valid proxy for commercially successful pre-trained model products, even though the authors concede this corpus never addresses financial viability or market acceptance.

Editorial extensions

If this is right

  • Product teams should start with content-understanding features: summarizing, answering, refining, finding similar, identifying, interpreting, or ranking existing content.
  • Moderate model performance is usually sufficient to create user value, so teams can ship with good-enough quality and avoid costly over-engineering toward excellent performance.
  • Leisure and education are where pre-trained model applications cluster; finance, government, transportation, and manufacturing remain open spaces this corpus does not yet cover.
  • There is a largely unexplored opportunity to use pre-trained models to produce many tailored versions of a single artifact for many people, a space between mass manufacturing and craft.
  • The seven interaction design patterns offer ready-made starting points for human-AI interaction, each solving a specific problem like helping users express vague desires or overcome blank-page paralysis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's distribution holds, the public label 'generative AI' is strategically misleading: the value layer is often the understanding layer that supports or precedes generation, so funding and engineering should be allocated accordingly.
  • The near-total absence of sensor, time-series, and graph data in the corpus suggests pre-trained models are not yet substituting for narrow AI in low-level sensing and forecasting, leaving a defensible niche for traditional approaches.
  • A direct test of the proxy would be to repeat the capability analysis on app-store lists of popular AI consumer products; such a study could confirm or overturn the research-corpus assumption, but that test is our suggestion, not the paper's.
  • The 'many things for many people' gap suggests that personalization-at-scale may be the next wave of pre-trained model applications, even though the paper does not develop this direction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper aims to help innovators identify low-risk opportunities for building products and services with pre-trained models. The authors analyze 85 applications drawn from CHI and DIS full papers between 2022 and 2024, using artifact analysis to extract 294 individual capabilities, cluster them into 33 capability clusters, 13 specific actions, and three high-level themes, and then rate each application on task expertise and minimum model performance. They report that 69.4% of capabilities are content-understanding capabilities, that no application required excellent model performance, and they propose seven interaction design patterns. The paper frames these findings as a resource for innovators, with content understanding as a promising starting place for innovation at moderate-to-good performance thresholds.

Significance. If the empirical claims hold, the paper would provide a useful, systematically organized map of where pre-trained model applications have been demonstrated in HCI research, complementing prior work on narrow AI innovation. The strengths are the transparent search procedure, the full appendix listing of 85 artifacts and 294 capabilities, the explicit comparison to the Yildirim et al. task-expertise/model-performance matrix, and the candid discussion of limitations. The main significance, however, rests on the proxy assumption that CHI/DIS research applications are a valid stand-in for commercially successful applications; the paper itself acknowledges this is imperfect, and the conclusions about capability prevalence and performance thresholds are directly sensitive to that assumption and to the counting methodology.

major comments (4)
  1. [§3.1] The corpus selection is the load-bearing foundation for the headline findings, but the inclusion criteria are not fully reproducible. The paper reports searching ACM SIGCHI databases with six terms, obtaining 1140 publications, then filtering 'full papers' to 196, then identifying '85 full papers that designed applications with pre-trained models,' but the step from 196 to 85 is described only via a flow diagram with no operationalization of what counts as 'designs applications' or how disagreements were resolved. Because the 69.4% content-understanding figure and the 'none required excellent performance' claim are corpus-level statistics, a different reading of the inclusion criterion could change both. The authors should provide a coding protocol, dual-coding with inter-rater reliability, or at least a detailed list of excluded papers with reasons.
  2. [§4.2] The capability counts are at the level of individual capabilities, not applications, which can inflate the share of 'Understand Content.' Each application contributes multiple capabilities, and an image-generation application can contribute one 'Render' capability plus several 'Summarize,' 'Refine,' or 'Identify' subcapabilities, even though the user-facing value may be the generated artifact. The paper's central claim that 'more than half of capabilities, 69.4%... fit understand content' is vulnerable to this granularity choice. The authors should report the distribution of primary or main capabilities per application, or weight the capabilities by their role in the application, to test whether the dominance of content understanding persists.
  3. [§4.3] The claim that 'none of the applications required excellent model performance' is an expected consequence of the CHI/DIS venue filter rather than evidence about the commercial opportunity space. CHI and DIS reward novel user-facing interaction and feasible demonstrations; applications requiring high-stakes performance—medical imaging, fraud detection, autonomous systems—are likely to be unpublished or published in venues with different criteria. The comparison with Yildirim et al.'s narrow AI matrix in Figure 4a is confounded by venue and time. The authors should soften the prescriptive reading of this finding or provide external evidence, such as an analysis of commercial or industry applications, that moderate-to-good performance is indeed where pre-trained model value predominantly lies.
  4. [§5.1 and §6] There is a tension between the headline takeaway and the paper's own discussion. Section 5.1 states that 'Most of the applications we analyzed create value for users by helping them generate or improve new artifacts made up largely of text and images,' which is difficult to reconcile with the claim that content understanding is the dominant source of value. The paper then attributes the domination to the fact that generation requires understanding, but this reframes the finding: value may still be generated through generation, with understanding as a supporting component. Section 6 correctly acknowledges the 'huge blind spot in terms of financial risks' in the HCI corpus, but this acknowledgment is not carried into the presentation of the headline claims. The authors should either restate the conclusions to reflect the value-through-generation interpretation or provide evidence that the understanding capabilities themselves are what users value.
minor comments (5)
  1. [§3.2.3] The paper states that task-expertise and model-performance ratings are 'collective, subjective inferences' with no reported inter-rater reliability or codebook; adding a short reliability check or at least a definition of the rating scale would strengthen the reproducibility of the matrix in Figure 4b.
  2. [§4.2] There is a broken figure reference in Section 4.4.3: 'The CreativeConnect (Figure ??)' should be replaced with the correct figure number.
  3. [Throughout] Several typographical errors appear, including 'Yilidirim' for 'Yildirim', 'cateogrized' for 'categorized', 'dUsers' at the start of Section 4.4.2, and 'summmarize' in Table 4; these should be corrected in a final pass.
  4. [§4.2.1] The percentages for data types do not sum to a clear total because an application can use multiple input or output types; a note clarifying that categories are not mutually exclusive would improve interpretability.
  5. [§6] The limitations section acknowledges reliance on Yildirim et al. for domain and capability structures, but the paper does not discuss how this reliance might have biased the capability taxonomy; a sentence or two on this would help calibrate the reader's confidence in the 33 capability clusters.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the 69.4% and 'no excellent performance' findings are corpus observations from inductive coding, not fitted inputs or conclusions imported from self-cited prior work.

full rationale

The paper's central quantitative claims are empirical counts and subjective inferences from its own corpus, not parameters fitted to produce those claims. Section 4.2 reports 'More than half of capabilities, 69.4% (204 of 294) fit understand content' after the authors 'first detailed specific capabilities for each application. This resulted in 294 specific capabilities' and then clustered them bottom-up (Section 3.2.2); the percentage is a count, not a construction. Section 4.3 reports 'none of the applications required excellent model performance' after 'we made collective, subjective inferences for task-expertise and for model-performance' (Section 3.2.3); this is an acknowledged judgment about the corpus, not a renamed fit. The choice of CHI/DIS research applications as a proxy is a selection and representativeness assumption, which the authors explicitly flag in Section 6: 'the HCI research corpus has a huge blind spot in terms of financial risks.' Selection bias is a validity concern, not a circularity. The main caveat is the heavy acknowledged reliance on Yildirim et al. [65], whose co-author list includes Jodi Forlizzi: 'our research heavily relies on a single prior work—specifically, Yildirim et al.'s study on AI capabilities' (Section 6). That prior work supplies the domain list, capability-analysis process, and task-expertise/model-performance matrix, but it is an external empirical study of 40 commercial AI features; the present paper's headline distributions are computed independently from the 85-application corpus and are not equivalent to [65]'s outputs. The self-citation overlap therefore does not make the central claim reduce to its own inputs, so no circular step is identified. Score 2 reflects the acknowledged self-citation overlap without treating it as load-bearing circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on three qualitative assumptions: the proxy validity of HCI research applications, the coverage of the search strategy, and the reliability of artifact analysis. No numeric free parameters are fitted, and no new entities are invented.

assumptions (3)
  • domain assumption HCI research applications from CHI and DIS are a valid proxy for commercially successful pre-trained model applications.
    Section 3.1 states this explicitly; the authors acknowledge in Section 6 that financial viability is almost never addressed, making the proxy imperfect. If research applications do not reflect commercial success, the resource loses its stated purpose.
  • domain assumption The search terms and venues (CHI, DIS) capture the population of pre-trained model applications relevant for innovation.
    Section 3.1 defines the search; excluding UIST and CSCW, and limiting to papers with specific keywords, may miss relevant applications and bias domain and capability distributions.
  • domain assumption Artifact analysis of published papers reliably reveals user needs, model performance, and value creation.
    Section 3.2 introduces the method; the ratings are 'collective, subjective inferences' (Section 3.2.3), so the analysis assumes that reading papers is sufficient to infer these qualities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring the Innovation Opportunities for Pre-trained Models." pith.science (2026). https://pith.science/paper/4CO2L6YL

@misc{pith2026250515790,
  author       = {Pith},
  title        = {Pith review of: Exploring the Innovation Opportunities for Pre-trained Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4CO2L6YL}},
  note         = {Machine review of arXiv:2505.15790}
}
read the original abstract

Innovators transform the world by understanding where services are successfully meeting customers' needs and then using this knowledge to identify failsafe opportunities for innovation. Pre-trained models have changed the AI innovation landscape, making it faster and easier to create new AI products and services. Understanding where pre-trained models are successful is critical for supporting AI innovation. Unfortunately, the hype cycle surrounding pre-trained models makes it hard to know where AI can really be successful. To address this, we investigated pre-trained model applications developed by HCI researchers as a proxy for commercially successful applications. The research applications demonstrate technical capabilities, address real user needs, and avoid ethical challenges. Using an artifact analysis approach, we categorized capabilities, opportunity domains, data types, and emerging interaction design patterns, uncovering some of the opportunity space for innovation with pre-trained models.

Figures

Figures reproduced from arXiv: 2505.15790 by the authors.

Figure 1
Figure 1. The number of papers presenting applications made [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Illustrating the search, filtering, inclusion, and ex [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Mapping of pretrained model applications to indus [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Task-expertise and Model-performance (e.g., relationships, clusters), or sensor data (e.g., motion, non-vocal sound, humidity, radar) as input or output, even though these are frequently used for narrow AI systems. 4.3 Task-expertise and Model-performance We plotted th…
Figure 5
Figure 5. Figure 5: Examples of Chatbot Interview Examples from Our Corpus : The Selenite application helps users make a purchase decision for products that they are new to or unfamiliar with (Figure 6a) [37]. It reveals the criteria commonly used by prior users when completing this task.…
Figure 6
Figure 6. Figure 6: Examples of Reveal Dimensions (a) Photoscout (Artifact Number: 38, CHI’ 24) [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Examples of Something Like This know what they want when they see it. The Dessert Cart pattern provides users with several versions (images, color palettes, stories, or documents) that they can choose from. By offering a variety of options, the Dessert Cart pattern hel…
Figure 9
Figure 9. Figure 9: Examples of Refine This Examples from Our Corpus : ChatScratch (Figure 10a) [10] is a learning tool which teaches programming through interactive storyboards and digital drawings. Based on an initial sketch created by a child, ChatScratch creates a polished version. Pr…
Figure 10
Figure 10. Figure 10: Examples of Complete This (a) GlassMail (Artifact Number: 70, DIS’ 24) (b) DynaVis (Artifact Number: 29, CHI’ 24) [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 12
Figure 12. Figure 12: The overview of when our insights can be useful to innovators. Yang et al,[ [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 58 canonical work pages

  1. [1]

    Philip Adeoye. 2023. Artifact Analysis. Retrieved Jan 30, 2023 from https: //philipadeoye.com/100_days_of_ux/artifact_analysis.html#:~:text=Artifact% 20Analysis%20is%20the%20study,in%20which%20it%20typically%20exists

  2. [2]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–13

  3. [3]

    Apple. 2025. Human Interface Guidelines: Machine Learning. Retrieved 2025 from https://developer.apple.com/design/human-interface-guidelines/ technologies/machine-learning/introduction/

  4. [4]

    Phillip Areeda and Donald F Turner. 1975. Predatory pricing and related practices under Section 2 of the Sherman Act. J. Reprints Antitrust L. & Econ. 6 (1975), 219

  5. [5]

    Celeste Barnaby, Qiaochu Chen, Chenglong Wang, and Isil Dillig. 2024. Photo- Scout: Synthesis-Powered Multi-Modal Image Search. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15

  6. [6]

    Sara Bly and Elizabeth F Churchill. 1999. Design through matchmaking: technol- ogy in search of users. interactions 6, 2 (1999), 23–31

  7. [7]

    Jan O Borchers. 2000. A pattern approach to interaction design. In Proceedings of the 3rd conference on Designing interactive systems: processes, practices, methods, and techniques. 369–378

  8. [8]

    Businesswire. 2021. Gartner Identifies Key Emerging Technologies Spurring Innovation Through Trust, Growth and Change . Retrieved August, 2021 from https://www.businesswire.com/news/home/20210823005367/en/Gartner- Identifies-Key-Emerging-Technologies-Spurring-Innovation-Through-Trust- Growth-and-Change

Show all 71 references
  1. [9]

    Runze Cai, Nuwan Janaka, Yang Chen, Lucia Wang, Shengdong Zhao, and Can Liu. 2024. PANDALens: Towards AI-Assisted In-Context Writing on OHMD Dur- ing Travels. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–24

  2. [10]

    Liuqing Chen, Shuhong Xiao, Yunnong Chen, Yaxuan Song, Ruoyu Wu, and Lingyun Sun. 2024. ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–19

  3. [11]

    DaEun Choi, Sumin Hong, Jeongeon Park, John Joon Young Chung, and Juho Kim. 2024. CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–25

  4. [12]

    John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching stories with generative pretrained language models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–19

  5. [13]

    Ozgur Dedehayir and Martin Steinert. 2016. The hype cycle model: A review and future directions. Technological Forecasting and Social Change 108 (2016), 28–41

  6. [14]

    C DiSalvo. 2002. All Robots Are Not Created Equal: The Design and Perception of Humanoid Robot Heads. Human Computer Interaction Institute and school of Design, Carnegie Mellon University (2002)

  7. [15]

    T Dotan and D Seetharaman. 2023. Big Tech struggles to turn AI hype into profits. The Wall Street Journal 1, 1 (2023), 1

  8. [16]

    Graham Dove, Kim Halskov, Jodi Forlizzi, and John Zimmerman. 2017. UX design innovation: Challenges for working with machine learning as a design material. In Proceedings of the 2017 chi conference on human factors in computing systems . 278–288

  9. [17]

    Peter F Drucker et al. 2002. The discipline of innovation. Harvard business review 80, 8 (2002), 95–102

  10. [18]

    KJ Feng, Q Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X Zhang, and David W McDonald. 2024. Canvil: Designerly Adaptation for LLM-Powered User Experiences. arXiv preprint arXiv:2401.09051 (2024)

  11. [19]

    KJ Kevin Feng, Maxwell James Coppock, and David W McDonald. 2023. How Do UX Practitioners Communicate AI as a Design Material? Artifacts, Conceptions, and Propositions. In Proceedings of the 2023 ACM Designing Interactive Systems Conference. 2263–2280

  12. [20]

    Frederic Gmeiner and Nur Yildirim. 2023. Dimensions for Designing LLM-based Writing Support. In In2Writing Workshop at CHI

  13. [21]

    Juhye Ha, Hyeon Jeon, Daeun Han, Jinwook Seo, and Changhoon Oh. 2024. CloChat: Understanding How People Customize, Interact, and Experience Per- sonas in Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–24

  14. [22]

    Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Yuan Yao, Ao Zhang, Liang Zhang, et al. 2021. Pre-trained models: Past, present and future. AI Open 2 (2021), 225–250

  15. [23]

    Bruce Hanington and Bella Martin. 2019. Universal methods of design expanded and revised: 125 Ways to research complex problems, develop innovative ideas, and design effective solutions. Rockport publishers

  16. [24]

    Yihan Hou, Manling Yang, Hao Cui, Lei Wang, Jie Xu, and Wei Zeng. 2024. C2Ideas: Supporting Creative Interior Color Design Ideation with a Large Lan- guage Model. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–18

  17. [25]

    Kristen Howell, Gwen Christian, Pavel Fomitchov, Gitit Kehat, Julianne Marzulla, Leanne Rolston, Jadin Tredup, Ilana Zimmerman, Ethan Selfridge, and Joseph Bradley. 2023. The economic trade-offs of large language models: A case study. arXiv preprint arXiv:2306.07402 (2023)

  18. [26]

    HAOMIAO HUANG. 2023.The generative AI revolution has begun—how did we get here? Retrieved Jan 30, 2023 from https://arstechnica.com/gadgets/2023/01/the- generative-ai-revolution-has-begun-how-did-we-get-here/

  19. [27]

    IBM. 2022. Design for AI. Retrieved 2022 from https://www.ibm.com/design/ai/

  20. [28]

    2017.Things that keep us busy: The elements of interaction

    Lars-Erik Janlert and Erik Stolterman. 2017.Things that keep us busy: The elements of interaction. MIT Press

  21. [29]

    Taewan Kim, Donghoon Shin, Young-Ho Kim, and Hwajung Hong. 2024. Diary- Mate: Understanding User Perceptions and Experience in Human-AI Collabo- ration for Personal Journaling. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15

  22. [30]

    Stephen J Kline and Nathan Rosenberg. 2010. An overview of innovation.Studies on science and the innovation process: Selected works of Nathan Rosenberg (2010), 173–203

  23. [31]

    Rafal Kocielnik, Saleema Amershi, and Paul N Bennett. 2019. Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–14

  24. [32]

    Sean Kross and Philip Guo. 2021. Orienting, framing, bridging, magic, and counseling: How data scientists navigate the outer loop of client collaborations in industry and academia.Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–28

  25. [33]

    Michelle S Lam, Zixian Ma, Anne Li, Izequiel Freitas, Dakuo Wang, James A Landay, and Michael S Bernstein. 2023. Model sketching: centering concepts in early-stage machine learning model design. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–24

  26. [34]

    Brenna Li, Ofek Gross, Noah Crampton, Mamta Kapoor, Saba Tauseef, Mohit Jain, Khai N Truong, and Alex Mariakakis. 2024. Beyond the Waiting Room: Patient’s Perspectives on the Conversational Nuances of Pre-Consultation Chatbots. In Proceedings of the CHI Conference on Human Fac...

  27. [35]

    Q Vera Liao, Hariharan Subramonyam, Jennifer Wang, and Jennifer Wort- man Vaughan. 2023. Designerly understanding: Information needs for model Park, et al. transparency to support design ideation for AI-powered user experience. In Proceedings of the 2023 CHI conference on huma...

  28. [36]

    Houjiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou, Daisy Pinaroc, Matthew Lease, and Min Kyung Lee. 2024. Human-centered NLP Fact-checking: Co- Designing with Fact-checkers using Matchmaking for AI. Proceedings of the ACM on Human-Computer Interaction 8, CSCW2 (2024), 1–44

  29. [37]

    Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2024. Selenite: Scaffolding Online Sensemak- ing with Comprehensive Overviews Elicited from Large Language Models. In Proceedings of the CHI Conference on Human Factors in ...

  30. [38]

    Yimeng Liu and Misha Sra. 2024. DanceGen: Supporting Choreography Ideation and Prototyping with Generative AI. In Proceedings of the 2024 ACM Designing Interactive Systems Conference. 920–938

  31. [39]

    Tobias Mann. 2023. Microsoft reportedly runs GitHub’s AI Copilot at a loss . Re- trieved 2023 from https://www.theregister.com/2023/10/11/github_ai_copilot_ microsoft/

  32. [40]

    Yaoli Mao, Dakuo Wang, Michael Muller, Kush R Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović. 2019. How data scientistswork together with domain experts in scientific collaborations: To find the right answer or to ask the right question? Proceedings of the ACM...

  33. [41]

    Steven Moore, Q Vera Liao, and Hariharan Subramonyam. 2023. fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision Tasks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19

  34. [42]

    Donald A Norman. 1986. User-centered System Design: New Perspectives on Human–computer Interaction

  35. [43]

    William Odom, Erik Stolterman, and Amy Yo Sue Chen. 2022. Extending a theory of slow technology for design through artifact analysis.Human–Computer Interaction 37, 2 (2022), 150–179

  36. [44]

    OpenAI. 2025. OpenAI. Retrieved 2025 from https://openai.com/chatgpt/

  37. [45]

    Google PAIR. 2019. People + AI Guidebook. Retrieved 2019 from pair.withgoogle. com/guidebook

  38. [46]

    Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, and Danielle Bragg. 2025. Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review. arXiv preprint arXiv:2501.12557 (2025)

  39. [47]

    Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020. Pre-trained models for natural language processing: A survey. Science China technological sciences 63, 10 (2020), 1872–1897

  40. [48]

    Woosuk Seo, Chanmo Yang, and Young-Ho Kim. 2024. ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20

  41. [49]

    Craig S. Smith. 2023. What Large Models Cost You – There Is No Free AI Lunch . Retrieved Sep 08, 2023 from https://www.forbes.com/sites/craigsmith/2023/09/ 08/what-large-models-cost-you--there-is-no-free-ai-lunch/

  42. [50]

    UI-Patterns. 2007. UI-Patterns.com. Retrieved 2007 from https://ui-patterns.com

  43. [51]

    Usabilityfirst. 2015. Artifact Analysis. Retrieved Jan 30, 2015 from https://www. usabilityfirst.com/glossary/artifact-analysis/

  44. [52]

    UXPin. 2023. Examples of Interaction Design — Patterns and Best Practices . Re- trieved May 30, 2023 from https://www.uxpin.com/studio/blog/examples-of- interaction-design/

  45. [53]

    Priyan Vaithilingam, Elena L Glassman, Jeevana Priya Inala, and Chenglong Wang. 2024. DynaVis: Dynamically Synthesized UI Widgets for Visualization Editing. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  46. [54]

    Brian Wang. 2024. IF AI LLM Queries Replace Google Internet Search . Retrieved April 9, 2024 from https://www.nextbigfuture.com/2024/04/if-ai-llm-queries- replace-google-internet-search.html

  47. [55]

    Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, and Tianyi Zhang. 2024. PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21

  48. [56]

    Joyce Weiner. 2022. Why AI/data science projects fail: how to avoid project pitfalls . Springer Nature

  49. [57]

    Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang. 2016. A survey of transfer learning. Journal of Big data 3 (2016), 1–40

  50. [58]

    Derek White. 2024. Future-Proofing Banking: The Transition From Digital To Intelligent . Retrieved 2024 from https://www.forbes.com/councils/ forbestechcouncil/2024/11/01/future-proofing-banking-the-transition-from- digital-to-intelligent/

  51. [59]

    Anna Xygkou, Chee Siang Ang, Panote Siriaraya, Jonasz Piotr Kopecki, Alexandra Covaci, Eiman Kanjo, and Wan-Jou She. 2024. MindTalker: Navigating the Complexities of AI-Enhanced Social Engagement for People with Early-Stage Dementia. In Proceedings of the CHI Conference on Hum...

  52. [60]

    Qian Yang, Nikola Banovic, and John Zimmerman. 2018. Mapping machine learn- ing advances from hci research to reveal starting places for design innovation. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–11

  53. [61]

    Qian Yang, Justin Cranshaw, Saleema Amershi, Shamsi T Iqbal, and Jaime Teevan

  54. [62]

    Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. 2020. Re- examining whether, why, and how human-AI interaction is uniquely difficult to design. In Proceedings of the 2020 chi conference on human factors in computing systems. 1–13

  55. [63]

    Qian Yang, John Zimmerman, Aaron Steinfeld, and Anthony Tomasic. 2016. Planning adaptive mobile experiences when wireframing. In Proceedings of the 2016 ACM Conference on Designing Interactive Systems . 565–576

  56. [64]

    Nur Yildirim, Alex Kass, Teresa Tung, Connor Upton, Donnacha Costello, Robert Giusti, Sinem Lacin, Sara Lovic, James M O’Neill, Rudi O’Reilly Meehan, et al

  57. [65]

    Nur Yildirim, Changhoon Oh, Deniz Sayar, Kayla Brand, Supritha Challa, Violet Turri, Nina Crosby Walton, Anna Elise Wong, Jodi Forlizzi, James McCann, et al. 2023. Creating design resources to scaffold the ideation of AI concepts. In Proceedings of the 2023 ACM Designing Inter...

  58. [66]

    Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fer- nanda Viégas. 2023. Investigating how practitioners use human-ai guidelines: A case study on the people+ ai guidebook. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–13

  59. [67]

    Nur Yildirim, Susanna Zlotnikov, Deniz Sayar, Jeremy M Kahn, Leigh A Bukowski, Sher Shah Amin, Kathryn A Riman, Billie S Davis, John S Minturn, Andrew J King, et al. 2024. Sketching AI Concepts with Capabilities and Examples: AI Innovation in the Intensive Care Unit. In Procee...

  60. [68]

    JD Zamfirescu-Pereira, Heather Wei, Amy Xiao, Kitty Gu, Grace Jung, Matthew G Lee, Bjoern Hartmann, and Qian Yang. 2023. Herding AI cats: Lessons from de- signing a chatbot by prompting GPT-3. In Proceedings of the 2023 ACM Designing Interactive Systems Conference. 2206–2220

  61. [69]

    What It Wants Me To Say

    Chen Zhou, Zihan Yan, Ashwin Ram, Yue Gu, Yan Xiang, Can Liu, Yun Huang, Wei Tsang Ooi, and Shengdong Zhao. 2024. GlassMail: Towards Personalised Wearable Assistant for On-the-Go Email Creation on Smart Glasses. In Proceed- ings of the 2024 ACM Designing Interactive Systems Co...

  62. [2019]

    InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems

    Sketching nlp: A case study of exploring the right things to design with language intelligence. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–12

  63. [2022]

    In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems

    How experienced designers of enterprise applications engage AI as a design material. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–13

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.