Pith. sign in

REVIEW 3 major objections 1 minor 1 cited by

Data-driven Progressive Discovery of Physical Laws

T0 review · 3 major / 1 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Physical laws are recovered from data by chaining simple, meaningful symbolic units rather than inventing one long expression at once.

desk verdict We only have the CoSR abstract; the supplied full text is an unrelated CHI UX paper, so every progressive-discovery claim is uncheckable. read the letter →

arxiv 2603.13727 v2 pith:4BHBVVHU submitted 2026-03-14 cs.LG physics.data-an

classification cs.LGphysics.data-an
keywords symbolicregressionphysicallawdiscoveryprogressiveknowledgechainscalinglawsinterpretablemachinelearningscientificaerodynamiccoefficients
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Conventional symbolic regression tries to invent a complete mathematical law in a single step and often produces long, unphysical formulas that fail to generalize. The paper argues that real physical discovery proceeds hierarchically, from simple relations to more complex ones, and therefore models discovery itself as a progressive chain of symbolic knowledge units that each carry clear physical meaning. By combining these units step by step along a logical path, the method recovers known laws (Kepler to Newton) and improves classical scaling relations in fluid and laser problems, while also extracting new scaling knowledge for aircraft aerodynamics. A sympathetic reader cares because the approach restores interpretability and generalization precisely where pure end-to-end search collapses into meaningless expressions.

What carries the argument

Chain of Symbolic Regression (CoSR): a progressive sequence of symbolic knowledge units that are combined along a fixed logical path, each unit remaining physically interpretable.

What would settle it

Apply CoSR and a standard one-step symbolic regressor to a system whose progressive discovery path is already known (for example the Kepler-to-Newton sequence) and check whether CoSR alone recovers the intermediate units and the final law while the one-step method produces only lengthy unphysical expressions.

Watch

Extended reading notes

Core claim

The authors claim that physical laws do not appear as monolithic expressions but as hierarchical chains; their Chain of Symbolic Regression therefore discovers them by successively assembling discrete knowledge units that already possess clear physical meanings, ultimately recovering the correct underlying law from data.

Load-bearing premise

The claim rests on the premise that every physical law of interest can be broken into a fixed hierarchy of simple, physically meaningful symbolic units whose progressive combination is both necessary and sufficient to recover the true law.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The abstract proposes Chain of Symbolic Regression (CoSR), a framework that discovers physical laws by progressively combining discrete symbolic knowledge units with clear physical meanings, rather than one-shot end-to-end symbolic regression. It claims to recapitulate the historical path from Kepler’s third law to Newton’s law of universal gravitation, to improve classical scaling relations for turbulent Rayleigh–Bénard convection, pipe flow, and laser–metal interaction, and to discover new aerodynamic-coefficient scalings for aircraft. The supplied full manuscript body, however, is an entirely unrelated CHI ’26 multi-session HCI study of UX evaluators collaborating with novice versus experienced conversational AI assistants (arXiv:2603.13717). Consequently no methods, algorithms, intermediate expressions, datasets, error metrics, baselines, or ablations for CoSR are present.

Significance. If the progressive-chain idea were correctly implemented and validated, CoSR would address a recognized failure mode of conventional symbolic regression (lengthy, unphysical expressions with poor generalization) and would constitute a meaningful methodological contribution to data-driven scientific discovery. The historical Kepler-to-Newton reconstruction and the claimed improvements to classical scaling laws would be especially valuable if independently recovered from data. None of these contributions can be assessed from the manuscript as supplied.

major comments (3)
  1. The full text provided under the CoSR title and abstract is the complete, unrelated CHI ’26 paper “It Became My Buddy, But I’m Not Afraid to Disagree” (arXiv:2603.13717). No section, equation, algorithm, figure, or table belonging to CoSR appears. Every empirical claim in the abstract—recovery of the gravitational law, improved Rayleigh–Bénard / pipe-flow / laser-metal scalings, and new aircraft aerodynamic coefficients—is therefore uninspectable.
  2. Because the method body is absent, it is impossible to determine whether the “knowledge units” and the “specific logic” that combine them are recovered from data or are seeded by the same classical knowledge the method claims to rediscover. This circularity risk, already flagged by the abstract’s motivating premise, cannot be evaluated.
  3. No datasets, intermediate symbolic expressions, quantitative error metrics, baselines (e.g., standard genetic-programming or sparse-regression SR), or ablation studies are supplied. The central claim that progressive chaining overcomes the unphysical expressions of one-step SR therefore rests solely on an unsupported abstract.
minor comments (1)
  1. The arXiv identifier printed in the body (2603.13717) does not match the identifier under review (2603.13727), confirming a manuscript-swap error.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be assessed: supplied full text is an unrelated CHI UX/CA study; CoSR derivation chain is absent.

full rationale

The claimed paper (arXiv 2603.13727, CoSR / progressive symbolic discovery of physical laws) is represented only by its abstract. The body that follows under FULL MANUSCRIPT TEXT is an entirely different manuscript (CHI ’26 multi-session UX-evaluator study with conversational AI assistants, arXiv 2603.13717). Consequently there are no CoSR equations, intermediate symbolic units, fitting procedures, Kepler-to-Newton reconstruction steps, scaling-law improvements, or aircraft-coefficient results that can be inspected. Per the analyzer rules, circularity may be flagged only when a specific reduction can be quoted and exhibited (self-definitional identity, fitted parameter renamed as prediction, load-bearing self-citation, etc.). With no derivation present, no such reduction exists. The abstract’s narrative claim that CoSR “fully recapitulates” a known historical path is therefore uncheckable rather than circular. Score is 0; steps list is empty.

Assumptions & free parameters 0 free parameters · 2 assumptions · 1 invented entities

With only the abstract, the ledger is necessarily incomplete. The central claim rests on the domain assumption that physical discovery is hierarchical and progressive, plus the ad-hoc construction of “knowledge units” and a “chain” logic whose precise definition is not given. No free parameters or invented physical entities (particles, forces, etc.) are named in the abstract.

assumptions (2)
  • domain assumption Physical laws do not exist in a single form but follow a hierarchical and progressive pattern from simplicity to complexity.
    Stated as the fundamental motivation; if false for the target systems, the chain construction loses its justification.
  • ad hoc to paper A chain of symbolic knowledge units with clear physical meanings can be progressively combined along a specific logic to recover the true underlying law.
    Core modeling choice of CoSR; the abstract does not derive why this particular chaining is complete or unique.
invented entities (1)
  • Chain of Symbolic Regression (CoSR) / knowledge chain / knowledge units
    purpose: To structure symbolic discovery as progressive combination of simple units rather than one-shot search.
    These are the paper’s central constructs; no independent evidence outside the claimed recoveries is supplied in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data-driven Progressive Discovery of Physical Laws." pith.science (2026). https://pith.science/paper/4BHBVVHU

@misc{pith2026260313727,
  author       = {Pith},
  title        = {Pith review of: Data-driven Progressive Discovery of Physical Laws},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4BHBVVHU}},
  note         = {Machine review of arXiv:2603.13727}
}
read the original abstract

Symbolic regression is a powerful tool for knowledge discovery, enabling the extraction of interpretable mathematical expressions directly from data. However, conventional symbolic discovery typically follows an end-to-end, "one-step" process, which often generates lengthy and physically meaningless expressions when dealing with real physical systems, leading to poor model generalization. This limitation fundamentally stems from its deviation from the basic path of scientific discovery: physical laws do not exist in a single form but follow a hierarchical and progressive pattern from simplicity to complexity. Motivated by this principle, we propose Chain of Symbolic Regression (CoSR), a novel framework that models the discovery of physical laws as a chain of symbolic knowledge. This knowledge chain is formed by progressively combining multiple knowledge units with clear physical meanings along a specific logic, ultimately enabling the precise discovery of the underlying physical laws from data. CoSR fully recapitulates the progressive discovery path from Kepler's third law to the law of universal gravitation in classical mechanics, and is applied to three types of problems: turbulent Rayleigh-Benard convection, viscous flows in a circular pipe, and laser-metal interaction, demonstrating its ability to improve classical scaling theories. Finally, CoSR showcases its capability to discover new knowledge in the complex engineering problem of aerodynamic coefficients scaling for different aircraft.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Skin friction prediction for attached flows based on two-dimensional inviscid solutions

    physics.flu-dyn 2026-07 conditional novelty 6.0 of 10

    A progressive symbolic regression framework discovers an interpretable analytical formula chain for skin friction prediction across subsonic to hypersonic attached flows using inviscid surface flow features.

Reference graph

Works this paper leans on

162 extracted references · 41 canonical work pages · cited by 1 Pith paper

  1. [1]

    Ajenaghughrure, Sonia C

    Ighoyota Ben. Ajenaghughrure, Sonia C. Sousa, Ilkka Johannes Kosunen, and David Lamas. 2019. Predictive model to assess user trust: a psycho-physiological approach. InProceedings of the 10th Indian Conference on Human-Computer Interaction (IndiaHCI ’19). Association for Computing Machinery, New York, NY, USA, 1–10. https://doi.org/10.1145/3364183.3364195

  2. [2]

    AltexSoft. 2019. UX Researcher: Methods, Skills and Process. https://www. altexsoft.com/blog/ux-researcher-methods-skills-process/

  3. [3]

    Theo Araujo and Nadine Bol. 2024. From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents.Computers in Human Behavior: Artificial Humans2, 1 (Jan. 2024), 100030. https://doi.org/10.1016/j.chbah.2023.100030

  4. [4]

    Andrea Batch, Yipeng Ji, Mingming Fan, Jian Zhao, and Niklas Elmqvist. 2023. uxSense: Supporting User Experience Analysis with Visualization and Com- puter Vision.IEEE Transactions on Visualization and Computer Graphics(2023), 1–15. https://doi.org/10.1109/TVCG.2023.3241581 Conference Name: IEEE Transactions on Visualization and Computer Graphics

  5. [5]

    Patricia Benner. 1982. From Novice To Expert.AJN The American Journal of Nursing82, 3 (1982), 402–407. https://journals.lww.com/ajnonline/citation/ 1982/82030/from_novice_to_expert.4.a

  6. [6]

    Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond. 2023. Generative AI at Work(Working Paper Series). National Bureau of Economic Research. https://doi.org/10.3386/w31161

  7. [7]

    Félix Buendía-García and Javier Piris-Ruano. 2025. Using Generative AI to Support UX Design Students in Web Development Courses.Applied Sciences15, 13 (June 2025), 7389. https://doi.org/10.3390/app15137389 Publisher: Multidis- ciplinary Digital Publishing Institute

  8. [8]

    Gajos, and Elena L

    Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. InProceedings of the 25th International Conference on Intelligent User Interfaces (IUI ’20). Association for Computing Machinery, New York, NY, USA, 454–464. https://doi.org/10.1145/3377325.3377498

Show all 162 references
  1. [9]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.Proc. ACM Hum.-Comput. Interact.5, CSCW1 (April 2021), 188:1–188:21. https://doi.org/10.1145/3449287

  2. [10]

    2006.Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis

    Kathy Charmaz. 2006.Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis. SAGE. Google-Books-ID: v1qP1KbXz1AC. Multi-Session Study of UX Evaluators Using Conversational AI Agents CHI ’26, April 13–17, 2026, Barcelona, Spain

  3. [11]

    Chase, Doris B

    Catherine C. Chase, Doris B. Chin, Marily A. Oppezzo, and Daniel L. Schwartz

  4. [12]

    2009), 334–352

    Teachable Agents and the Protégé Effect: Increasing the Effort Towards Learning.Journal of Science Education and Technology18, 4 (Aug. 2009), 334–352. https://doi.org/10.1007/s10956-009-9180-4

  5. [13]

    Hsi-Jen Chen, Yan-Ting Chen, and Chia-Han Yang. 2022. Behaviors of Novice and Expert Designers in the Design Process: From Discovery to Design.Inter- national Journal of Design16, 3 (2022), 59–76. https://doi.org/10.57698/v16i3.04

  6. [14]

    Chilana, Jacob O

    Parmit K. Chilana, Jacob O. Wobbrock, and Andrew J. Ko. 2010. Understanding Usability Practices in Complex Domains. InProceedings of the 28th International Conference on Human Factors in Computing Systems - CHI ’10. ACM Press, Atlanta, Georgia, USA, 2337–2346. https://doi.org/...

  7. [15]

    Dorothee Clasen and Marc Hassenzahl. 2024. Fostering people’s autonomy by foregrounding and questioning daily choices. InAdjunct Proceedings of the 2024 Nordic Conference on Human-Computer Interaction (NordiCHI ’24 Adjunct). Association for Computing Machinery, New York, NY, U...

  8. [16]

    Andy Cockburn, Carl Gutwin, Joey Scarr, and Sylvain Malacria. 2014. Supporting Novice to Expert Transitions in User Interfaces.ACM Comput. Surv.47, 2 (Nov. 2014), 31:1–31:36. https://doi.org/10.1145/2659796

  9. [17]

    Tyler Corwin, Mehmet Kosa, Mahsa Nasri, Christoffer Holmgård, and Casper Harteveld. 2023. The Teaching Efficacy of the Protégé Effect in Gamified Edu- cation. In2023 IEEE Conference on Games (CoG). 1–8. https://doi.org/10.1109/ CoG57401.2023.10333166 ISSN: 2325-4289

  10. [18]

    Emmelyn A. J. Croes and Marjolijn L. Antheunis. 2021. Can we be friends with Mitsuku? A longitudinal study on the process of relationship formation between humans and a social chatbot - Emmelyn A. J. Croes, Marjolijn L. Antheunis, 2021.Journal of Social and Personal Relationsh...

  11. [19]

    Fleur Deken, Maaike Kleinsmann, Marco Aurisicchio, Kristina Lauche, and Rob Bracewell. 2011. Tapping into past design experiences: knowledge sharing and creation during novice–expert design consultations.Research in Engineering Design23, 3 (Oct. 2011), 203–218. https://doi.org...

  12. [20]

    2022.Intelligent Adaptive Flight Training System- A Human Performance in the Loop for Real-Time Decision Making

    Jean-François Delisle. 2022.Intelligent Adaptive Flight Training System- A Human Performance in the Loop for Real-Time Decision Making. phd. Polytechnique Montréal. https://publications.polymtl.ca/10551/

  13. [21]

    Dietvorst, Joseph P

    Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after seeing them err.Journal of Experimental Psychology: General144, 1 (2015), 114–126. https://doi.org/10. 1037/xge0000033 Place: US Publisher: American P...

  14. [22]

    Shiying Ding, Xinyi Chen, Yan Fang, Wenrui Liu, Yiwu Qiu, and Chunlei Chai

  15. [23]

    In2023 16th Interna- tional Symposium on Computational Intelligence and Design (ISCID)

    DesignGPT: Multi-Agent Collaboration in Design. In2023 16th Interna- tional Symposium on Computational Intelligence and Design (ISCID). 204–208. https://doi.org/10.1109/ISCID59865.2023.00056 ISSN: 2473-3547

  16. [24]

    Peitong Duan, Jeremy Warner, Yang Li, and Bjoern Hartmann. 2024. Generating Automatic Feedback on UI Mockups with Large Language Models. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY,...

  17. [25]

    Mingming Fan, Yue Li, and Khai N. Truong. 2020. Automatic Detection of Usability Problem Encounters in Think-Aloud Sessions.ACM Transactions on Interactive Intelligent Systems10, 2 (June 2020), 1–24. https://doi.org/10.1145/ 3385732

  18. [26]

    Mingming Fan, Serina Shi, and Khai N Truong. 2020. Practices and Challenges of Using Think-Aloud Protocols in Industry: An International Survey.Journal of Usability Studies15, 2 (2020), 85–102

  19. [27]

    Mingming Fan, Ke Wu, Jian Zhao, Yue Li, Winter Wei, and Khai N. Truong

  20. [28]

    2020), 343–352

    VisTA: Integrating Machine Intelligence with Visualization to Support the Investigation of Think-Aloud Sessions.IEEE Transactions on Visualization and Computer Graphics26, 1 (Jan. 2020), 343–352. https://doi.org/10.1109/TVCG. 2019.2934797

  21. [29]

    Vera Liao, and Jian Zhao

    Mingming Fan, Xianyou Yang, TszTung Yu, Q. Vera Liao, and Jian Zhao. 2022. Human-AI Collaboration for UX Evaluation: Effects of Explanation and Synchro- nization. 6 (2022), 96:1–96:32. Issue CSCW1. https://doi.org/10.1145/3512943

  22. [30]

    James Fisher. 1991. Defining the novice user.Behaviour & Information Technology 10, 5 (1991), 437–441. https://doi.org/10.1080/01449299108924301 Place: United Kingdom Publisher: Taylor & Francis

  23. [31]

    Asbjørn Følstad, Effie Lai-Chong Law, and Kasper Hornbæk. 2010. Analysis in Usability Evaluations: An Exploratory Study. InProceedings of the 6th Nordic Conference on Human-Computer Interaction: Extending Boundaries (NordiCHI ’10). Association for Computing Machinery, New York...

  24. [32]

    Asbjørn Følstad, Effie Lai-Chong Law, and Kasper Hornbæk. 2012. Analysis in Practical Usability Evaluation: A Survey Study. InProceedings of the 30th SIGCHI Conference on Human Factors in Computing Systems - CHI ’12. ACM Press, Austin, Texas, 2127–2136. https://doi.org/10.1145...

  25. [33]

    Eureka Foong, Darren Gergle, and Elizabeth M. Gerber. 2017. Novice and Expert Sensemaking of Crowdsourced Design Feedback.Proc. ACM Hum.-Comput. Interact.1, CSCW (Dec. 2017), 45:1–45:18. https://doi.org/10.1145/3134680

  26. [34]

    2010.Longitudinal Research in Human-Computer Interaction

    Jens Gerken. 2010.Longitudinal Research in Human-Computer Interaction. Ph. D. Dissertation. Universität Konstanz, Konstanz, Germany. https://kops.uni- konstanz.de/entities/publication/ef10c26c-3f4c-41a9-8045-da53586fc360

  27. [35]

    Kaklamanos

    Kostis Giannakopoulos, Argyro Kavadella, Anas Aaqel Salim, Vassilis Stam- atopoulos, and Eleftherios G. Kaklamanos. 2023. Evaluation of the Perfor- mance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry...

  28. [36]

    Julián Grigera, Alejandra Garrido, José Matías Rivero, and Gustavo Rossi. 2017. Automatic detection of usability smells in web applications.International Journal of Human-Computer Studies97 (2017), 129–148

  29. [37]

    Siddharth Gulati, Joe McDonagh, Sonia Sousa, and David Lamas. 2024. Trust models and theories in human–computer interaction: A systematic literature review.Computers in Human Behavior Reports16 (Dec. 2024), 100495. https: //doi.org/10.1016/j.chbr.2024.100495

  30. [38]

    An Error Occurred!

    Kasper Hald, Katharina Weitz, Elisabeth André, and Matthias Rehm. 2021. “An Error Occurred!” - Trust Repair With Virtual Robot Using Levels of Mistake Explanation. InProceedings of the 9th International Conference on Human-Agent Interaction (HAI ’21). Association for Computing...

  31. [39]

    Patrick Harms. 2019. Automated usability evaluation of virtual reality applica- tions.ACM Transactions on Computer-Human Interaction (TOCHI)26, 3 (2019), 1–36

  32. [40]

    Morten Hertzum and Niels Ebbe Jacobsen. 2001. The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods.International Journal of Human-Computer Interaction15, 1 (2001), 183–204. https://doi.org/10.1207/ S15327590IJHC1501_14

  33. [41]

    Morten Hertzum, Niels Ebbe Jacobsen, and Bonnie E. John. 1998. The Evaluator Effect in Usability Tests. InCHI 98 Conference Summary on Human Factors in Computing Systems (CHI ’98). Association for Computing Machinery, New York, NY, USA, 255–256. https://doi.org/10.1145/286498.286737

  34. [42]

    Annabell Ho, Jeff Hancock, and Adam S Miner. 2018. Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot. Journal of Communication68, 4 (Aug. 2018), 712–733. https://doi.org/10.1093/ joc/jqy026

  35. [43]

    Hoffman, Matthew Johnson, Jeffrey M

    Robert R. Hoffman, Matthew Johnson, Jeffrey M. Bradshaw, and Al Underbrink

  36. [44]

    2013), 84–88

    Trust in Automation.IEEE Intelligent Systems28, 1 (Jan. 2013), 84–88. https://doi.org/10.1109/MIS.2013.24

  37. [45]

    Jane Hogan, Gary Grant, Fiona Kelly, and Jennie O’Hare. 2020. Factors influ- encing acceptance of robotics in hospital pharmacy: a longitudinal study using the Extended Technology Acceptance Model.International Journal of Pharmacy Practice28, 5 (Oct. 2020), 483–490. https://do...

  38. [46]

    Ironhack. 2023. The Evolving Role of a UX/UI Designer: Trends and Skills for Success. https://www.ironhack.com/gb/blog/the-evolving-role-of-a-ux-ui- designer-trends-and-skills-for-success

  39. [47]

    JongWook Jeong, NeungHoe Kim, and Hoh Peter In. 2020. Detecting usability problems in mobile applications on the basis of dissimilarity in user behavior. International Journal of Human-Computer Studies139 (2020), 102364

  40. [48]

    Kahr, Gerrit Rooks, Martijn C

    Patricia K. Kahr, Gerrit Rooks, Martijn C. Willemsen, and Chris C. P. Sni- jders. 2024. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Pre- vious Interactions.ACM Trans. Interact. Intell. Syst.(Aug....

  41. [49]

    It Requires Interest, Time, Patience and Struggle

    Mahmut Kalman. 2019. “It Requires Interest, Time, Patience and Struggle”: Novice Researchers’ Perspectives on and Experiences of the Qualitative Research Journey.Qualitative Research in Education8, 3 (Oct. 2019), 341–377. https: //doi.org/10.17583/qre.2019.4483 Number: 3

  42. [50]

    Takayuki Kanda, Rumi Sato, Naoki Saiwaki, and Hiroshi Ishiguro. 2007. A Two-Month Field Trial in an Elementary School for Long-Term Human–Robot Interaction.IEEE Transactions on Robotics23, 5 (Oct. 2007), 962–971. https: //doi.org/10.1109/TRO.2007.904904 Conference Name: IEEE T...

  43. [51]

    Evangelos Karapanos, Jhilmil Jain, and Marc Hassenzahl. 2012. Theories, meth- ods and case studies of longitudinal HCI research. InCHI ’12 Extended Abstracts on Human Factors in Computing Systems (CHI EA ’12). Association for Com- puting Machinery, New York, NY, USA, 2727–2730...

  44. [52]

    Abdolvahab Khademi. 2023. Can ChatGPT and Bard Generate Aligned Assess- ment Items? A Reliability Analysis against Human Performance.Journal of Applied Learning & Teaching6, 1 (May 2023). https://doi.org/10.37074/jalt.2023. 6.1.28 arXiv:2304.05372 [cs]. CHI ’26, April 13–17, 2...

  45. [53]

    Ahmad Khawaji, Jianlong Zhou, Fang Chen, and Nadine Marcus. 2015. Using Galvanic Skin Response (GSR) to Measure Trust and Cognitive Load in the Text- Chat Environment. InProceedings of the 33rd Annual ACM Conference Extended Abstracts on Human Factors in Computing Systems (CHI...

  46. [54]

    Jay Kim. 2020. Rubber Ducking- What It Is and Why It Works. https://medium.com/@jkarma0920/rubber-ducking-what-it-is-and-why-it- works-5d026fd9ae58

  47. [55]

    Sunyoung Kim and Abhishek Choudhury. 2021. Exploring older adults’ per- ception and use of smart speaker-based voice assistants: A longitudinal study. Computers in Human Behavior124 (Nov. 2021), 106914. https://doi.org/10.1016/ j.chb.2021.106914

  48. [56]

    Skov, Peter Axel Nielsen, Jesper Kjeldskov, Jens Gerken, and Harald Reiterer

    Maria Kjærup, Mikael B. Skov, Peter Axel Nielsen, Jesper Kjeldskov, Jens Gerken, and Harald Reiterer. 2021. Longitudinal Studies in HCI Research: A Review of CHI Publications From 1982–2019. InAdvances in Longitudinal HCI Research, Evangelos Karapanos, Jens Gerken, Jesper Kjel...

  49. [57]

    Eva Krapp, Robin Neuhaus, Marc Hassenzahl, and Matthias Laschke. 2024. In a Quasi-Social Relationship With ChatGPT. An Autoethnography on Engaging With Prompt-Engineered LLM Personas. InProceedings of the 13th Nordic Confer- ence on Human-Computer Interaction (NordiCHI ’24). A...

  50. [58]

    Emily Kuang. 2025. Evaluating Usability Challenges in VR Games for Older Adults: A Comparison With and Without AI Assistance. InAdjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST Adjunct ’25). Association for Computing Machiner...

  51. [59]

    Emily Kuang, Ehsan Jahangirzadeh Soure, Mingming Fan, Jian Zhao, and Kristen Shinohara. 2023. Collaboration with Conversational AI Assistants for UX Evaluation: Questions and How to Ask them (Voice vs. Text). InPro- ceedings of the 2023 CHI Conference on Human Factors in Compu...

  52. [60]

    Merging Results Is No Easy Task

    Emily Kuang, Xiaofu Jin, and Mingming Fan. 2022. "Merging Results Is No Easy Task": An International Survey Study of Collaborative Data Analysis Practices Among UX Practitioners. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems.Association for Com...

  53. [61]

    Emily Kuang, Minghao Li, Mingming Fan, and Kristen Shinohara. 2024. Enhanc- ing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing. InProceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24)....

  54. [62]

    Vera Liao, Alison Smith-Renner, and Chenhao Tan

    Vivian Lai, Chacha Chen, Q. Vera Liao, Alison Smith-Renner, and Chenhao Tan

  55. [63]

    https://doi.org/10.48550/arXiv.2112.11471 arXiv:2112.11471 [cs]

    Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies. https://doi.org/10.48550/arXiv.2112.11471 arXiv:2112.11471 [cs]

  56. [64]

    Richard Landis and Gary G

    J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data.Biometrics33, 1 (March 1977), 159. https: //doi.org/10.2307/2529310

  57. [65]

    Claire Lauer, Danielle Storey, and Romit Soley. 2024. Vector Personas: How UX Researchers Can Use AI to Bring New Dimension to Traditional Persona Development. InProceedings of the 42nd ACM International Conference on Design of Communication (SIGDOC ’24). Association for Compu...

  58. [66]

    Nadine Lessio and Alexis Morris. 2020. Toward Design Archetypes for Conver- sational Agent Personality. In2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC). 3221–3228. https://doi.org/10.1109/SMC42975. 2020.9283254 ISSN: 2577-1655

  59. [67]

    Tobias Lieberei, Virginia Deborah Elaine Welter, Leroy Großmann, and Moritz Krell. 2023. Findings from the expert-novice paradigm on differential response behavior among multiple-choice items of a pedagogical content knowledge test – implications for test development. 14 (2023...

  60. [68]

    Yiren Liu, Pranav Sharma, Mehul Oswal, Haijun Xia, and Yun Huang. 2025. PersonaFlow: Designing LLM-Simulated Expert Perspectives for Enhanced Research Ideation. InProceedings of the 2025 ACM Designing Interactive Systems Conference (DIS ’25). Association for Computing Machiner...

  61. [69]

    Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2023. Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing. https://doi.org/10.48550/arXiv. 2305.09434 arXiv:2305.09434 [cs]

  62. [70]

    Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2024. Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Deci- sions. InProceedings of the IEEE/ACM 46th International Confe...

  63. [71]

    Tao Long, Katy Ilonka Gero, and Lydia B Chilton. 2024. Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI Workflow. In Proceedings of the 2024 ACM Designing Interactive Systems Conference (DIS ’24). Association for Computing Machinery, New York, NY, U...

  64. [72]

    Yuwen Lu, Yuewen Yang, Qinyi Zhao, Chengzhi Zhang, and Toby Jia-Jun Li

  65. [73]

    https://doi.org/10.48550/arXiv.2402.06089 arXiv:2402.06089 [cs]

    AI Assistance for UX: A Literature Review Through Human-Centered AI. https://doi.org/10.48550/arXiv.2402.06089 arXiv:2402.06089 [cs]

  66. [74]

    Like Having a Really Bad PA

    Ewa Luger and Abigail Sellen. 2016. "Like Having a Really Bad PA": The Gulf between User Expectation and Experience of Conversational Agents. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (New York, NY, USA)(CHI ’16). Association for Computing...

  67. [75]

    https://doi.org/10.1145/2858036.2858288

  68. [76]

    Kim Man Lui and Keith C. C. Chan. 2006. Pair programming productivity: Novice–novice vs. expert–expert.International Journal of Human-Computer Studies64, 9 (Sept. 2006), 915–925. https://doi.org/10.1016/j.ijhcs.2006.04.010

  69. [77]

    Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2025. Towards Human-AI Deliberation: Design and Evalua- tion of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making. In Proceedings of the 2025 CHI Conference on Human Factors ...

  70. [78]

    Ellie Martin. 2016. Why 5 is the magic number for UX usability testing | Inside Design Blog. https://www.invisionapp.com/inside-design/ux-usability- research-testing/

  71. [79]

    McGrath, Oliver Lack, James Tisch, and Andreas Duenser

    Melanie J. McGrath, Oliver Lack, James Tisch, and Andreas Duenser. 2025. Measuring trust in artificial intelligence: validation of an established scale and its short form.Frontiers in Artificial Intelligence8 (May 2025). https: //doi.org/10.3389/frai.2025.1582880 Publisher: Frontiers

  72. [80]

    Valerie Mendoza and David G. Novick. 2005. Usability over time. InPro- ceedings of the 23rd annual international conference on Design of communi- cation: documenting & designing for pervasive information (SIGDOC ’05). As- sociation for Computing Machinery, New York, NY, USA, 1...

  73. [81]

    Kate Moran. 2019. Usability Testing 101. https://www.nngroup.com/articles/ usability-testing-101/

  74. [82]

    Mutsuddi and Kay Connelly

    Adity U. Mutsuddi and Kay Connelly. 2012. Text messages for encouraging physi- cal activity Are they effective after the novelty effect wears off?. In2012 6th Inter- national Conference on Pervasive Computing Technologies for Healthcare (Perva- siveHealth) and Workshops. 33–40...

  75. [83]

    Nelson and Erik Stolterman

    Harold G. Nelson and Erik Stolterman. 2012. The Ultimate Particular. InThe Design Way: Intentional Change in an Unpredictable World. MIT Press, 27–40. https://ieeexplore.ieee.org/document/6354147 Conference Name: The Design Way: Intentional Change in an Unpredictable World

  76. [84]

    Jakob Nielsen. 1994. Heuristic Evaluation. InUsability Inspection Methods. John Wiley & Sons, Ltd, New York, NY

  77. [85]

    Jakob Nielsen. 2000. Why You Only Need to Test with 5 Users. https://www. nngroup.com/articles/why-you-only-need-to-test-with-5-users/

  78. [86]

    Jakob Nielsen. 2012. Usability 101: Introduction to Usability. https://www. nngroup.com/articles/usability-101-introduction-to-usability/

  79. [87]

    Jakob Nielsen. 2024. What AI Can and Cannot Do for UX. https://www.uxtigers. com/post/ai-can-cannot-do-ux

  80. [88]

    Mack (Eds.)

    Jakob Nielsen and Robert L. Mack (Eds.). 1994.Usability inspection methods. John Wiley & Sons, Inc

  81. [89]

    Mie Nørgaard and Kasper Hornbæk. 2006. What Do Usability Evaluators Do in Practice? An Explorative Study of Think-Aloud Testing. InProceedings of the 6th Conference on Designing Interactive Systems (DIS ’06). Association for Computing Machinery, New York, NY, USA, 209–218. htt...

  82. [90]

    Don Norman and Jakob Neilsen. 1998. The Definition of User Experience (UX). https://www.nngroup.com/articles/definition-user-experience/

  83. [91]

    OpenAI. 2023. Introducing GPTs. https://openai.com/blog/introducing-gpts

  84. [92]

    OpenAI. 2024. Creating a GPT | OpenAI Help Center. https://help.openai.com/ en/articles/8554397-creating-a-gpt

  85. [93]

    OpenAI. 2024. OpenAI Platform. https://platform.openai.com

  86. [94]

    OpenAI. 2024. Prompt engineering. https://platform.openai.com/docs/guides/ prompt-engineering

  87. [95]

    Asil Oztekin, Dursun Delen, Ali Turkyilmaz, and Selim Zaim. 2013. A Machine Learning-Based Usability Evaluation Method for eLearning Systems.Decision Support Systems56 (Dec. 2013), 63–73. https://doi.org/10.1016/j.dss.2013.05.003

  88. [96]

    Debajyoti Pal, Vajirasak Vanijja, Himanshu Thapliyal, and Xiangmin Zhang

  89. [97]

    2023), 107788

    What affects the usage of artificial conversational agents? An agent personality and love theory perspective.Computers in Human Behavior145 (Aug. 2023), 107788. https://doi.org/10.1016/j.chb.2023.107788 Multi-Session Study of UX Evaluators Using Conversational AI Agents CHI ’2...

  90. [98]

    Fabio Paternò, Antonio Giovanni Schiavone, and Antonio Conti. 2017. Cus- tomizable automatic detection of bad usability smells in mobile accessed web applications. InProceedings of the 19th International Conference on Human- Computer Interaction with Mobile Devices and Services. 1–11

  91. [99]

    Payscale. 2024. UX Researcher Salary in 2024 | PayScale. https://www.payscale. com/research/US/Job=UX_Researcher/Salary

  92. [100]

    Iryna Pentina, Tianling Xie, Tyler Hancock, and Ainsworth Bailey. 2023. Consumer–machine relationships in the age of artificial intelligence: Sys- tematic literature review and research directions.Psychology & Market- ing40, 8 (2023), 1593–1614. https://doi.org/10.1002/mar.218...

  93. [101]

    Persky and Jennifer D

    Adam M. Persky and Jennifer D. Robinson. 2017. Moving from Novice to Exper- tise and Its Implications for Instruction.American Journal of Pharmaceutical Education81, 9 (Nov. 2017), 6065. https://doi.org/10.5688/ajpe6065

  94. [102]

    Ployhart and Robert J

    Robert E. Ployhart and Robert J. Vandenberg. 2010. Longitudinal research: The theory, design, and analysis of change.Journal of Management36, 1 (2010), 94–120. https://doi.org/10.1177/0149206309352110 Place: US Publisher: Sage Publications

  95. [103]

    Alisha Pradhan and Amanda Lazar. 2021. Hey Google, Do You Have a Per- sonality? Designing Personality and Personas for Conversational Agents. In Proceedings of the 3rd Conference on Conversational User Interfaces (CUI ’21). Association for Computing Machinery, New York, NY, US...

  96. [104]

    Pereira, Armando M

    Luiz Rodrigues, Filipe D. Pereira, Armando M. Toda, Paula T. Palomino, Marcela Pessoa, Leandro Silva Galvão Carvalho, David Fernandes, Elaine H. T. Oliveira, Alexandra I. Cristea, and Seiji Isotani. 2022. Gamification suffers from the novelty effect but benefits from the famil...

  97. [105]

    Rohrer, James Wendt, Jeff Sauro, Frederick Boyle, and Sara Cole

    Christian P. Rohrer, James Wendt, Jeff Sauro, Frederick Boyle, and Sara Cole

  98. [106]

    InProceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI EA ’16)

    Practical Usability Rating by Experts (PURE): A Pragmatic Approach for Scoring Product Usability. InProceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI EA ’16). Association for Computing Machinery, New York, NY, USA, 786–795. ht...

  99. [107]

    Krishna Ronanki, Beatriz Cabrero-Daniel, and Christian Berger. 2024. ChatGPT as a Tool for User Story Quality Evaluation: Trustworthy Out of the Box?. In Agile Processes in Software Engineering and Extreme Programming – Workshops (Lecture Notes in Business Information Processi...

  100. [108]

    Rynes, Marc O

    Sara L. Rynes, Marc O. Orlitzky, and Robert D. Bretz Jr. 1997. Experienced Hiring Versus College Recruiting: Practices and Emerging Trends.Personnel Psychology 50, 2 (1997), 309–339. https://doi.org/10.1111/j.1744-6570.1997.tb00910.x _eprint: https://onlinelibrary.wiley.com/do...

  101. [109]

    Ramteja Sajja, Yusuf Sermet, Muhammed Cikmaz, David Cwiertny, and Ibrahim Demir. 2024. Artificial Intelligence-Enabled Intelligent Assistant for Personalized and Adaptive Learning in Higher Education.Information15, 10 (Sept. 2024), 596. https://doi.org/10.3390/info15100596 Num...

  102. [110]

    Jeff Sauro. 2018. How Large Is the Evaluator Effect in Usability Testing? https: //measuringu.com/evaluator-effect/

  103. [111]

    Schmidt, A

    M. Schmidt, A. A. Tawfik, I. Jahnke, and Y. Earnshaw. 2020. Learner and User Experience Research.Learner and User Experience Research(2020). https: //doi.org/10.59668/36 Publisher: EdTech Books

  104. [112]

    Schmuckler

    Mark A. Schmuckler. 2001. What Is Ecological Validity? A Dimensional Analy- sis.Infancy2, 4 (2001), 419–436. https://doi.org/10.1207/S15327078IN0204_02 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1207/S15327078IN0204_02

  105. [113]

    Valentin Schwind, Stefan Resch, and Jessica Sehrt. 2023. The HCI User Studies Toolkit: Supporting Study Designing and Planning for Undergraduates and Novice Researchers in Human-Computer Interaction. InExtended Abstracts of the 2023 CHI Conference on Human Factors in Computing...

  106. [114]

    2021.An Exploration of Knowledge Practices Among User Experience (UX) Professionals

    Hyojung Julia Seo. 2021.An Exploration of Knowledge Practices Among User Experience (UX) Professionals. Thesis. https://tspace.library.utoronto.ca/handle/ 1807/108705 Accepted: 2021-11-30T16:44:14Z

  107. [115]

    Luyao Shen, Qing Shi, Emily Kuang, Linjie Qiu, Shixu Zhou, Pan Hui, and Ming- ming Fan. 2026. MultiUX: A Human-AI Collaborative Tool to Facilitate Multiple Usability Test Video Analyses.International Journal of Human–Computer Inter- action0, 0 (2026), 1–26. https://doi.org/10....

  108. [116]

    Grace Shin, Yuanyuan Feng, Mohammad Hossein Jarrahi, and Nicci Gafinowitz

  109. [117]

    https: //doi.org/10.1093/jamiaopen/ooy048

    Beyond novelty effect: a mixed-methods exploration into the motivation for long-term activity tracker use.JAMIA Open2, 1 (April 2019), 62–72. https: //doi.org/10.1093/jamiaopen/ooy048

  110. [118]

    Hedderich, AndréS Lucero, and Antti Oulasvirta

    Joongi Shin, Michael A. Hedderich, AndréS Lucero, and Antti Oulasvirta. 2022. Chatbots Facilitating Consensus-Building in Asynchronous Co-Design. InPro- ceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (UIST ’22). Association for Computing Ma...

  111. [119]

    Hedderich, Bartłomiej Jakub Rey, Andrés Lucero, and Antti Oulasvirta

    Joongi Shin, Michael A. Hedderich, Bartłomiej Jakub Rey, Andrés Lucero, and Antti Oulasvirta. 2024. Understanding Human-AI Workflows for Generating Personas. InProceedings of the 2024 ACM Designing Interactive Systems Con- ference (DIS ’24). Association for Computing Machinery...

  112. [120]

    Subin Shin, Jeesun Oh, and Sangwon Lee. 2025. Can LLMs See What I See? A Study on Five Prompt Engineering Techniques for Evaluating UX on a Shop- ping Site. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). Associ...

  113. [121]

    Keng Siau and Weiyu Wang. 2018. Building Trust in Artificial Intelligence, Machine Learning, and Robotics.Cutter Business Technology Journal31, 2 (2018), 47–53. https://ink.library.smu.edu.sg/sis_research/9371

  114. [122]

    It is there, and you need it, so why do you not use it?

    Auste Simkute, Ewa Luger, Michael Evans, and Rhianne Jones. 2024. "It is there, and you need it, so why do you not use it?" Achieving better adoption of AI systems by domain experts, in the case study of natural science research. https://doi.org/10.48550/arXiv.2403.16895 arXiv...

  115. [123]

    Kahn, and Adam Perer

    Venkatesh Sivaraman, Leigh A Bukowski, Joel Levin, Jeremy M. Kahn, and Adam Perer. 2023. Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health Care. InProceedings of the 2023 CHI Conference on Human Factors in Computing...

  116. [124]

    Smith and Cindy Kheng

    Thomas J. Smith and Cindy Kheng. 2021. Reliability of Heuristic Evaluation During Usability Analysis. https://doi.org/10.1007/978-3-030-74614-8_88

  117. [125]

    Ehsan Jahangirzadeh Soure, Emily Kuang, Mingming Fan, and Jian Zhao. 2021. CoUX: Collaborative Visual Analysis of Think-Aloud Usability Test Videos for Digital Interfaces.IEEE Transactions on Visualization and Computer Graphics (2021), 1–11. https://doi.org/10.1109/TVCG.2021.3114822

  118. [126]

    Christensen, and Rebecca E

    JaYoung Sung, Henrik I. Christensen, and Rebecca E. Grinter. 2009. Robots in the wild: Understanding long-term use. In2009 4th ACM/IEEE International Conference on Human-Robot Interaction (HRI). 45–52. https://ieeexplore.ieee. org/document/6256092 ISSN: 2167-2148

  119. [127]

    Movellan, Bret Fortenberry, and Kazuki Aisaka

    Fumihide Tanaka, Javier R. Movellan, Bret Fortenberry, and Kazuki Aisaka

  120. [128]

    InProceedings of the 1st ACM SIGCHI/SIGART conference on Human-robot interaction (HRI ’06)

    Daily HRI evaluation at a classroom environment: reports from dance interaction experiments. InProceedings of the 1st ACM SIGCHI/SIGART conference on Human-robot interaction (HRI ’06). Association for Computing Machinery, New York, NY, USA, 3–9. https://doi.org/10.1145/1121241.1121245

  121. [129]

    Armin Tanovic. 2025. AI Thematic Analysis + the Best Tools for the Job. https: //maze.co/collections/ai/thematic-analysis/

  122. [130]

    Sirui Tao, Ivan Liang, Cindy Peng, Zhiqing Wang, Srishti Palani, and Steven P. Dow. 2025. DesignWeaver: Dimensional Scaffolding for Text-to-Image Product Design. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing...

  123. [131]

    Phuoc Thai, Fernando Montalvo, Promise Stephens, Jordan Sasser, and Sean Hinkle. 2025. Smarter UX Evaluations? Comparing AI and Human Experts in Usability Analysis.Proceedings of the Human Factors and Ergonomics So- ciety Annual Meeting69, 1 (Sept. 2025), 1219–1222. https://do...

  124. [132]

    curse of knowledge

    Jonathan G. Tullis and Brennen Feder. 2022. The “curse of knowledge” when predicting others’ knowledge.Memory & Cognition51, 5 (Dec. 2022), 1214–1234. https://doi.org/10.3758/s13421-022-01382-3 Publisher: Springer

  125. [133]

    Janne Tyni, Aatu Turunen, Juho Kahila, Roman Bednarik, and Matti Tedre. 2024. Can ChatGPT Match the Experts? A Feedback Comparison for Serious Game Development.International Journal of Serious Games11, 2 (June 2024), 87–106. https://doi.org/10.17083/ijsg.v11i2.744

  126. [134]

    UserTesting. 2020. UserTesting: How to Analyze and Share Results. https://help.usertesting.com/hc/en-us/articles/115003371891-How-to- analyze-and-share-results

  127. [135]

    UserTesting. 2021. UserTesting: The Human Insight Platform. https://www. usertesting.com/

  128. [136]

    2022.UXTesting: The Best and Most Intelligent User Experience Solution for Enterprises

    UXTesting. 2022.UXTesting: The Best and Most Intelligent User Experience Solution for Enterprises. https://www.uxtesting.io/user-experience-program

  129. [137]

    Zombor Varnagy-Toth. 2022. 2021 — In-demand skills for UX Researchers. https://uxplanet.org/2021-in-demand-skills-for-ux-researchers-f997753a9f8

  130. [138]

    Bernstein, and Ranjay Krishna

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making.Proc. ACM Hum.-Comput. Interact.7, CSCW1 (April 2023), 129:1–129:38. ht...

  131. [139]

    Dana Vertsberger, Navot Naor, and Mirène Winsberg. 2022. Adolescents’ Well-being While Using a Mobile Artificial Intelligence–Powered Acceptance Commitment Therapy Tool: Evidence From a Longitudinal Study.JMIR AI 2022;1(1):e38171 https://ai.jmir.org/2022/1/e38171(Nov. 2022). h...

  132. [140]

    Smith, and Tom Carey

    Karel Vredenburg, Ji-Ye Mao, Paul W. Smith, and Tom Carey. 2002. A Survey of User-Centered Design Practice. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’02). Association for Computing Machinery, New York, NY, USA, 471–478. https://doi.org/...

  133. [141]

    Wen-Fan Wang, Chien-Ting Lu, Nil Ponsa i Campanyà, Bing-Yu Chen, and Mike Y. Chen. 2025. AIdeation: Designing a Human-AI Collaborative Ideation System for Concept Designers. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association f...

  134. [142]

    White, Pejman Mirza-babaei, Graham McAllister, and Judith Good

    Gareth R. White, Pejman Mirza-babaei, Graham McAllister, and Judith Good

  135. [143]

    In CHI ’11 Extended Abstracts on Human Factors in Computing Systems (CHI EA ’11)

    Weak inter-rater reliability in heuristic evaluation of video games. In CHI ’11 Extended Abstracts on Human Factors in Computing Systems (CHI EA ’11). Association for Computing Machinery, New York, NY, USA, 1441–1446. https://doi.org/10.1145/1979742.1979788

  136. [144]

    Kate Whiting. 2023. Is AI making you suffer from FOBO? Here’s what can help. https://www.weforum.org/stories/2023/12/ai-fobo-jobs-anxiety/

  137. [145]

    Kelvin K. L. Wong, Yong Han, Yifeng Cai, Wumin Ouyang, Hemin Du, and Chao Liu. 2025. From Trust in Automation to Trust in AI in Healthcare: A 30-Year Longitudinal Review and an Interdisciplinary Framework.Bioengineering12, 10 (Oct. 2025), 1070. https://doi.org/10.3390/bioengin...

  138. [146]

    Wei Xiang, Hanfei Zhu, Suqi Lou, Xinli Chen, Zhenghua Pan, Yuping Jin, Shi Chen, and Lingyun Sun. 2024. SimUser: Generating Usability Feed- back by Simulating Various Users Interacting with Mobile Applications. In Proceedings of the CHI Conference on Human Factors in Computing...

  139. [147]

    Cindy Xiong, Lisanne Van Weelden, and Steven Franconeri. 2020. The Curse of Knowledge in Visual Data Communication.IEEE Transactions on Visualization and Computer Graphics26, 10 (Oct. 2020), 3051–3062. https://doi.org/10.1109/ TVCG.2019.2917689

  140. [148]

    Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. 2020. Re- examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design. InProceedings of the 2020 CHI Conference on Human Factors in Com- puting Systems(New York, NY, USA)(CHI ’20). Associati...

  141. [149]

    Kun Yu, Shlomo Berkovsky, Ronnie Taib, Dan Conway, Jianlong Zhou, and Fang Chen. 2017. User Trust Dynamics: An Investigation Driven by Differences in System Performance. InProceedings of the 22nd International Conference on Intelligent User Interfaces (IUI ’17). Association fo...

  142. [150]

    Michelle Zhao, Reid Simmons, and Henny Admoni. 2022. The Role of Adaptation in Collective Human–AI Teaming.Topics in Cognitive Sciencen/a, n/a (2022). https://doi.org/10.1111/tops.12633 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/tops.12633

  143. [151]

    Zicheng Zhu, Yugin Tan, Naomi Yamashita, Yi-Chieh Lee, and Renwen Zhang

  144. [152]

    InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25)

    The Benefits of Prosociality towards AI Agents: Examining the Effects of Helping AI Agents on Human Well-Being. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, 1–18. https://doi.o...

  145. [153]

    moon” or “sun

    John Zimmerman, Changhoon Oh, Nur Yildirim, Alex Kass, Teresa Tung, and Jodi Forlizzi. 2020. UX Designers Pushing AI in the Enterprise: A Case for Adaptive UIs.Interactions28, 1 (Dec. 2020), 72–77. https://doi.org/10.1145/ 3436954 A Appendix Table 5: Prompts used to generate a...

  146. [154]

    Customize the home screen so that UV is visible at the top

  147. [155]

    Vacation

    Add Copenhagen to the list of saved locations and change the nickname of this city to “Vacation. ”

  148. [156]

    Find the weather forecast for Copenhagen on Oct 16

  149. [157]

    Vacation

    Remove the nickname “Vacation” and reset it to the default name. 7:27 14:15 7:56 7:29 11:08 4 6 5 7 5 6 6 6 9 7 GenAI Website

  150. [158]

    Create a poster to advertise a fundraising bake sale in a local community center that includes the date, time, and location of the bake sale and at least 3 different visualizations

  151. [159]

    18:08 10:14 13:49 17:07 18:00 6 7 6 7 9 8 9 8 9 13 VR Game

    Export the poster, then generate a catchy caption and hashtags to post on Instagram. 18:08 10:14 13:49 17:07 18:00 6 7 6 7 9 8 9 8 9 13 VR Game

  152. [160]

    Blocking

    Select the “Blocking” game from the menu and play it

  153. [161]

    Volleyball

    Select the “Volleyball” game from the menu and play it

  154. [162]

    Just wanted to let you know you should report each and every minor and major usability issue to me

    Select the “Skeet” game from the menu and play it. 7:51 9:42 10:56 9:06 11:27 5 5 7 6 6 7 8 7 7 6 Multi-Session Study of UX Evaluators Using Conversational AI Agents CHI ’26, April 13–17, 2026, Barcelona, Spain Table 7: Categories of Messages sent by participants. Session Appe...

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.