REVIEW 3 major objections 5 minor 69 references
VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Proactive, source-traceable scaffolding helps fiction writers discover and fill domain-knowledge gaps they cannot name on their own.
desk verdict A thoughtful mixed-initiative writing study with an honest baseline and one honest design flaw: the knowledge the system surfaces is never checked for factual accuracy, so the headline gap-filling claim rests on perception, not verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the division of cognitive labor. VeriForge's front end couples three interactions: proactive inline highlights triggered after a five-second pause when an Alignment Engine finds a source-backed term whose graph neighborhood suggests useful knowledge; dual-stream queries that present the same retrieval as both a conversational text stream and source-anchored Knowledge Cards; and a spatial Knowledge Canvas whose nodes carry provenance and whose topology feeds back into the Semantic Frame for later retrievals. Behind them, a graph-based retrieval-augmented generation pipeline stores domain facts as nodes and edges with exact source passages bound at ingestion, so every surfaced fact remains traceable to its primary source. The Alignment Engine is the piece that decides what counts as a candidate gap, using a hand-coded relational-complexity taxonomy rather than error detection.
What would settle it
Run VeriForge on drafts that have been seeded with known domain errors, such as a still-bleeding corpse described as decomposing or an anachronistic vaccine rollout. If the proactive highlights fail to flag a large share of seeded errors, or flag a large share of correct statements, the claim that the system reveals latent gaps instead of injecting noise would be refuted.
Extended reading notes
Core claim
The central claim is that proactive, flow-aware, source-traceable scaffolding can make authors aware of latent domain-knowledge gaps that reactive tools cannot surface, without taking narrative authorship away. The authors ground this in the knowledge-transforming model of composition: writing well in an unfamiliar domain requires shuttling between factual content and narrative rhetoric, and the illusion of explanatory depth blocks that shuttle precisely where it matters most. VeriForge operationalizes a discovery-synthesis split, in which the system takes initiative over finding and structuring external facts while the author takes initiative over composing, and the study's pattern of results, higher blind-spot alerting, more exploration, higher perceived creativity support, higher expert-rated domain competence, and unchanged autonomy, is offered as preliminary evidence that the split works.
Load-bearing premise
The load-bearing premise is that the proactive highlights point at real, useful knowledge gaps most of the time; the paper does not measure how often they are right, so if highlights are frequently irrelevant or misleading the reported benefits would not survive.
Editorial extensions
If this is right
- Authors using VeriForge in the study rated blind-spot alerting far higher than they did in a reactive baseline sharing the same retrieval backend, and they issued more queries, so proactive cues changed search behavior rather than merely improving satisfaction.
- Blind expert raters scored VeriForge passages higher on perceived authorial domain competence, with positive but uncorrected trends on natural integration and specificity, suggesting the paradigm can affect cold-start prose quality.
- Participants' perceived authorial autonomy was identical across conditions, supporting the claim that a discovery-synthesis split preserves agency even when the system takes initiative over knowledge discovery.
- One-week delayed recall showed that participants retained more domain terms and a higher proportion of the terms they had used under VeriForge, consistent with deeper encoding rather than mere exposure.
- Only 1 of 12 participants noticed a deliberately mismatched source passage, so visible provenance can act as an authority cue that suppresses verification even when the underlying facts are trustworthy.
Reading between the lines
- If the discovery-synthesis split is the active ingredient, the same scaffolding logic should transfer to journalism, documentary filmmaking, legal drafting, and clinical writing, wherever the output's value depends on a human voice interpreting expert material; the paper gestures at this transfer but does not test it.
- The unmeasured precision of highlight selection is the critical knob: a risk-adaptive design that raises epistemic friction when consequences are higher could convert the provenance-as-authority failure into a feature while preserving flow in low-stakes writing.
- A direct testable extension would replace the AI description rather than the source passage in the deception experiment, isolating whether authors can detect fabricated domain claims; the paper notes this is not tested.
- The retention-rate gains suggest that writing with proactive factual scaffolding may function like writing-to-learn, but the 30-minute cold-start task cannot show whether the effect compounds over a full manuscript; longitudinal drafts would settle that.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces VeriForge, a mixed-initiative writing environment for fiction authors working in unfamiliar domains. The system combines proactive inline highlighting that flags candidate knowledge gaps, dual-stream querying that pairs conversational responses with source-anchored Knowledge Cards, and a spatial Knowledge Canvas for organizing discovered facts, all backed by a graph-based retrieval-augmented generation pipeline over user-uploaded source material. The design is motivated by formative interviews with nine fiction writers and evaluated in a within-subjects user study (N=12) against a strengthened baseline that shares the same retrieval backend but omits the proactive and canvas-integration mechanisms. The authors report that VeriForge increases blind-spot awareness, query volume, canvas use, creativity support, delayed recall, and one of three expert-rated prose dimensions (perceived authorial domain competence), while preserving perceived authorial autonomy.
Significance. If the findings hold, VeriForge addresses a genuinely underexplored problem: helping writers discover domain knowledge they cannot self-diagnose, rather than only retrieving answers to explicit queries. The manuscript has several real strengths: the formative study grounds the design in observed writer behavior; the within-subjects design with a strengthened baseline is a thoughtful attempt to isolate interaction design from retrieval quality; FDR correction is applied within predefined families; the deception experiment directly probes epistemic trust; and the delayed-recall analysis is a useful exploratory addition. The reported effect sizes are large and consistent across several behavioral and self-report measures. However, the central construct—mitigating latent knowledge gaps—is never validated against ground truth: highlight precision, Knowledge Card factual accuracy, and the factual correctness of final passages are not measured. Given the deception result showing that users do not inspect provided sources, the observed benefits could in principle stem from plausible but inaccurate scaffolding.
major comments (3)
- [Section 4.2.1 and Section 6.3] The central mechanism of VeriForge is proactive highlighting, defined in Section 4.2.1 as a 'low-cost hypothesis about knowledge value.' The paper reports no precision, recall, or false-positive analysis for the Alignment Engine's highlight suggestions, and no factual-accuracy evaluation of the LLM-generated Knowledge Card descriptions against their quoted source passages. The deception experiment in Section 6.3 shows that 11 of 12 participants accepted a card whose source passage was replaced with unrelated text, so users are unlikely to catch inaccurate AI-generated descriptions. Without an audit showing that highlights and cards correspond to genuine, source-supported knowledge, the claim that VeriForge 'mitigates latent knowledge gaps' is not yet established; the evidence supports only the weaker claim that it changes perceived gap awareness and exploration behavior. Please add an objective evaluation (e.g., expert or LLM-assisted fact-checking of a sample of highlights and cards against the source corpora, with error rates) or explicitly restrict the contribution to perceived gap recognition.
- [Section 5.5.3 and Section 6.4] RQ4 is the only outcome measure tied to the 'domain grounding' claim, but it is operationalized solely through two raters' holistic judgments of 'perceived authorial domain competence.' The raters are described as published authors, not as HEMA or TCMA domain experts, and they were not asked to verify any factual claims in the passages against the uploaded source corpora. The manuscript therefore never measures whether the knowledge gaps were actually filled correctly. Given that the deception study shows readers do not verify sources, the observed difference in perceived competence could reflect surface-level use of domain terminology rather than accurate grounding. I recommend adding a fact-based outcome measure—for example, the proportion of domain-specific claims in the final passages that are accurate according to the provided sources, or a count of factual errors—or, failing that, changing the abstract's 'stronger domain grounding' to 'higher perceived authorial domain competence.'
- [Section 5.5.4, Section 6.4, and Figure 5] The expert-rating evidence is thinner than the abstract implies. Of the three rated dimensions, only 'perceived authorial domain competence' survives FDR correction (q = .015); natural integration (q = .065) and specificity (q = .065) are positive trends that do not meet the corrected threshold, and the composite score reported in Section 6.4 is explicitly uncorrected. The single surviving dimension is also the most perceptual and least tied to factual accuracy. This is not a fatal flaw for a preliminary study, but it should be stated clearly in the abstract and conclusion so that readers do not infer broad support for 'stronger domain grounding' across all quality dimensions.
minor comments (5)
- [Section 5.5.3] The term 'expert raters' in the abstract is stronger than the description in Section 5.5.3, where the raters are 'two published authors who were not study participants' and not necessarily domain experts. Please align the terminology throughout.
- [Figure 5 caption and Section 6.4] The caption notes that the expert composite is uncorrected, but the main text should also prominently state that natural integration and specificity did not survive FDR correction, since these are reported in the same paragraph as the composite result.
- [Section 5.5.2 and Section 6.2] The delayed-recall analyses are described as exploratory and are not FDR-corrected, yet they are presented with strong effect sizes and p-values. Please clearly label these results as exploratory in the figures or tables, not only in the prose.
- [Section 6.3] For the perceived authorial autonomy result (M = 6.17 vs. M = 6.17, p = 1.0), please report the number of nonzero differences and the distribution of signed ranks; with N = 12, a p-value of exactly 1.0 can arise from mostly tied responses, and this information is needed to interpret the null result.
- [Section 4.4] The implementation section names Qwen3.5-Flash and BAAI/bge-m3 but does not report the entity-extraction or alignment prompt templates or any examples of highlight failures. Including representative positive and negative highlight examples in the supplemental material would help readers calibrate the reliability of the proactive mechanism.
Circularity Check
No significant circularity: VeriForge's claims rest on an empirical within-subjects comparison against a deliberately conservative baseline, not on fitted parameters or self-citations.
full rationale
The paper contains no mathematical derivation or fitted-parameter feedback loop, so none of the enumerated circularity patterns apply. The central evaluation (Sections 5-6) is a controlled within-subjects study comparing VeriForge with a 'strengthened baseline' that shares 'the same graph-based retrieval and Semantic Frame backend as VeriForge' (Section 5.1); sharing infrastructure across conditions is a conservative control condition, not circular reasoning. Proactive highlighting is explicitly scoped as 'a low-cost hypothesis about knowledge value, not a claim that the author has made an error' (Section 4.2.1), so self-reported blind-spot alerting measures the intervention itself rather than a prediction forced from fitted data. RQ4's 'domain grounding' is operationalized as raters' perceived authorial competence (Section 5.5.3), and the paper itself cautions that these ratings are 'initial evidence of domain grounding' (Section 7.3); this is a measurement-validity limitation, not a circular reduction. The design goals (Section 3.2) were derived from a separate formative study, not from the outcome data, and the load-bearing psychological constructs (IOED, knowledge-transforming) are cited from external literature ([55], [11]). No self-citation chain is load-bearing, and the deception experiment (Section 6.3) provides an independent, potentially disconfirming observation about trust calibration. Overall, the derivation chain is empirical and self-contained; concerns about highlight precision and factual verification are threats to construct validity, not circularity.
Assumptions & free parameters
free parameters (3)
- highlight_delay_seconds =
5
- graph_expansion_hops =
1
- relational_complexity_filter
assumptions (5)
- domain assumption The Illusion of Explanatory Depth applies to fiction writers in unfamiliar domains, causing latent knowledge gaps.
- domain assumption The knowledge-transforming model of writing (Bereiter and Scardamalia) accurately describes fiction drafting.
- domain assumption Source traceability is the main lever for restoring author trust in retrieved facts.
- domain assumption The strengthened baseline with a shared retrieval backend isolates interaction-design effects.
- domain assumption Self-reported and expert-rated metrics are valid proxies for gap recognition, creativity support, and domain grounding.
Cite this review
Pith. "Pith review of VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding." pith.science (2026). https://pith.science/paper/QTVBE5DY
@misc{pith2026260809698,
author = {Pith},
title = {Pith review of: VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding},
year = {2026},
howpublished = {\url{https://pith.science/paper/QTVBE5DY}},
note = {Machine review of arXiv:2608.09698}
}
read the original abstract
Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot yet smell of decay, weaving domain expertise into the fabric of invented worlds. Current AI writing tools offer limited support for discovering and integrating unfamiliar domain knowledge into narrative. They require explicit queries that authors cannot formulate, generate finished prose that risks homogenizing voice, or assist only within the boundaries of what authors already know. We argue that AI should reveal latent knowledge gaps to writers while preserving their agency to transform discovered knowledge into authentic prose. Grounded in formative interviews with 9 fiction writers, we present VeriForge, a mixed-initiative writing system that divides cognitive labor so that the system assumes initiative over domain discovery while the author retains full initiative over narrative synthesis. VeriForge realizes this through three complementary mechanisms. Proactive inline highlighting flags potential knowledge gaps as authors draft. Dual-stream querying pairs conversational responses with source-anchored Knowledge Cards for direct fact extraction. A spatial Knowledge Canvas allows authors to organize and connect discovered knowledge across their writing. These mechanisms are powered by a graph-based retrieval-augmented generation pipeline grounded in domain-specific source materials. A within-subjects user study (N=12) provides preliminary evidence that this paradigm helps authors recognize previously overlooked knowledge gaps, supports creative exploration, and is perceived by expert raters to produce passages with stronger domain grounding in a controlled cold-start writing task.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[2]
Erik Albæk. 2011. The Interaction Between Experts and Journalists in News Journalism.Journalism12, 3 (2011), 335–348. doi:10.1177/1464884910392851
-
[3]
Alibaba Cloud. 2026. Qwen3.5-Flash Model Service. https://bailian.console.aliyun. com/. Accessed: 2026-03-29. Available via Alibaba Cloud Model Studio
work page 2026
-
[4]
Kholod Alsufiani, Simon Attfield, and Leishi Zhang. 2017. Towards an instru- ment for measuring sensemaking and an assessment of its theoretical features. In Proceedings of the 31st British Computer Society Human Computer Interaction Con- ference(Sunderland, UK)(HCI ’17). BCS Learning & Development Ltd., Swindon, GBR, Article 86, 5 pages. doi:10.14236/ewi...
-
[5]
Rifat Mehreen Amin, Oliver Hans Kühle, Daniel Buschek, and Andreas Butz
-
[6]
Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. 2024. Homog- enization Effects of Large Language Models on Human Creative Ideation. In Proceedings of the 16th Conference on Creativity & Cognition(Chicago, IL, USA) (C&C ’24). Association for Computing Machinery, New York, NY, USA, 413–425. doi:10.1145/3635636.3656204
arXiv 2024
-
[8]
Kathleen Arnold, Sharda Umanath, Kara Thio, Walter Reilly, Mark Mcdaniel, and Elizabeth Marsh. 2017. Understanding the Cognitive Processes Involved in Writing to Learn.Journal of Experimental Psychology: Applied23 (04 2017), 115–127. doi:10.1037/xap0000119
-
[9]
Roland Barthes. 1989. The Reality Effect. InThe Rustle of Language. University of California Press, Berkeley, CA, USA, 141–148
work page 1989
-
[10]
Nicholas J. Belkin. 1980. Anomalous states of knowledge as a basis for information retrieval.Canadian Journal of Information Science5, 1 (1980), 133–143
work page 1980
Show all 69 references
-
[11]
1987.The Psychology of Written Compo- sition
Carl Bereiter and Marlene Scardamalia. 1987.The Psychology of Written Compo- sition. Lawrence Erlbaum Associates, Hillsdale, NJ
1987
-
[12]
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024. Art or Artifice? Large Language Models and the False Promise of Creativity. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’2...
2024
-
[13]
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu
-
[15]
arXiv:2402.03216 [cs.CL] https://arxiv.org/abs/2402.03216
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv:2402.03216 [cs.CL] https://arxiv.org/abs/2402.03216
-
[16]
Jean-Peïc Chou, Alexa Fay Siu, Nedim Lipka, Ryan Rossi, Franck Dernoncourt, and Maneesh Agrawala. 2023. TaleStream: Supporting Story Ideation with Trope Knowledge. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology(San Francisco, CA, USA)(...
2023
-
[17]
Daixuan Cheng, Shaohan Huang, and Furu Wei. 2024. Adapting Large Language Models to Domains via Reading Comprehension. arXiv:2309.09530 [cs.CL] https: //arxiv.org/abs/2309.09530
2024 arXiv
-
[19]
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching Stories with Generative Pretrained Language Models. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for ...
2022
-
[20]
Yours is better!
Nicola Dell, Vidya Vaidyanathan, Indrani Medhi, Edward Cutrell, and William Thies. 2012. "Yours is better!": participant response bias in HCI. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Austin, Texas, USA)(CHI ’12). Association for Computing M...
2012
-
[21]
Susan De La Paz. 2005. Effects of Historical Reasoning Instruction and Writing Strategy Mastery in Culturally and Academically Diverse Middle School Class- rooms.Journal of Educational Psychology97, 2 (2005), 139–156. doi:10.1037/0022- 0663.97.2.139
2005 doi
-
[22]
Michael Fleming and Robin Cohen. 2001. A User Modeling Approach to Deter- mining System Initiative in Mixed-Initiative AI Systems. InProceedings of the 8th International Conference on User Modeling 2001 (UM ’01). Springer-Verlag, Berlin, Heidelberg, 54–63
2001
-
[23]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2025. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 [cs.CL] https://arxiv....
2025 arXiv
-
[24]
Green and Timothy C
Melanie C. Green and Timothy C. Brock. 2000. The Role of Transportation in the Persuasiveness of Public Narratives.Journal of Personality and Social Psychology 79, 5 (2000), 701–721. doi:10.1037/0022-3514.79.5.701
2000 doi
-
[25]
Kexue Fu, Jingfei Huang, Long Ling, Sumin Hong, Yihang Zuo, Ray LC, and Toby Jia jun Li. 2026. Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing. arXiv:2509.13646 [cs.HC] https://arxiv.org/abs/2509.13646
2026
-
[26]
Marijn Haverbeke. 2015. ProseMirror: A toolkit for building rich-text editors on the web. https://prosemirror.net/. Accessed: 2026-03-29
2015
-
[27]
Tanay Kumar Gupta, Tushar Goel, Ishan Verma, Lipika Dey, and Sachit Bhardwaj
-
[28]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Pittsburgh, Pennsylvania, USA)(CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. doi:10.1145/302979.303030
1999
-
[29]
Iqbal and Brian P
Shamsi T. Iqbal and Brian P. Bailey. 2007. Understanding and developing models for detecting and differentiating breakpoints during interactive tasks. InProceed- ings of the SIGCHI Conference on Human Factors in Computing Systems(San Jose, California, USA)(CHI ’07). Associatio...
2007
-
[30]
Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi
Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-Retriever: Retrieval- Augmented Generation for Textual Graph Understanding and Question An- swering. arXiv:2402.07630 [cs.LG] https://arxiv.org/abs/2402.07630
2024 arXiv
-
[32]
James Kaufman, John Baer, Jason Cole, and Janel Sexton. 2008. A Comparison of Expert and Nonexpert Raters Using the Consensual Assessment Technique. Creativity Research Journal - CREATIVITY RES J20 (04 2008), 171–178. doi:10. 1080/10400410802059929
2008
-
[33]
Iren Irbe. 2025. Investigating Tacit Knowledge Transfer in Public Sector Work- places. InProceedings of the 36th Annual Conference of the European Association of Cognitive Ergonomics (ECCE ’25). Association for Computing Machinery, New York, NY, USA, Article 32, 5 pages. doi:1...
2025
-
[34]
Matthias Kraus, Nicolas Wagner, and Wolfgang Minker. 2020. Effects of Proactive Dialogue Strategies on Human-Computer Trust. InProceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization(Genoa, Italy)(UMAP ’20). Association for Computing Machinery, ...
2020
-
[35]
Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of Explanatory Debugging to Personalize Interactive Machine Learning. InProceedings of the 20th International Conference on Intelligent User Interfaces (IUI ’15). Association for Computing Mach...
2015
-
[36]
Burak Korkmaz. 2021. Vue Flow: A customizable Vue 3 component for building node-based editors and diagrams. https://vueflow.dev/. Accessed: 2026-03-29
2021
-
[37]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.1140...
2021 arXiv
-
[38]
Lim and Anind K
Brian Y. Lim and Anind K. Dey. 2009. Assessing Demand for Intelligibility in Context-Aware Applications. InProceedings of the 11th International Conference VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding UIST ’26, November 02–...
2009
-
[39]
Florian Lehmann. 2023. Mixed-Initiative Interaction with Computational Gen- erative Systems. InExtended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany)(CHI EA ’23). Associa- tion for Computing Machinery, New York, NY, USA, Article 5...
2023
-
[40]
Tom Lumley. 2002. Assessment Criteria in a Large-Scale Writing Test: What Do They Really Mean to the Raters?Language Testing - LANG TEST19 (07 2002), 246–276. doi:10.1191/0265532202lt230oa
2002 doi
-
[41]
Haoran Luo, Haihong E, Guanting Chen, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng, Zemin Kuang, Meina Song, Yifan Zhu, and Luu Anh Tuan. 2025. HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation. arXiv:2503.21322 [cs.AI] ...
2025
-
[42]
Literature & Latte. 2026. Scrivener: Overview. https://www.literatureandlatte. com/scrivener/overview. Accessed: 2026-07-08
2026
-
[43]
Yu Mei, Yuanxi Wang, Shiyi Wang, Qingyang Wan, Zhuojun Li, Chun Yu, Weinan Shi, and Yuanchun Shi. 2025. InterQuest: A Mixed-Initiative Framework for Dynamic User Interest Modeling in Conversational Search. InProceedings of the 38th Annual ACM Symposium on User Interface Softwa...
2025
-
[44]
Microsoft Corporation. 2012. TypeScript: JavaScript with syntax for types. https: //www.typescriptlang.org/. Accessed: 2026-03-29
2012
-
[46]
George E. Newell. 2006. Writing to Learn: How Alternative Theories of School Writing Account for Student Performance. InHandbook of Writing Research. Guilford Press, New York, NY, USA, 235–247
2006
-
[47]
OpenAI. 2026. ChatGPT. https://chatgpt.com/. Accessed: 2026-03-29
2026
-
[48]
Neo4j, Inc. 2026. Neo4j Graph Database. https://neo4j.com/. Accessed: 2026-03-29. Enterprise or Community Edition
2026
-
[49]
Syemin Park, Soobin Park, and Youn-kyung Lim. 2026. Constella: Supporting Sto- rywriters’ Interconnected Character Creation through LLM-based Multi-Agents. ACM Trans. Comput.-Hum. Interact.33, 3 (Feb. 2026), 1–56. doi:10.1145/3796234
2026 doi
-
[50]
Petty and John T
Richard E. Petty and John T. Cacioppo. 1986. The Elaboration Likelihood Model of Persuasion.Advances in Experimental Social Psychology19 (1986), 123–205. doi:10.1016/S0064-2601(08)60214-2
1986 doi
-
[51]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu
-
[52]
IEEE Transactions on Knowledge and Data Engineering36, 7 (July 2024), 3580–3599
Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering36, 7 (July 2024), 3580–3599. doi:10.1109/tkde.2024.3352100
2024
-
[53]
Sebastián Ramírez. 2018. FastAPI framework, high performance, easy to learn, fast to code, ready for production. https://fastapi.tiangolo.com/. Accessed: 2026-03-29
2018
-
[54]
Anyi Rao, Jean-Peïc Chou, and Maneesh Agrawala. 2024. ScriptViz: A Visualiza- tion Tool to Aid Scriptwriting based on a Large Movie Database. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburgh, PA, USA)(UIST ’24). Association f...
2024
-
[56]
Hua Xuan Qin, Guangzhi Zhu, Mingming Fan, and Pan Hui. 2025. Toward Per- sonalizable AI Node Graph Creative Writing Support: Insights on Preferences for Generative AI Features and Information Presentation Across Story Writing Processes. InProceedings of the 2025 CHI Conference...
2025
-
[57]
Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu...
2024
-
[58]
Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling Multilevel Exploration and Sensemaking with Large Language Models. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology(San Francisco, CA, USA)(UIST ’23). Ass...
2023
-
[59]
Leonid Rozenblit and Frank Keil. 2002. The misunderstood limits of folk science: An illusion of explanatory depth.Cognitive Science26, 5 (2002), 521–562. doi:10. 1207/s15516709cog2605_1
2002
-
[60]
2000.Picturing Culture: Explorations of Film and Anthropology
Jay Ruby. 2000.Picturing Culture: Explorations of Film and Anthropology. Univer- sity of Chicago Press, Chicago
2000
- [61]
-
[62]
Tiptap GmbH. 2019. Tiptap: The headless editor framework for web artisans. https://tiptap.dev/. Accessed: 2026-03-29
2019
-
[63]
Shyam Sundar
S. Shyam Sundar. 2008. The MAIN Model: A Heuristic Approach to Understanding Technology Effects on Credibility. InDigital Media, Youth, and Credibility. The MIT Press, Cambridge, MA, USA, 73–100. doi:10.1162/dmal.9780262562324.073
2008 doi
-
[64]
Shayan Talaei, Meijin Li, Kanu Grover, James Kent Hippler, Diyi Yang, and Amin Saberi. 2025. StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST ’25)....
2025 doi
-
[65]
It Felt Like Having a Second Mind
Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2024. “It Felt Like Having a Second Mind”: Investigating Human-AI Co-creativity in Prewriting with Large Language Models.Proc. ACM Hum.-Comput. Interact.8, CSCW1 (2024), 1–26. doi:10.1145/3637361
2024 doi
-
[66]
Xiyuan Wang, Yi-Fan Cao, Junjie Xiong, Sizhe Chen, Wenxuan Li, Junjie Zhang, and Quan Li. 2025. ClueCart: Supporting Game Story Interpretation and Narrative Inference from Fragmented Clues. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25...
2025
-
[67]
Rama Adithya Varanasi, Batia Mishan Wiesenfeld, and Oded Nov. 2025. AI Rivalry as a Craft: How Resisting and Embracing Generative AI Are Reshaping the Writing Profession. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for ...
2025
-
[68]
Artem Vizniuk, Grygorii Diachenko, Ivan Laktionov, Agnieszka Siwocha, Min Xiao, and Jacek Smoląg. 2025. A Comprehensive Survey of Retrieval-Augmented Large Language Models for Decision Making in Agriculture: Unsolved Problems and Research Opportunities.Journal of Artificial In...
2025 doi
-
[69]
Jochen Wulf and Juerg Meierhofer. 2024. Exploring the Potential of Large Language Models for Automation in Technical Customer Service. arXiv:2405.09161 [econ.GN] https://arxiv.org/abs/2405.09161
2024 arXiv
-
[70]
Evan You and Vue Core Team. 2014. Vue.js: The Progressive JavaScript Frame- work. https://vuejs.org/. Accessed: 2026-03-29
2014
-
[71]
Mark J. P. Wolf. 2012.Building Imaginary Worlds: The Theory and History of Subcreation. Routledge, New York, NY
2012
-
[72]
Zeqiu Wu, Ryu Parish, Hao Cheng, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. INSCIT: Information-Seeking Con- versations with Mixed-Initiative Interactions. arXiv:2207.00746 [cs.CL] https: //arxiv.org/abs/2207.00746
2023 arXiv
-
[75]
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: Story Writing With Large Language Models. InProceedings of the 27th International Conference on Intelligent User Interfaces(Helsinki, Finland)(IUI ’22). Association for Computing Machinery, New York, NY, ...
2022 doi
-
[76]
Chengbo Zheng, Yuanhao Zhang, Zeyu Huang, Chuhan Shi, Minrui Xu, and Xiaojuan Ma. 2024. DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology(Pit...
2024
-
[2024]
InProceedings of the 2nd International Workshop on Knowledge Graphs for Sus- tainability (KG4S 2024) (CEUR Workshop Proceedings, Vol
Knowledge Graph aided LLM based ESG Question-Answering from News. InProceedings of the 2nd International Workshop on Knowledge Graphs for Sus- tainability (KG4S 2024) (CEUR Workshop Proceedings, Vol. 3753). CEUR-WS.org, Hersonissos, Greece, paper6. https://ceur-ws.org/Vol-3753...
2024
-
[2025]
InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25)
Composable Prompting Workspaces for Creative Writing: Exploration and Iteration Using Dynamic Widgets. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). Association for Computing Machinery, New York, NY, USA, 1–11...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.