REVIEW 3 major objections 4 minor 99 references
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Dragging two triangle symbols like toys can control an AI storyteller in both directions.
desk verdict Genuinely new toy-playing interaction for AI storytelling, honestly evaluated in the user study, but the abstract overstates the technical results by claiming bidirectional superiority over GPT-4o when the text-to-motion pipeline was never tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The translational layer is the system's core object. It consists of a continuous action embedding—a vector in a sentence-embedding space (SBERT) representing one of 31 two-character action verbs such as 'chase' or 'hug'—together with a boolean active-character indicator that says which character is the agent. LSTM models project recorded symbol trajectories into this layer; two further LSTM models generate frame-by-frame symbol coordinates from it; and a soft-prompt pipeline interpolates the top-k relevant base-action tokens into the LLM's input embedding space to condition story text on the recognized action. The reverse direction, text-to-action, embeds the user's sentence in the same space, interpolates the nearest base actions, and asks the LLM which character is active. That shared layer is what lets motion and text condition each other without training one joint multimodal model.
What would settle it
Ask users to act out a domain-specific two-character event, such as a basketball pass or a tackle, in Toyteller and rate whether the generated story sentence matches their intent; if recognition routinely maps the motion to a base action it does not fit, as the reported athletic-story bias suggests, the coverage assumption fails and the central claim goes with it.
Extended reading notes
Core claim
Toyteller's central claim is that anthropomorphized motion can serve as a bidirectional interface to story generation. The system recognizes which of 31 base two-character actions a user's symbol motion expresses, converting that motion into a continuous action embedding and an active-character indicator; from the same representation it can generate matching story text and can generate new symbol motions conditioned on text. In the technical evaluation, the paper reports that this pipeline ranks the gold-standard action higher and assigns it more weight than GPT-4o given either rendered frames or coordinate text, that its action-conditioned motion generation scores higher on alignment and realism (proactive case), and that motion-conditioned story text is more aligned and more novel/interesting, all with drastically lower latency. The user study adds that toy-playing is perceived as vague and good for half-formed ideas, and that users naturally combine it with natural language prompts to pin down specifics. The claim, stated in the abstract, is that Toyteller is significantly better at enabling toy-playing interaction than a GPT-4o-backed baseline while providing fluent real-time interaction.
Load-bearing premise
The system assumes that the 31 base action labels and the 924 training two-character motion instances cover the range of interactions users will express in open-ended story co-creation; events outside that vocabulary, such as a sports play, get mapped to whatever base action is nearest, and the generated text or motion drifts from user intent.
Editorial extensions
If this is right
- A user can steer story text by moving character symbols alone, expressing events that are hard to articulate in words.
- The system can fill in whichever part the user did not create, so initiative can be divided flexibly: the user moves one symbol while AI moves the other, or AI drafts both motion and text.
- Frame-by-frame motion generation stays under 0.1 seconds, so story playback with synchronized text and symbol motion is interactive in real time.
- Because toy-playing is perceived as vague and language as specific, combining a short text prompt with a motion input becomes a natural way to constrain the AI's interpretation.
- The proposed five-dimensional design space implies the same interaction principle can extend to more characters, props, 3D form factors, and different mappings between motion time and story time.
Reading between the lines
- Editorial inference: the shared action-embedding layer is a template for any gesture-to-text interaction, not just storytelling; any abstract motion that can be labeled with a small action vocabulary could be fused into an LLM via soft-prompt interpolation.
- Editorial inference: the paper reports that interpolating top-k actions gave little benefit over the top-1 action, which suggests the continuous embedding's nuance-preserving advantage is under-exploited by the current training data and loss; with softer action labels the advantage might materialize.
- Editorial inference: the observed bias toward common actions (e.g., against athletic stories) predicts that the approach's scalability depends on broadening the 31-action vocabulary; a testable extension is fine-tuning the action space on user-added, domain-specific actions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Toyteller is an AI-powered visual storytelling system in which users manipulate two triangular character symbols to express story events; the system maps these motions into a shared action-embedding space (based on SBERT and base action labels) and uses that space to condition LLM text generation (motion-to-text) and to generate character motions from text-derived actions (text-to-motion). The paper reports a technical evaluation of motion-to-action, motion-to-text, and action-to-motion components against GPT-4o-based baselines, a user study with 12 participants comparing Toyteller to a natural-language-only baseline, and a five-dimension design space for toy-playing interactions. The central claim is that Toyteller significantly outperforms GPT-4o at enabling toy-playing interaction and supports bidirectional motion/text steering.
Significance. If the central claim holds, Toyteller makes a valuable contribution to HCI and AI-powered storytelling: it demonstrates a genuinely new interaction modality (gestural toy-playing) for steering generative models, achieves real-time latencies with small custom models, and provides a design-space framing that could guide future systems. The paper has notable strengths: the technical evaluation uses a held-out split of an existing dataset with human ratings for text-motion alignment and motion realism; the user study is well-conducted (12 participants, think-aloud, CSI, qualitative coding); the authors are explicit about several limitations (e.g., the unvalidated text-to-action link and the system's action bias); and the design-space discussion is thoughtful. The main weakness is that the paper's headline 'outperforms GPT-4o' claim is only partially supported by the evaluations actually reported, and the text-steering direction central to the system's bidirectional design is never evaluated end-to-end.
major comments (3)
- [Section 5 (Technical Evaluation), Section 4.2.4, Section 3.2.2] The central claim of bidirectional motion-text steering is not directly supported because the full text-to-action-to-motion pipeline is never evaluated. Section 5 explicitly states that text2action+char is not evaluated, and Section 5.3 evaluates action-to-motion only with gold-standard action labels, not with actions inferred from story text. However, the user study (Section 6, Figure 15 and Figure 17b) shows that text-first creation is a common mode of use. The justification that SBERT and Llama-3 are separately strong does not cover the error propagation in the composite text-to-action-to-motion pathway. The authors should either add an evaluation of the full pipeline (or at least of text2action+char on held-out story sentences) or explicitly restrict the outperformance claim to the motion-steered direction and to gold-conditioned action-to-motion generation.
- [Section 5.2 (Figure 11) and Section 5.3 (Figure 12)] The abstract and Section 1 state that Toyteller is 'significantly better' than GPT-4o on motion-conditioned text generation and motion generation, but the reported statistics only partially support this. In Section 5.2, alignment with motion is significantly better than GPT-4o-C (p=0.00484) but not significantly different from GPT-4o-V; on coherence/grammaticality, GPT-4o-V is significantly better than Toyteller. In Section 5.3, the reactive motion generation condition shows a non-significant alignment difference (U=2120.0, p=0.321). These results should be reported more precisely in the abstract and introduction, distinguishing where Toyteller is and is not statistically superior.
- [Section 4.1 and Section 6.3.4] The generality of the 'outperforms GPT-4o' claim is limited by the narrow action space: the system is trained on 31 base action labels from 924 instances of the Roemmele et al. dataset. The user study itself notes that Toyteller 'seemed to have its own bias in interpreting actions, not well adapting to stories with specialized domains, such as athletic stories' (Section 6.3.4). The paper should prominently acknowledge that the technical outperformance is demonstrated only for this limited action space and that its applicability to open-ended story domains remains untested. A broader evaluation or a clear scope statement would strengthen the central claim.
minor comments (4)
- [Appendix B.7] The prompt contains a typo: 'The second character is named {second character's description}' should read 'named {second character's name}'.
- [Section 4.1] The dataset is referred to as 'Charades dataset,' but the cited references [69] and [71] describe 'Triangle Charades'; please use the dataset's actual name to avoid confusion with the unrelated video dataset of the same name.
- [Figure 15] The stacked-bar legend is dense and the color/label mapping is hard to parse in print; consider a clearer categorical color scheme or a supplementary table of per-participant interaction counts.
- [Section 5.2] The diversity metric (MST dispersion) is reported as a single number without any uncertainty estimate; at minimum, a note that this is a descriptive statistic would be appropriate.
Circularity Check
No significant circularity: the central claims rest on held-out data, an external GPT-4o baseline, and human ratings.
full rationale
Toyteller's central claims are supported by evaluations that do not reduce to its own fitted parameters. The motion2action model is trained on 924 Charades instances and scored on the held-out 232-instance test split (Section 4.1, Section 5.1), so the fact that both the training target and the Section 5.1 ranking metric live in the same SBERT embedding space is a standard supervised evaluation rather than a constructional equivalence. The motion-to-text and action-to-motion evaluations (Sections 5.2, 5.3) use human 7-point Likert ratings and compare directly against GPT-4o, an external baseline; no fitted parameter is renamed as a prediction. The one explicitly omitted component, text2action+char (Section 5), is not evaluated, and the paper justifies this with external benchmarks for SBERT and Llama-3 rather than with a self-citation chain; this is a coverage/limitation concern, not circularity. The user study similarly measures real user experience against a baseline tool using the same LLM. Self-citations appear (e.g., [13], [14], [15], [46], [82]) but are contextual and not load-bearing: no central premise depends solely on an unverified result imported from the authors' prior work, and no uniqueness theorem is invoked to force the design. The acknowledged domain-bias limitation (Section 6.3.4) is an empirical weakness of the dataset coverage, not a circular derivation.
Assumptions & free parameters
free parameters (4)
- top-k for action+char2text soft prompt =
4
- top-k for text2action+char =
2
- LLM generation temperature =
0.7 for action+char2text, 0.0 for text2action+char
- Action embedding weights for soft prompt =
cosine similarity ratios among top-k base actions
assumptions (4)
- domain assumption The 31 base action categories from Roemmele et al.'s Charades dataset are sufficient to express the two-character story interactions users will perform.
- domain assumption Cosine similarity in the all-MiniLM-L6-v2 SBERT embedding space reflects semantic proximity of action descriptions.
- ad hoc to paper A weighted sum of base action token embeddings in the Llama-3 input embedding space conditions the LLM to generate the intended action.
- domain assumption The Charades motion dataset, recorded via a data-collection game, generalizes to real interactive toy-playing with touchscreens.
Cite this review
Pith. "Pith review of Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols." pith.science (2026). https://pith.science/paper/RJA47E6K
@misc{pith2026250113284,
author = {Pith},
title = {Pith review of: Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols},
year = {2026},
howpublished = {\url{https://pith.science/paper/RJA47E6K}},
note = {Machine review of arXiv:2501.13284}
}
read the original abstract
We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
-
[1]
F Abell, F Happé, and U Frith. 2000. Do triangles play tricks? Attribution of mental states to animated shapes in normal and abnormal development.Cognitive Development 15, 1 (2000), 1–16. https://doi.org/10.1016/S0885-2014(00)00014-9
-
[3]
2024 (accessed August 1, 2024)
Ian Arawjo. 2024 (accessed August 1, 2024). LLM Wrapper Papers are Hurting HCI Research. https://ianarawjo.medium.com/llm-wrapper-papers-are-hurting- hci-research-8ad416a5d59a
2024
-
[4]
H. Clark Barrett, Peter M. Todd, Geoffrey F. Miller, and Philip W. Blythe. 2005. Accurate judgments of intention from motion cues alone: A cross-cultural study. Evolution and Human Behavior 26, 4 (2005), 313–331. https://doi.org/10.1016/j. evolhumbehav.2004.08.015
doi:10.1016/j 2005
-
[5]
Margaret S. Benson. 1993. The structure of four- and five-year- olds’ narratives in pretend play and storytelling. First Language 13, 38 (1993), 203–223. https://doi.org/10.1177/014272379301303803 arXiv:https://doi.org/10.1177/014272379301303803
-
[6]
Oloff C. Biermann, Ning F. Ma, and Dongwook Yoon. 2022. From Tool to Compan- ion: Storywriters Want AI Writers to Respect Their Personal Values and Writing Strategies. In Proceedings of the 2022 ACM Designing Interactive Systems Confer- ence (Virtual Event, Australia) (DIS ’22). Association for Computing Machinery, New York, NY, USA, 1209–1227. https://do...
arXiv 2022
-
[7]
Alex Calderwood, Vivian Qiu, Katy Ilonka Gero, and Lydia B Chilton. 2020. How novelists use generative language models: An exploratory user study.. In HAI-GEN+ user2agent@ IUI
2020
-
[8]
Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, and Smaranda Mure- san. 2024. Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional Writers. InProceedings of the 16th Conference on Cre- ativity & Cognition (Chicago, IL, USA) (C&C ’24). Association for Computing Ma- chinery, New York, NY, USA, 132–155. http...
arXiv 2024
-
[9]
Ankur Chemburkar, Andrew Gordon, and Andrew Feng. 2024. Evaluating Vision- Language Models on the TriangleCOPA Benchmark. The International FLAIRS Conference Proceedings 37, 1 (May 2024). https://journals.flvc.org/FLAIRS/article/ view/135485
2024
Show all 99 references
-
[10]
Erin Cherry and Celine Latulipe. 2014. Quantifying the Creativity Support of Digital Tools through the Creativity Support Index. ACM Trans. Comput.-Hum. Interact. 21, 4, Article 21 (jun 2014), 25 pages. https://doi.org/10.1145/2617588
2014 doi
-
[11]
Jean-Peïc Chou, Alexa Fay Siu, Nedim Lipka, Ryan Rossi, Franck Dernoncourt, and Maneesh Agrawala. 2023. TaleStream: Supporting Story Ideation with Trope Knowledge. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA...
2023
-
[12]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebas- tian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vi...
2022 arXiv
-
[13]
John Joon Young Chung, Minsuk Chang, and Eytan Adar. 2022. Gestural inputs as control interaction for generative human-AI co-creation. In Workshops at the International Conference on Intelligent User Interfaces (IUI)
2022
-
[14]
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching Stories with Generative Pretrained Language Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CH...
2022
-
[15]
John Joon Young Chung and Max Kreminski. 2024. Patchview: LLM-powered Worldbuilding with Generative Dust and Magnet Visualization. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburg, PA, USA)(UIST ’24). Association for Computin...
2024
-
[16]
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A. Smith. 2018. Creative Writing with a Machine in the Loop: Case Studies on Slo- gans and Stories. In Proceedings of the 23rd International Conference on Intelligent User Interfaces (Tokyo, Japan) (IUI ’18...
2018
-
[17]
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021. Directed Diversity: Leveraging Language Embedding Dis- tances for Collective Creativity in Crowd Ideation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Syste...
2021
-
[18]
Christopher Crick and Brian Scassellati. 2008. Inferring narrative and intention from playground games. In2008 7th IEEE International Conference on Development and Learning. 13–18. https://doi.org/10.1109/DEVLRN.2008.4640798
2008
-
[20]
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang
-
[21]
Tao Gao, Gregory McCarthy, and Brian J. Scholl. 2010. The Wolfpack Effect: Perception of Animacy Irresistibly Influences Interactive Behavior. Psychological Science 21, 12 (2010), 1845–1853. https://doi.org/10.1177/0956797610388814 arXiv:https://doi.org/10.1177/095679761038881...
2010 doi
-
[22]
Newman, and Brian J
Tao Gao, George E. Newman, and Brian J. Scholl. 2009. The psychophysics of chasing: A case study in the perception of animacy. Cognitive Psychology 59, 2 (2009), 154–179. https://doi.org/10.1016/j.cogpsych.2009.03.001
2009 doi
-
[23]
Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023. Social Dynamics of AI Support in Creative Writing. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Artic...
2023
-
[24]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind: One Embedding Space To Bind Them All. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 15180–15190
2023
-
[25]
Yuan Gong, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang, Xintao Wang, Ying Shan, and Yujiu Yang
-
[26]
Andrew Gordon. 2016. Commonsense Interpretation of Triangle Behavior. Proceedings of the AAAI Conference on Artificial Intelligence 30, 1 (Mar. 2016). https://doi.org/10.1609/aaai.v30i1.9881
2016 doi
-
[27]
In SIGGRAPH Asia 2023 Conference Papers (Sydney, NSW, Australia)(SA ’23)
Interactive Story Visualization with Multiple Characters. In SIGGRAPH Asia 2023 Conference Papers (Sydney, NSW, Australia)(SA ’23). Association for Computing Machinery, New York, NY, USA, Article 101, 10 pages. https://doi. org/10.1145/3610548.3618184
2023
-
[28]
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2022. Audioclip: Extending Clip to Image, Text and Audio. InICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 976–980. https: //doi.org/10.1109/ICASSP43922.2022.9747631
2022
-
[29]
Gordon and Melissa Roemmele
Andrew S. Gordon and Melissa Roemmele. 2014. An Authoring Tool for Movies in the Style of Heider and Simmel. In Interactive Storytelling, Alex Mitchell, Clara Fernández-Vara, and David Thue (Eds.). Springer International Publishing, Cham, 49–60
2014
-
[30]
Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. 2024. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey. arXiv:2403.14608 [cs.LG] https://arxiv.org/abs/2403.14608
2024 arXiv
-
[31]
David Ha and Douglas Eck. 2018. A Neural Representation of Sketch Drawings. In International Conference on Learning Representations . https://openreview.net/ forum?id=Hy6GHpkCW
2018
-
[32]
Hergenrader
T. Hergenrader. 2018.Collaborative Worldbuilding for Writers and Gamers. Blooms- bury Academic. https://books.google.co.kr/books?id=z-_7swEACAAJ
2018
-
[33]
Fritz Heider and Marianne Simmel. 1944. An Experimental Study of Apparent Behavior. The American Journal of Psychology 57, 2 (1944), 243–259. http: //www.jstor.org/stable/1416950
1944
-
[34]
Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-Term Memory. Neural Comput. 9, 8 (nov 1997), 1735–1780. https://doi.org/10.1162/neco.1997.9. 8.1735
1997 doi
-
[35]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022. Imagen Video: High Definition Video Generation with Diffusion Models. arXiv:2210.02303 [cs.CV] https://a...
2022 arXiv
-
[36]
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao. 2023. Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.. InICML (Proceedings of Machine Learning Research, Vol. 202), Andr...
2023
-
[37]
Carollee Howes. 1992. The collaborative construction of pretend: Social pretend play functions. State University of New York Press
1992
-
[38]
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. 2024. Mo- tiongpt: Human motion as a foreign language. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[39]
Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell. 2016. Visual Storytell...
2016
-
[40]
Hiroki Kaimoto, Kyzyl Monteiro, Mehrad Faridan, Jiatong Li, Samin Farajian, Yasuaki Kakehi, Ken Nakagaki, and Ryo Suzuki. 2022. Sketched Reality: Sketching Bi-Directional Interactions Between Virtual and Physical Worlds with AR and Actuated Tangible UI. In Proceedings of the 3...
2022
-
[41]
Ze Jin and Zorina Song. 2023. Generating coherent comic with rich story using ChatGPT and Stable Diffusion. arXiv:2305.11067 [cs.CV] https://arxiv.org/abs/ 2305.11067
2023 arXiv
-
[42]
Ross, Bryan Seybold, and Lu Jiang
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...
-
[43]
Taewook Kim, Hyomin Han, Eytan Adar, Matthew Kay, and John Joon Young Chung. 2024. Authors’ Values and Attitudes Towards AI-bridged Scalable Person- alization of Creative Language Arts. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, ...
2024
-
[44]
Max Kreminski, Melanie Dickinson, Noah Wardrip-Fruin, and Michael Mateas
-
[45]
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2023. AudioGen: Textually Guided Audio Generation. In The Eleventh International Conference on Learning Representations. https://openreview.net/...
2023
-
[46]
Max Kreminski and John Joon Young Chung. 2024. Intent Elicitation in Mixed- Initiative Co-Creativity.. In IUI Workshops
2024
-
[47]
Mina Lee, Percy Liang, and Qian Yang. 2022. CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for...
2022
-
[48]
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih ...
2021
-
[49]
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast Inference from Transformers via Speculative Decoding. In Proceedings of the 40th Interna- tional Conference on Machine Learning (Proceedings of Machine Learning Re- search, Vol. 202), Andreas Krause, Emma Brunskill, K...
2023
-
[50]
Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C
Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, S...
2024
-
[51]
When He Feels Cold, He Goes to the Seahorse
Di Liu, Hanqing Zhou, and Pengcheng An. 2024. "When He Feels Cold, He Goes to the Seahorse"—Blending Generative AI into Multimaterial Storymaking for Family Expressive Arts Therapy. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) ...
2024
-
[52]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744 [cs.CV] https://arxiv.org/abs/ 2310.03744
2024 arXiv
-
[53]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning
2023
-
[54]
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023. GPT Understands, Too. arXiv:2103.10385 [cs.CL] https: //arxiv.org/abs/2103.10385
2023 arXiv
-
[55]
Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Associa- tion for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: ...
2021 doi
-
[56]
Nicole Maslan, Melissa Roemmele, and Andrew S. Gordon. 2015. One Hundred Challenge Problems for Logical Formalizations of Commonsense Psychology. In Proceedings of the Twelfth International Symposium on Logi- cal Formalizations of Commonsense Reasoning (Commonsense-2015) . Sta...
2015
-
[57]
2024 (accessed July 22, 2024)
Meta. 2024 (accessed July 22, 2024). Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta-llama-3/
2024
-
[58]
Mathewson, Jaylen Pittman, and Richard Evans
Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany)...
2023
-
[59]
Neale and John M
Dennis C. Neale and John M. Carroll. 1997. Chapter 20 - The Role of Metaphors in User Interface Design. In Handbook of Human-Computer Interaction (Second Edition) (second edition ed.), Marting G. Helander, Thomas K. Landauer, and Prasad V. Prabhu (Eds.). North-Holland, Amsterd...
1997 doi
-
[60]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d 'Alché-Buc, E. Fo...
2019
-
[61]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Bal- com, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jef...
2025 arXiv
-
[62]
Bernstein
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology(San Franci...
2023
-
[63]
Esben Warming Pedersen and Kasper Hornbæk. 2011. Tangible bots: interaction with active tangibles in tabletop interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11) . Association for Computing Machinery, New Y...
2011
-
[64]
Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Aposto- los Dedeloudis, Jackson Sargent, and David Jurgens. 2022. POTATO: The Portable Text Annotation Tool. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Dem...
2022
-
[65]
2024 (accessed July 22, 2024)
OpenAI. 2024 (accessed July 22, 2024). Hello GPT-4o. https://openai.com/index/ hello-gpt-4o/
2024
-
[66]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...
2021
-
[67]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen
-
[68]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics. https://arxiv.org/abs/1908.10084
2019 arXiv
-
[69]
Melissa Roemmele, Haley Archer-McClellan, and Andrew S. Gordon. 2014. Trian- gle charades: a data-collection game for recognizing actions in motion trajectories. In Proceedings of the 19th International Conference on Intelligent User Interfaces (Haifa, Israel) (IUI ’14). Assoc...
2014
-
[70]
Hua Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, and Pan Hui. 2024. Char- acterMeet: Supporting Creative Writers’ Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars. In Pro- ceedings of the CHI Conference on Human Factors in Comput...
2024
-
[71]
Gordon, and Louis-Philippe Morency
Melissa Roemmele, Soja-Marie Morgens, Andrew S. Gordon, and Louis-Philippe Morency. 2016. Recognizing Human Actions in the Motion Trajectories of Shapes. In Proceedings of the 21st International Conference on Intelligent User Interfaces (Sonoma, California, USA) (IUI ’16). Ass...
2016
-
[72]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022. Photo- realistic Text-to-Image Diffusion Models with Deep Lan...
2022
-
[73]
arXiv:2204.06125 [cs.CV] https://arxiv.org/abs/2204.06125
Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125 [cs.CV] https://arxiv.org/abs/2204.06125
-
[74]
Kunwar Singh, Nicholas Davis, Chih-Pin Hsiao, Mikhail Jacob, Krunal Patel, and Brian Magerko. 2016. Recognizing Actions in Motion Trajectories Using Deep Neural Networks. Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment 12, 1 ...
2016
-
[75]
Nilüfer Talu. 2018. Symbolic creativity in play activity: a critique on playthings from daily life objects to toys. International Journal of Play 7, 1 (2018), 81–96
2018
-
[76]
Melissa Roemmele and Andrew S. Gordon. 2018. Automated Assistance for Creative Writing with an RNN Language Model. In Companion Proceedings of the 23rd International Conference on Intelligent User Interfaces (Tokyo, Japan) (IUI ’18 Companion). Association for Computing Machine...
2018
-
[77]
Hao Tan and Mohit Bansal. 2019. LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP...
2019 doi
-
[78]
2024 (accessed July 22, 2024)
Sentence Transformers. 2024 (accessed July 22, 2024). Sentence Transformers- Pretrained Models-Original Models. https://www.sbert.net/docs/sentence_ transformer/pretrained_models.html#original-models
2024
-
[79]
Helen Sharp, Yvonne Rogers, and Jenny Preece. 2007. Interaction Design: Beyond Human Computer Interaction. John Wiley & Sons, Inc., Hoboken, NJ, USA
2007
-
[80]
Nick Walton. 2019. AI Dungeon 2. https://aidungeon.cc/
2019
-
[81]
Qian Wan, Xin Feng, Yining Bei, Zhiqi Gao, and Zhicong Lu. 2024. Metamorpheus: Interactive, Affective, and Creative Dream Narration Through Metaphorical Visual Storytelling. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24...
2024
-
[82]
Felicia Fang-Yi Tan, Peisen Xu, Ashwin Ram, Wei Zhen Suen, Shengdong Zhao, Yun Huang, and Christophe Hurter. 2024. AudioXtend: Assisted Reality Visual Accompaniments for Audiobook Storytelling During Everyday Routine Tasks. In Proceedings of the CHI Conference on Human Factors...
2024
-
[83]
Williams and David Zipser
Ronald J. Williams and David Zipser. 1989. A Learning Algorithm for Continually Running Fully Recurrent Neural Networks. Neural Computation 1, 2 (1989), 270–280. https://doi.org/10.1162/neco.1989.1.2.270
1989 doi
-
[84]
Wobbrock and Julie A
Jacob O. Wobbrock and Julie A. Kientz. 2016. Research contributions in human- computer interaction. Interactions 23, 3 (April 2016), 38–44. https://doi.org/10. 1145/2907069
2016
-
[85]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. ...
2017
-
[86]
Kaige Xie and Mark Riedl. 2024. Creating Suspenseful Stories: Iterative Planning with Large Language Models. In Proceedings of the 18th Conference of the Euro- pean Chapter of the Association for Computational Linguistics (Volume 1: Long Pa- pers), Yvette Graham and Matthew Pu...
2024
-
[87]
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. In Proceedings of the 2021 Conference on Empirical Methods in Nat...
2021
-
[88]
Wang and Max Kreminski
Phoebe J. Wang and Max Kreminski. 2024. Guiding and Diversifying LLM- Based Story Generation via Answer Set Programming. arXiv:2406.00554 [cs.CL] https://arxiv.org/abs/2406.00554
2024 arXiv
-
[89]
Young, Takeo Igarashi, and Ehud Sharlin
James E. Young, Takeo Igarashi, and Ehud Sharlin. 2008. Puppet Master: designing reactive character behavior by demonstration. In Proceedings of the 2008 ACM SIGGRAPH/Eurographics Symposium on Computer Animation (Dublin, Ireland) (SCA ’08). Eurographics Association, Goslar, DE...
2008
-
[90]
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: Story Writing With Large Language Models. In 27th International Conference on Intelli- gent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Ma- chinery, New York, NY, USA, 841–852...
2022
-
[91]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[92]
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. 2022. PointCLIP: Point Cloud Understanding by CLIP. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8552–8562
2022
-
[93]
avoid,” “escape,
Yiming Zhang, Avi Schwarzschild, Nicholas Carlini, Zico Kolter, and Daphne Ippolito. 2024. Forcing Diffuse Distributions out of Language Models. arXiv:2404.10859 [cs.CL] https://arxiv.org/abs/2404.10859 A Technical Details A.1 Prompt: action+char2text My story has the followin...
2024 arXiv
-
[94]
Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, and Alex Wong. 2024. Binding Touch to Everything: Learning Unified Multimodal Tactile Representations. In Proceedings of the IEEE/CVF Conferenc...
2024
-
[97]
Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. 2024. Task Me Anything. arXiv preprint arXiv:2406.11775 (2024)
2024 arXiv
-
[100]
{Series of images appended} B.2 GPT-4o-C Prompt for vs
examine - 0.2 ... {Series of images appended} B.2 GPT-4o-C Prompt for vs. motion2action The following sequence of coordinates is describing an action described in character symbols. {sequence of coordinates in ((x1, y1, r1), (x2, y2, r2))} For each frame, the first item is for...
-
[101]
B.3 GPT-4o-V Prompt for vs
examine - 0.2 ... B.3 GPT-4o-V Prompt for vs. motion2char The following sequence of image is describing an action described in character symbols denoted in white and black triangles. The described action is{action}. The frame rate of image sequence is 10 frames per second. Dec...
2025
-
[2022]
In Proceedings of the Eighteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (Pomona, CA, USA) (AIIDE’22)
Loose ends: a mixed-initiative creative interface for playful storytelling. In Proceedings of the Eighteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (Pomona, CA, USA) (AIIDE’22). AAAI Press, Article 15, 9 pages. https://doi.org/10.1609/...
-
[2023]
In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
CLAP Learning Audio Concepts from Natural Language Supervision. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1–5. https://doi.org/10.1109/ICASSP49357.2023.10095889
2023
-
[2024]
arXiv:2312.14125 [cs.CV] https://arxiv.org/abs/2312.14125
VideoPoet: A Large Language Model for Zero-Shot Video Generation. arXiv:2312.14125 [cs.CV] https://arxiv.org/abs/2312.14125
-
[3059]
https://doi.org/10.18653/v1/2021.emnlp-main.243
2021 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.