REVIEW 3 major objections 5 minor 89 references
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Object-based audio editing, where characters and scenes are editable objects linked to their audio assets, lets a single high-level edit replace dozens of manual timeline edits and cuts professional creation time from 16.7 to 3.7 hours.
desk verdict A genuinely useful systems paper: object-based audio editing is a real abstraction with good evidence, but the headline 'automatic propagation' claim rests on an LLM link that the paper never directly measures. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the tripartite object model: objects (narrative constructs such as characters and scenes), attributes (natural-language designs such as a voice design like 'a middle-aged woman with a measured cadence' or an ambience design like 'a busy market square'), and object-dependent assets (each speech clip, sound effect, music cue, and ambience layer on the timeline, each linked to the object or objects it belongs to). The argument runs on propagation: editing an attribute triggers an LLM to identify all linked assets requiring updates and to propose concrete modifications, so one high-level edit replaces many low-level clip edits. Supporting machinery includes best-of-25 script generation scored by an LLM judge against an 11-item rubric, batched text-to-speech synthesis of all of a character's lines in a single call to ensure voice consistency, and anchoring music and ambience to word-level dialogue timestamps so that assets stay aligned with dialogue across regenerations.
What would settle it
Run Dramarrator on several books, apply one object edit such as making a specified character's voice raspier, and count how many linked speech, sound effect, music, and ambience clips actually change; if the changed set is smaller than the linked set, propagation is incomplete. Separately, generate a character's lines batched in one text-to-speech call versus line-by-line and have listeners judge whether both sets sound like the same character; if batched and unbatched lines are equally consistent, the batching rationale for the propagation model is not load-bearing.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that narrative constructs in audio drama—characters and scenes—can be reified as first-class editable objects whose attributes generate and regenerate their dependent assets, and that this structure converts high-level creative decisions from expensive project-wide manual overhauls into single parameter changes. Dramarrator extracts these objects from raw book text with an LLM, instantiates a voice design per character and an ambience design per scene, generates 25 candidate adapted scripts and selects the best by an 11-item rubric scored by an LLM judge, then synthesizes speech, sound effects, music, and ambience layers into a four-track timeline where every clip carries a link to its object or objects. When a creator edits an attribute, such as making a character nervous, an LLM proposes which dependent assets need changes, and accepted proposals regenerate or rewrite them across the whole project. The paper reports that this yields a significant task-load reduction for professionals, a 9.7x average amplification of editing effort (one voice edit propagated to all 30 of a character's lines in one case), and listener ratings for refined output that are statistically comparable to professional productions on engagement, overall quality, and character/scene consistency, while a gap remains on audio element placement, timing, clarity, and distraction.
Load-bearing premise
The whole efficiency gain depends on the assumption that one attribute value—a single AI voice prompt for a character, one ambience prompt for a scene—can faithfully represent that character or scene everywhere it appears, so regenerating all linked assets from that attribute preserves consistency; the paper reports no accuracy measurement of this extraction or of voice consistency on the raw pipeline.
Editorial extensions
If this is right
- If correct, object-based editing turns late-stage creative revisions in audio drama from near-restarts into cheap experiments, since changing a character or scene attribute regenerates every dependent asset automatically.
- A single production can support many alternative versions—different voices for the same character, different ambience for the same scene—because objects parameterize the whole composition rather than individual clips.
- The object model makes consistency a structural property: because all of a character's lines are generated from one voice attribute and batched in one text-to-speech call, characters and scenes stay recognizably the same across scenes, which the listener study suggests is achieved.
- The paradigm implies a workflow shift for professionals from sequential pre-production, production, and post-production to iterative listen-edit-listen cycles, which participants in the study reported as increased creative exploration.
- Object-based editing could extend to other narrative-driven media where characters and scenes cascade across heterogeneous assets, such as game audio and immersive stories, an extension the paper's exploratory study begins to probe.
Reading between the lines
- The object model could generalize beyond audio to video production, where character and scene objects would propagate lighting, framing, and expression changes—though visual regeneration risks cross-frame consistency in a way audio regeneration does not; the paper gestures at this possibility but does not test it.
- A shared emotion space such as valence and arousal could make cross-modal propagation more transparent and less reliant on LLM judgment: after an edit shifts a scene's emotional valence, assets out of alignment could be flagged deterministically, an extension the paper discusses but does not implement.
- Object inheritance (such as child and adult versions of the same character) and time-keyframed attributes would let the object model handle characters and environments that change over a story, which the current always-one-attribute model cannot express.
- The measured fourfold time reduction likely understates the gain for full-length books, because the 16.7-hour baseline came from short excerpts; amortizing object creation and prompt tuning over a whole production could make the ratio larger—this is speculative and should be tested.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Dramarrator, an end-to-end authoring system that adapts books into multi-track audio dramas using object-based audio editing. Characters and scenes are extracted by an LLM into editable objects with attributes such as voice and ambience design; generated speech, SFX, music, and ambience are linked to these objects, and edits to an object are claimed to propagate automatically to all dependent assets. The system is evaluated in three studies: a within-subjects user study with 8 professional producers comparing Dramarrator to their existing tools, a listener study with 300 participants comparing automatically generated, professionally refined, and existing-tool productions, and an exploratory study with 3 novice storytellers. The reported results are that Dramarrator significantly lowers task load, that a single object edit replaces up to 30 manual edits (a 9.7x amplification), that self-reported creation time falls from 16.7 to 3.7 hours, and that refined output approaches the quality of existing-tool productions on several dimensions.
Significance. If the automatic-propagation mechanism works as described, the contribution is significant for creativity support tools: it offers a genuinely semantic editing abstraction over heterogeneous audio assets, grounded in a careful formative analysis of professional practice. The paper has real strengths: the professional user study collects both quantitative ratings and rich qualitative evidence, the listener study is large (N=300), and the authors are transparent about several study limitations. The central weakness is that the load-bearing mechanism—object extraction, asset-to-object linking, and LLM-proposed edit propagation—is never evaluated for correctness or completeness. The headline efficiency claims therefore rest on an unmeasured mechanism, and the causal attribution to object-based editing is weakened by the fixed condition order and by the absence of a condition that isolates object-based editing from the fully automated pipeline. The paper is a promising systems contribution, but the evidence as presented does not yet make the central claim secure.
major comments (3)
- [Section 4.1 / Algorithm 1] The claim that edits 'automatically propagate to all dependent assets' rests on an LLM identifying which assets require updates and proposing modifications, but the paper reports no accuracy evaluation of (a) character and scene extraction from raw book text (Algorithm 1, line 2), (b) asset-to-object linking (Algorithm 1, line 7), or (c) the completeness and correctness of the LLM's propagation suggestions (Section 4.1). Consequently, the quantitative supports in Section 5.2—2.6 object edits 'propagating' into 25.0 manual edits, the 9.7x amplification, and the 16.7-to-3.7-hour time reduction—are not yet tied to a measured mechanism. If the LLM misses dependent assets, the claim is unsupported; if it proposes spurious edits, the amplification overstates the benefit. An end-to-end propagation accuracy test, measuring missed and spurious edits against a ground-truth object model on held-out books, is needed to make the central claim secure.
- [Section 5.1 / 5.1.3] The study cannot cleanly attribute the observed task-load reduction to object-based editing. Participants always completed the existing-tools condition first (Section 5.1), time estimates were self-reported (Section 5.1.3), and Dramarrator bundles object-based editing together with a fully automated script-generation, asset-generation, and mixing pipeline. The limitations paragraph acknowledges the ordering and self-report issues but not the attribution confound: there is no condition with object-based editing disabled, and no condition in which the automated pipeline is used without the object abstraction. A follow-up ablation or counterbalanced design is needed to support the causal claim; until then, the abstract's and Section 5.2's causal wording should be softened.
- [Section 4.3 / 5.2] The pipeline claims to ensure voice consistency by concatenating all of a character's dialogue lines into a single TTS call (Section 4.3), but no objective or perceptual evaluation of cross-scene voice consistency is reported, and participant P3 explicitly doubted that a single AI voice prompt produces consistent results across all lines (Section 5.2). Since consistency is a promised benefit of object propagation (DG2) and a central part of the listener-evaluation argument, the paper should report a consistency evaluation—for example, per-character listener ratings across different scenes, or acoustic similarity measures—along with the accuracy of character and scene extraction from the raw book text.
minor comments (5)
- [Section 5.2] The '9.7x amplification' is presented as if it were directly measured: '2.6 object edits that propagated into 25.0 manual edits.' As described, this appears to be a projection based on counting object-level edits and estimating the manual edits they would replace. Please label the metric as an estimate and explain how the counterfactual manual-edit count was derived.
- [Table 5 / Section 5.2] The table reports raw p-values for many Wilcoxon tests without any multiple-comparison correction. Reporting uncorrected values is acceptable, but it should be stated explicitly so readers can calibrate the number of significant findings.
- [Section 5.1.2 / Figure 5] The procedure says participants had 'unlimited time to refine' within five business days, while Figure 5 reports 2.9 hours of final mixing and mastering in existing tools. Please clarify how the 3.7-hour Dramarrator total was measured and whether it includes that 2.9-hour post-refinement mixing phase.
- [Section 5.3] The listener study randomly chose 6 stories from the 8-participant user study and assigned each listener one condition per story. The assignment of conditions to stories and listeners is not fully described; please specify the design (e.g., Latin square, number of ratings per condition) so the reader can assess balance.
- [Figure 2] The figure caption and labels appear to contain garbled text tokens (e.g., '/gid00048/gid00065/...') in the provided manuscript. If these are rendering artifacts, the published version should use clean, readable labels.
Circularity Check
No significant circularity: the paper's claims are supported by empirical user and listener studies, not by a derivation that reduces to its inputs.
full rationale
The paper is a systems-and-evaluation paper, not a formal derivation. Its central claim that object-based editing automatically propagates edits to dependent assets is implemented as a pipeline (Algorithm 1) and assessed through a within-subjects user study (N=8), a listener study (N=300), and an exploratory study (N=3). The reported time savings and the 9.7x editing amplification are empirical measurements from interaction logs and self-reports, not quantities fitted to reproduce the claim. The LLM-judge script selection uses a rubric derived from the same formative best practices that motivated the system, but this is an internal quality-control step, not a prediction compared against the paper's headline results. The one overlapping-author citation, SoundStager [88], is used only for adapting a soundscape taxonomy for ambience decomposition and is not load-bearing for the central propagation claim. The paper even discloses study limitations (Section 5.1.3) and reports P3's skepticism about voice consistency, which further indicates the results are presented as empirical findings rather than as inevitable consequences of definitions. Any concern that propagation accuracy is not separately benchmarked is a validity or completeness issue, not a circularity issue.
Assumptions & free parameters
free parameters (4)
- script_candidate_count_N =
25
- per_track_mix_gain =
speech 0 dB, SFX -6 dB, music -14 dB, ambience -20 dB; -16 LUFS normalize
- scene_buffer_duration =
5 s
- llm_judge_repetitions =
3
assumptions (4)
- domain assumption Gemini 3 Pro can reliably extract characters and scenes from raw book text and infer plausible voice and ambience attributes.
- domain assumption Batching all of a character's lines into one TTS call and splicing with word-level alignment maintains consistent voice across scenes.
- domain assumption An LLM-as-a-judge with an 11-item rubric yields script quality scores that correlate with human listener preference.
- domain assumption The audio drama mixing practices in Table 1 (e.g., dialogue at reference level, named LUFS targets) are valid and apply to AI-generated assets.
invented entities (1)
-
Character and Scene objects (object-based audio editing model)
Cite this review
Pith. "Pith review of Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books." pith.science (2026). https://pith.science/paper/ABWF7OH5
@misc{pith2026260808349,
author = {Pith},
title = {Pith review of: Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books},
year = {2026},
howpublished = {\url{https://pith.science/paper/ABWF7OH5}},
note = {Machine review of arXiv:2608.08349}
}
read the original abstract
Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains labor-intensive, requiring them to interpret source material, author scripts, generate audio assets, and assemble them on a timeline. Because story elements like characters and scenes manifest across many interdependent assets, a single change can ripple into manual updates across the entire project. We present Dramarrator, an audio drama authoring tool built around object-based audio editing, where these story elements are represented as editable objects. Dramarrator extracts these objects from a book, generates linked audio assets (speech, sound effects, and music), and composes a multi-track audio drama. Edits to any object (e.g., a character's voice) automatically propagate to all dependent assets. In a user study with professionals (N=8), Dramarrator significantly lowered task load when creating audio dramas. A listener study (N=300) shows that creator-refined output from Dramarrator approaches the quality of productions made with existing professional tools, and an exploratory study (N=3) suggests object-based editing lowers entry barriers and generalizes beyond audio dramas.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
ACX. 2026. ACX audio submission requirements. https://help.acx.com/s/article/ what-are-the-acx-audio-submission-requirements Accessed: 2026
2026
-
[2]
Adobe. 2026. Adobe Audition. https://www.adobe.com/ca/products/audition. html Accessed: 2026
2026
-
[3]
Adobe. 2026. After Effects. https://www.adobe.com/products/aftereffects.html Accessed: 2026
2026
-
[4]
Adobe. 2026. Firefly. https://www.adobe.com/products/firefly Accessed: 2026
2026
-
[5]
Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, An- toine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. 2023. MusicLM: Generating Music From Text. https://doi.org/10.48550/arXiv.2301.11325
-
[6]
Apple. 2026. Logic Pro. https://www.apple.com/logic-pro/ Accessed: 2026
2026
-
[7]
Avid. 2026. Pro Tools. https://www.avid.com/pro-tools Accessed: 2026
2026
-
[8]
YouTube: BBC. 2026. Kenneth Branagh and David Tennant on the art of Radio Drama (BBC Radio 4). https://www.youtube.com/watch?v=4IagWCqZ2ww Accessed: 2026
2026
Show all 89 references
-
[9]
Michel Beaudouin-Lafon. 2000. Instrumental interaction: an interaction model for designing post-WIMP user interfaces. InProceedings of the SIGCHI conference on Human factors in computing systems. 446–453
2000
-
[11]
2021.Audionarratology: lessons from radio drama
Lars Bernaerts and Jarmila Mildorf. 2021.Audionarratology: lessons from radio drama. Ohio State University Press
2021
-
[12]
Blender. 2026. Blender. https://www.blender.org/ Accessed: 2026
2026
-
[13]
Toon Boom. 2026. Harmony. https://www.toonboom.com/products/harmony Accessed: 2026
2026
- [14]
-
[15]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology3, 2 (Jan. 2006), 77–101. https://doi.org/10. 1191/1478088706qp063oa
2006
-
[16]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health11, 4 (Aug. 2019), 589–597. https://doi.org/10.1080/2159676X.2019.1628806
2019
-
[17]
J Brooke. 1996. SUS: A quick and dirty usability scale.Usability Evaluation in Industry(1996)
1996
-
[18]
Rick Busselle and Helena Bilandzic. 2009. Measuring Narrative Engage- ment.Media Psychology12, 4 (Nov. 2009), 321–347. https://doi.org/10.1080/ 15213260903287259
2009
-
[19]
positive fric- tion
Zeya Chen and Ruth Schmidt. 2024. Exploring a behavioral model of “positive fric- tion” in human-AI interaction. InInternational Conference on Human-Computer Interaction. Springer, 3–22
2024
-
[20]
Erin Cherry and Celine Latulipe. 2014. Quantifying the Creativity Support of Digital Tools through the Creativity Support Index.ACM Trans. Comput.-Hum. Interact.21, 4 (June 2014), 21:1–21:25. https://doi.org/10.1145/2617588
2014 doi
-
[21]
YouTube: Writing Comics. 2026. How to Write an Audio Drama and the Collabo- rations That Come With It. https://www.youtube.com/watch?v=z6OWH7r8bAU Accessed: 2026
2026
-
[22]
Panos Constantinides, Ola Henfridsson, and Geoffrey G Parker. 2018. Introduc- tion—platforms and infrastructures in the digital age. , 381–400 pages
2018
-
[23]
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023. Simple and controllable music generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23). Curran Associ...
2023
-
[24]
2014.Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory
Juliet Corbin and Anselm Strauss. 2014.Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory. SAGE Publications
2014
-
[25]
Anna L Cox, Sandy JJ Gould, Marta E Cecchinato, Ioanna Iacovides, and Ian Ren- free. 2016. Design frictions for mindful interactions: The case for microboundaries. InProceedings of the 2016 CHI conference extended abstracts on human factors in computing systems. 1389–1397
2016
-
[26]
2002.Radio drama
Tim Crook. 2002.Radio drama. Routledge
2002
-
[27]
2024.Human+ machine, updated and expanded: reimagining work in the age of AI
Paul R Daugherty and H James Wilson. 2024.Human+ machine, updated and expanded: reimagining work in the age of AI. Harvard Business Press
2024
-
[28]
2005.Writing and producing radio dramas
Esta De Fossard. 2005.Writing and producing radio dramas. Vol. 1. Sage
2005
-
[29]
Descript. 2026. Descript. https://www.descript.com Accessed: 2026
2026
-
[30]
ElevenLabs. 2026. Eleven v3. https://elevenlabs.io/blog/eleven-v3 Accessed: 2026
2026
-
[31]
ElevenLabs. 2026. ElevenLabs. https://elevenlabs.io Accessed: 2026
2026
- [32]
-
[33]
Jianyu Fan, Miles Thorogood, and Philippe Pasquier. 2017. Emo-soundscapes: A dataset for soundscape emotion recognition. In2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII). 196–201. https://doi.org/10.1109/ACII.2017.8273600
2017
-
[34]
Figma. 2026. Figma. https://www.figma.com/ Accessed: 2026
2026
-
[35]
Google. 2026. Gemini 3 Pro. https://gemini.google.com/ Accessed: 2026
2026
-
[36]
Google. 2026. NotebookLM. https://notebooklm.google.com Accessed: 2026
2026
-
[37]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, and others. 2024. A survey on llm-as-a-judge.The Innovation(2024)
2024
-
[38]
Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma, Yixiao Ge, Wei Zou, and Ying Shan
-
[39]
Han, Junhang Yu, Raphael Bournet, Alexandre Ciorascu, Wendy E
Han L. Han, Junhang Yu, Raphael Bournet, Alexandre Ciorascu, Wendy E. Mackay, and Michel Beaudouin-Lafon. 2022. Passages: Interacting with Text Across Documents. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing...
2022
-
[40]
Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. InAdvances in psy- chology. Vol. 52. Elsevier, 139–183
1988
-
[41]
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024. Llm-rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguisti...
2024
-
[42]
Le Van Huy, Hien TT Nguyen, Tan Vo-Thanh, Nguyen Huu Thai Thinh, Tran Thi Thu Dung, and others. 2024. Generative AI, why, how, and outcomes: A user adoption study.AIS Transactions on Human-Computer Interaction16, 1 (2024), 1–27
2024
-
[43]
Harry H Jiang, William Agnew, Tim Friedlander, Zhuolin Yang, Sarah E Fox, Michael S Bernstein, Josephine Charlie Passananti, Megumi Ogata, and Karla Ortiz. 2025. Forging an HCI Research Agenda with Artists Impacted by Generative AI. InProceedings of the Extended Abstracts of t...
2025
-
[44]
YouTube: Booth Junkie. 2026. Talking Audio Drama Sound Design With Kenny Neal | Booth Junkie. https://www.youtube.com/watch?v=NAO2byD9n1s Ac- cessed: 2026
2026
-
[45]
Patrik N Juslin and Klaus R Scherer. 2005. Vocal expression of affect. InThe New Handbook of Methods in Nonverbal Behavior Research, Jinni A Harrigan, Robert Rosenthal, and Klaus R Scherer (Eds.). Oxford University Press, 0. https: //doi.org/10.1093/oso/9780198529613.003.0003
2005
- [47]
-
[48]
Philippe Laban, Elicia Ye, Srujay Korlakunta, John Canny, and Marti Hearst
-
[49]
2013.Theatre studies: The basics
Robert Leach. 2013.Theatre studies: The basics. Routledge
2013
-
[50]
Jingyi Li, Eric Rawn, Jacob Ritchie, Jasper Tran O’Leary, and Sean Follmer. 2023. Beyond the Artifact: Power as a Lens for Creativity Support Tools. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing M...
2023
-
[51]
Plumbley, and Wenwu Wang
Xubo Liu, Zhongkai Zhu, Haohe Liu, Yi Yuan, Qiushi Huang, Meng Cui, Jinhua Liang, Yin Cao, Qiuqiang Kong, Mark D. Plumbley, and Wenwu Wang. 2025. WavJourney: Compositional Audio Creation With Large Language Models.IEEE Transactions on Audio, Speech and Language Processing33 (2...
2025
-
[52]
Ryan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry, and Carrie J Cai
-
[53]
Yi Luo. 2025. Designing with AI: A systematic literature review on the use, development, and perception of AI-enabled UX design tools.Advances in Human- Computer Interaction2025, 1 (2025), 3869207
2025
-
[54]
Philipp Mayring. 2021. Qualitative content analysis: A step-by-step guide.Quali- tative Content Analysis(2021), 1–100
2021
-
[55]
YouTube: Ryan McPherson and San Antonio Storytellers. 2026. How to Start Your Audio Drama Podcast With Producer Brooke Pillifant. https://www.youtube. com/watch?v=-MO97WJL8GY&t=41s Accessed: 2026
2026
-
[56]
YouTube: Cactus Coast Media. 2026. HOW TO CREATE AN AUDIO DRAMA | FICTIONAL PODCAST. https://www.youtube.com/watch?v=38nzdpONO1A Accessed: 2026
2026
-
[57]
Marc Pinski and Alexander Benlian. 2024. AI literacy for users–A comprehensive review and future research directions of learning methods, components, and effects.Computers in Human Behavior: Artificial Humans2, 1 (2024), 100062
2024
-
[58]
Andrew Schwartz, Gregory Park, Johannes Eich- staedt, Margaret Kern, Lyle Ungar, and Elisabeth Shulman
Daniel Preoţiuc-Pietro, H. Andrew Schwartz, Gregory Park, Johannes Eich- staedt, Margaret Kern, Lyle Ungar, and Elisabeth Shulman. 2016. Modelling Valence and Arousal in Facebook posts. InProceedings of the 7th Workshop on Computational Approaches to Subjectivity, Sentiment an...
2016 doi
-
[59]
YouTube: Fool & Scholar Productions. 2026. Writing for Audio Fiction - Au- dio Drama Podcasting with K. A. Statz. https://www.youtube.com/watch?v= HEdrzLPziR4 Accessed: 2026
2026
-
[60]
Proferes
N.T. Proferes. 2005.Film Directing Fundamentals: See Your Film Before Shooting. Focal Press. https://books.google.com/books?id=wp39nV-axRUC
2005
-
[61]
Prolific. 2026. Prolific. https://www.prolific.com Accessed: 2026
2026
-
[62]
2002.Theatre of Sound: Radio and the Dramatic Imagination
Dermot Rattigan. 2002.Theatre of Sound: Radio and the Dramatic Imagination. Carysfort Press
2002
-
[63]
Tim Rentsch. 1982. Object oriented programming.ACM Sigplan Notices17, 9 (1982), 51–57
1982
-
[64]
Nathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown, Hugo Romat, Michel Pahud, Nicolai Marquardt, Kori Inkpen, and Ken Hinckley. 2025. AI- Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools. InP...
2025
-
[65]
Steve Rubin and Maneesh Agrawala. 2014. Generating emotionally relevant musi- cal scores for audio stories. InProceedings of the 27th annual ACM symposium on User interface software and technology (UIST ’14). Association for Computing Ma- chinery, New York, NY, USA, 439–448. h...
2014
-
[66]
Mysore, Wilmot Li, and Maneesh Agrawala
Steve Rubin, Floraine Berthouzoz, Gautham J. Mysore, Wilmot Li, and Maneesh Agrawala. 2013. Content-based tools for editing audio stories. InProceedings of the 26th annual ACM symposium on User interface software and technology (UIST ’13). Association for Computing Machinery, ...
2013
-
[67]
James A Russell. 1980. A circumplex model of affect.Journal of personality and social psychology39, 6 (1980), 1161
1980
-
[68]
Murray Schafer
R. Murray Schafer. 1994.The Soundscape: Our Sonic Environment and the Tuning of the World. Destiny Books, Rochester
1994
-
[69]
Hijung Valentina Shin, Wilmot Li, and Frédo Durand. 2016. Dynamic Authoring of Audio with Linked Scripts. InProceedings of the 29th Annual Symposium on User Interface Software and Technology (UIST ’16). Association for Computing Ma- chinery, New York, NY, USA, 509–516. https:/...
2016
-
[70]
Epidemic Sound. 2026. Sound Effects. https://www.epidemicsound.com/sound- effects/ Accessed: 2026
2026
-
[71]
Mark Stefik and Daniel G Bobrow. 1985. Object-oriented programming: Themes and variations.AI magazine6, 4 (1985), 40–40
1985
-
[72]
Dow, and Tovi Grossman
Sangho Suh, Michael Lai, Kevin Pu, Steven P. Dow, and Tovi Grossman. 2025. StoryEnsemble: Enabling Dynamic Exploration & Iteration in the Design Process with AI and Forward-Backward Propagation. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Tech...
2025
-
[73]
Suno. 2026. AI Music Generator. https://suno.com/ Accessed: 2026
2026
-
[74]
YouTube: Radius: the religious drama society. 2026. How to write a radio drama. https://www.youtube.com/watch?v=C7XGGY6oxPs Accessed: 2026
2026
-
[75]
R. Toscan. 2023.Writing Audio Drama: Making Scripts that Work for Fiction & True Crime Podcasts. Independently published. https://a.co/d/03qMP7O2
2023
-
[76]
Tiffany Tseng, Ruijia Cheng, and Jeffrey Nichols. 2024. Keyframer: Empowering animation design using large language models.arXiv preprint arXiv:2402.06071 (2024)
2024 arXiv
-
[77]
Upwork. 2026. Upwork. https://www.upwork.com/hire/ Accessed: 2026
2026
-
[78]
Wayland, W
K. Wayland, W. Lucas, and S. Ryle. 2020.Bombs Always Beep - 2nd Edition - Revenge of the Beep: Creating Modern Audio Theater. Amazon Digital Services LLC - KDP Print US. https://books.google.com/books?id=R_M5zgEACAAJ
2020
-
[79]
Michael Wessel, Martin Adam, Alexander Benlian, Ann Majchrzak, and Ferdinand Thies. 2025. Generative AI and its transformative value for digital platforms. Journal of Management Information Systems42, 2 (2025), 346–369
2025
-
[80]
Robert F Woolson. 2007. Wilcoxon signed-rank test.Wiley encyclopedia of clinical trials(2007), 1–3
2007
-
[81]
YouTube: BBC Writers. 2026. Scriptwriting advice from the BBC Radio Drama North team. https://www.youtube.com/watch?v=h6sI3YpkO2Q Accessed: 2026
2026
-
[82]
YouTube: BBC Writers. 2026. Writing for Radio - advice from the team in BBC Ra- dio Drama North. https://www.youtube.com/watch?v=9uE3IHn5IuE Accessed: 2026
2026
-
[83]
Haijun Xia. 2020. Crosspower: Bridging Graphics and Linguistics. InProceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (UIST ’20). Association for Computing Machinery, New York, NY, USA, 722–734. https://doi.org/10.1145/3379337.3415845
2020
-
[84]
Haijun Xia, Bruno Araujo, Tovi Grossman, and Daniel Wigdor. 2016. Object- Oriented Drawing. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ’16). Association for Computing Machinery, New York, NY, USA, 4610–4621. https://doi.org/10.1145/2858...
2016
-
[85]
Haijun Xia, Bruno Araujo, and Daniel Wigdor. 2017. Collection Objects: En- abling Fluid Formation and Manipulation of Aggregate Selections. InProceed- ings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI Dramarrator: Object-Based Audio Editing for Audio D...
2017
- [86]
-
[87]
Tom Yeh and Jeeeun Kim. 2018. CraftML: 3D Modeling is Web Programming. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3173574.3174101
2018
-
[88]
Does this scene advance the central conflict?
Suhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon, Oriol Nieto, and Anh Truong. 2026. SoundStager: Interactive Design of Story- Driven GenAI Soundscapes for Video. (2026). UIST ’26, November 02–05, 2026, Detroit, MI, USA Benharrak, et al. A APPENDIX T...
2026
-
[2020]
InProceedings of the 2020 CHI conference on human factors in computing systems
Novice-AI music co-creation via AI-steering tools for deep generative models. InProceedings of the 2020 CHI conference on human factors in computing systems. 1–13
2020
-
[2022]
InProceedings of the 27th International Conference on Intelligent User Interfaces(Helsinki, Finland) (IUI ’22)
NewsPod: Automatic and Interactive News Podcasts. InProceedings of the 27th International Conference on Intelligent User Interfaces(Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 691–706. https://doi.org/10.1145/3490099.3511147
-
[2025]
AudioStory: Generating Long-Form Narrative Audio with Large Language UIST ’26, November 02–05, 2026, Detroit, MI, USA Benharrak, et al. Models. https://doi.org/10.48550/arXiv.2508.20088
2026 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.