{"id":"974ef364-ff0b-49b2-b32b-676a0b7f84bf","arxiv_id":"2608.10337","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Narrative keyframing uses plot, character, and perspective keyframes to let writers control AI-generated stories, with a 12-writer study reporting higher perceived control and richer characterization.","lead":"This paper introduces narrative keyframing, an interaction technique where writers set plot, character, and first-person perspective constraints at key story moments and let an AI fill in the rest. It reports that writers find the system more controllable, transparent, and engaging than a chatbot-style baseline, and that generated stories are rated as having richer characterization.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Study 1 compares an information-rich keyframing pipeline against an outline-only baseline, so the output-quality gain may come from extra conditioning rather than from narrative keyframing itself.","rationale":"The paper has two strands: an interaction design contribution and an output-quality claim. The interaction claims are supported by a reasonably well-conducted 12-participant within-subjects study using standardized instruments and qualitative analysis, and the authors explicitly flag the sample as preliminary; I credit that evidence. The load-bearing weakness is the output-quality claim, which appears in the abstract and conclusion and is the quantitative hook of the paper. It rests on Study 1, where the treatment differs from the baseline by more than keyframing: it also receives traits, first-person perspectives, and selected evidence. Those extra inputs are exactly the kind of content that could improve automatic quality scores and human judgments independently of the keyframing representation. The paper's own framing of Study 1 as validating 'narratologically-motivated conditioning' is narrower than the conclusion's claim about the approach as a whole. This is a fixable experimental-design gap rather than a fatal flaw, and the user-study evidence for the interaction experience still stands, so the reader's CONDITIONAL verdict is appropriate and no change is needed.","tokens_in":25297,"tokens_out":5081,"duration_ms":47649,"concrete_test":"Re-run Study 1 with an information-matched baseline: for each of the 20 prompts, give the vanilla GPT-4.1 baseline the same outline plus the same auto-suggested traits, first-person perspectives, and selected evidence as a flat prompt, without keyframing structure or evidence-selection UI; then recompute the WQRM-PRE pairwise preference and the human-rater comparison on 30 pairs. If the advantage disappears or shrinks to non-significance, the headline quality claim is attributable to extra conditioning rather than to narrative keyframing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central output-quality claim rests on Study 1, but the comparison in Section 6.1.1 is confounded: the keyframing condition receives the outline plus auto-suggested traits, generated first-person perspectives, and randomly selected evidence per character per plot, while the baseline receives only the outline and is asked to expand it. Any of those extra inputs, not the keyframing representation, could explain the higher WQRM-PRE scores and human preference rates reported in Section 6.1.2. Moreover, the pipeline generates traits for every plot rather than testing sparse keyframes plus interpolation, so the defining keyframing property is not actually exercised in Study 1. The paper's own Section 6.1 framing says the goal is to validate 'narratologically-motivated conditioning,' which is narrower than the conclusion's claim that 'our approach produces stories with higher overall quality.' Study 2 does evaluate the interaction, but with 12 participants and a baseline that differs on many interface dimensions, and the authors themselves call it preliminary. The paper therefore establishes that a richly conditioned pipeline can produce favorable stories and that users like the interface, but it does not establish that narrative keyframing, as a representation, causes the quality gain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces narrative keyframing, an interaction technique for AI-assisted creative writing in which writers specify plot events, character states, and first-person perspective keyframes at selected moments, and the system generates intervening third-person prose. The authors derive design goals from narratological theories of character arcs, characterization, and focalization; describe a three-view interface (Track, Table, Canvas); and evaluate the system through a technical comparison against a prompt-based LLM baseline using the WQRM-PRE metric and human ratings, plus a within-subjects user study with 12 writers comparing against a chatbot-based baseline. The paper claims that the approach produces stories with higher overall quality and richer characterization, and that it supports a more controllable, transparent, and engaging writing experience.","tokens_in":25526,"tokens_out":4556,"duration_ms":42259,"significance":"If the output-quality and user-experience claims held as stated, narrative keyframing—particularly perspective keyframes—would be a useful intermediate representation for controlling characterization and focalization in AI-assisted writing. The work is well grounded in narratology, uses an external quality model and external writing prompts for the technical evaluation, and the user study employs standardized instruments (CSI and AI System Experience). The perspective-keyframe idea is a genuinely interesting contribution to the design space of AI writing tools. However, the support for the output-quality claim is weakened by a confounded comparison, and the user study is small and system-specific; the conceptual contribution remains attractive but the evidence is not yet commensurate with the breadth of the conclusions.","major_comments":[{"comment":"The technical comparison is confounded. The keyframing condition receives the outline plus auto-suggested traits, generated first-person perspectives, and selected evidence, whereas the vanilla baseline receives only the outline. The observed preference (72/100 by the model; 83.3% human preference for overall quality) could therefore be caused by the extra conditioning information rather than by the keyframing representation. The authors themselves frame Study 1 as validating “narratologically-motivated conditioning” (§6.1), which is narrower than the conclusion that “our approach produces stories with higher overall quality.” An ablation—for example, giving the baseline the same trait lists and perspective-derived evidence, or removing perspective keyframes from the keyframing pipeline—is necessary to attribute the gain to the keyframing design.","section":"§6.1.1–6.1.2"},{"comment":"Study 1 does not exercise the defining property of keyframing, namely sparse user-specified constraints with interpolation between them. The pipeline generates traits for every character at every plot, then randomly selects two pieces of evidence per character per plot, with no user interaction and no comparison of keyframe density. Thus the result supports a richly conditioned pipeline, not the keyframing representation per se. The authors should either test sparse keyframes with interpolation against dense conditioning, or explicitly limit the technical claim to the conditioning scheme rather than to narrative keyframing as an interaction technique.","section":"§6.1.1"},{"comment":"The user study compares two systems that differ on many interface dimensions—keyframing views, color-coded evidence links, perspective generation, and selection mechanisms versus character sheets, character chatbots, and a story chatbot—so the observed differences in controllability, transparency, and enjoyment cannot be attributed specifically to narrative keyframing. With N=12 and the authors’ own characterization of the results as preliminary, the user-experience claim is suggestive but not conclusive. Additional interface-level ablations or a more matched baseline would be needed to support the broader claim that narrative keyframing, rather than the full system, causes the reported benefits.","section":"§6.2 and Appendix B.3"},{"comment":"The prompt titled “Interpolating Character Keyframes” in Appendix A.4 actually instructs the model to extract character traits from an existing narration, not to interpolate character states between keyframes. This makes the interpolation mechanism—a central feature of the keyframing analogy and a feature highlighted in Section 5.3—non-reproducible from the appendix. Please reconcile the prompt with the system description, or clarify whether interpolation is performed by a different prompt that is not shown.","section":"§5.3 vs. Appendix A.4"}],"minor_comments":[{"comment":"The abstract says “Through a user study” but the evaluation includes both a technical study and a user study; please refer to the two studies or adjust the wording.","section":"Abstract and §1"},{"comment":"There is a typo in the sentence “suggesting that the our system allowed users to better explore”; remove the extra “the”.","section":"§6.2.2"},{"comment":"The figure contains garbled and duplicated label text (for example, “Character A rcs” and repeated “Character & Perspective Keyframes” blocks). The final figure should be cleaned up so that the keyframe information flow is legible.","section":"Figure 1"},{"comment":"The sentence “To our knowledge, ours is the first work to explore first-person character narratives as an intermediate representation for controlling AI-assisted creative writing” is a strong novelty claim; consider softening it or providing a more systematic comparison with prior point-of-view or perspective-based writing tools.","section":"§3.3"},{"comment":"Reporting only W and p values for the Wilcoxon tests makes effect sizes hard to assess; adding a standardized effect size (e.g., rank-biserial correlation or matched rank-biserial) would strengthen the presentation.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Give this paper a real read. The core idea is genuinely new within AI writing tools: instead of global character sheets, you keyframe character states at story events, and you use first-person perspectives as an intermediate representation that can be selected and recombined into third-person prose. That is a real design-space opening, not just a prompt trick. The system is implemented well enough to run the studies, and the appendices include the actual prompts and participant details, which is more than most CHI/UIST papers ship.\n\nThe user study is the stronger half. 12 writers, within-subjects, against a chatbot-plus-character-sheet baseline. The reported differences on transparency, control, and exploration are plausible, and the interview quotes are specific enough to believe that the keyframing structure genuinely changed how participants worked. The authors are candid that the sample is small and call the results preliminary. I'd trust that part.\n\nThe soft spot is Study 1, the technical evaluation. The keyframing condition receives auto-suggested traits, generated first-person perspectives, and selected evidence; the baseline receives only the outline. So the 72/100 preference and the expert ratings may be driven by the extra conditioning material, not by the keyframing representation. The stress-test note is right about that. The paper partly disarms the concern by saying Study 1 validates 'narratologically-motivated conditioning' rather than the interaction. But the introduction and conclusion still say 'our approach produces stories with higher overall quality,' which overclaims relative to the design of Study 1. Also, the auto-generation and random evidence selection mean the technical study never exercises the actual interaction, so it's really a pipeline evaluation.\n\nThe fix is straightforward: either match the baseline for conditioning information (e.g., give both conditions the same traits, perspectives, and evidence), or drop the quality-difference claim from Study 1 and treat it only as a check that the conditioning doesn't hurt quality. The interaction claims rest on Study 2, and those are fine.\n\nI'd send this to reviewers. The concept is worth the field's time, and the issues are fixable with claim-tempering and one more study condition, not a rethink. I'd cite it if I were working on AI writing interfaces. Worth a reading group too, for the methodological lesson.","headline":"A real interaction concept for AI-assisted writing with solid user-study evidence, but the technical evaluation compares extra conditioning to an outline-only baseline rather than isolating the keyframing representation.","tokens_in":26011,"tokens_out":3156,"would_cite":true,"duration_ms":30415,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Narrative keyframing — setting plot, character, and first-person perspective constraints at key story moments — lets writers guide AI story generation with finer control, and the resulting stories rate higher in quality and…","keywords":["creative writing","human-AI interaction","narrative keyframing","perspective keyframes","characterization","large language models","user study","story generation"],"falsifier":"Conduct the technical evaluation with a flat-conditioning baseline that receives the auto-suggested traits, generated first-person perspectives, and selected evidence as additional prompt context but no keyframing interface; if that baseline matches the keyframing pipeline's scores on quality and characterization, the paper's claim that keyframing drives the improvement is refuted.","tokens_in":25105,"feed_emoji":"✍️","tokens_out":10158,"duration_ms":78095,"temperature":0.7,"pith_summary":"This paper introduces narrative keyframing, an interaction technique that transplants animation's keyframe-and-interpolate idea into AI-assisted creative writing. Instead of giving a language model a single global prompt, writers mark key moments in a story with plot keyframes, character keyframes, and first-person perspective keyframes, and the model fills in the intervening prose. The paper argues that this design gives writers a more controllable, transparent, and engaging way to work with generative AI, and reports evidence that it produces stories with richer characterization and higher overall quality than a standard LLM baseline. A sympathetic reader would care because it offers a concrete answer to a central problem in human-AI co-writing: how to keep authorial intent in control while letting the model do the drafting.","feed_headline":"Narrative keyframing makes AI stories richer and more controllable","feed_subtitle":"In tests with 12 writers, keyframed AI writing beat plain prompting on characterization, quality, and perceived control.","key_machinery":"The central mechanism is narrative keyframing, a three-track representation: plot keyframes anchor the timeline of events; character keyframes record a character's physiology, psychology, and sociology at a given plot point; and perspective keyframes are first-person narratives generated from those character states that externalize how a character experiences an event. These tracks are linked — editing a perspective keyframe updates the character keyframe and vice versa — and selected evidence from perspective keyframes is injected into the third-person generation, with color coding and snippet usage tracking making the influence visible. The interpolation step is supplied by the language model: given sparse constraints at key moments, the underlying model generates the intervening character development, first-person reflections, and final prose.","core_discovery":"The paper's central claim is that narrative keyframing supports a more controllable, transparent, and engaging way to use generative AI in creative writing, and that stories produced through it are rated as having higher overall quality and richer characterization than stories from a standard LLM baseline. The key move is treating first-person narratives as an intermediate representation: writers generate each character's perspective for an event, select textual evidence they want, and the model recombines that evidence into a third-person narrative, with color-coded highlights making the connection traceable. The paper further claims this approach is the first to use first-person character narratives as an intermediary for controlling AI-assisted creative writing.","pith_inferences":["Inference: A simpler tool that generates first-person drafts before third-person narration might reproduce the characterization gains without a full keyframing interface, which would imply the representation, not the timeline, is the essential ingredient.","Inference: Keyframing other narrative properties — tone, pacing, or scene atmosphere — is a natural next step; a reader willing to extend the paper would predict the same control and traceability benefits when those properties are keyframed event by event.","Inference: The traceability principle may transfer to other generative tasks: showing the exact mapping from user input to generated output could increase perceived transparency and reflection in code or image generation as well.","Inference: Testing keyframing on non-linear plots such as flashbacks would reveal whether the interpolation-based interaction degrades when the story timeline is not monotonic, since the paper's current design supports only sequential progression."],"forward_implications":["If narrative keyframing works as claimed, AI story tools would replace static global prompts with timeline-based interfaces in which writers localize control at key moments.","Because the same underlying model was used in both conditions, the measured gains in quality and characterization point to the keyframing representation itself as the source of improvement.","The color-coded traceability between selected evidence and final text could become a standard expectation for human-AI co-writing, supporting both verification and reflection.","The keyframing structure makes character arcs explicit and editable, allowing writers to shape how characters change across acts rather than maintaining one static persona."],"supporting_citations":[{"why":"supplies the animation keyframing principle from which the interaction metaphor is borrowed.","marker":"[29]"},{"why":"supplies the automatic quality model used for pairwise story comparison in the technical evaluation.","marker":"[5]"},{"why":"supplies ten of the writing prompts used in the technical evaluation.","marker":"[31]"},{"why":"supplies the other ten writing prompts for the technical evaluation.","marker":"[13]"},{"why":"defines the character-chatbot baseline approach that the user study compares against.","marker":"[42]"},{"why":"defines another character-chatbot baseline approach used in the user study comparison.","marker":"[38]"},{"why":"provides the typology of textual indicators used to generate and extract evidence from first-person perspectives.","marker":"[41]"},{"why":"supplies the three trait dimensions (physiology, psychology, sociology) that structure character keyframes.","marker":"[12]"},{"why":"supplies the Creativity Support Index questionnaire used to measure the user experience.","marker":"[6]"},{"why":"supplies the AI System Experience questionnaire used to measure perceived control and transparency.","marker":"[55]"}],"fun_headline_variants":["Keyframing AI stories: writers gain control, quality rises","Narrative keyframing: first-person tales steer AI prose","Study: narrative keyframing beats plain LLM prompting","AI writing tool uses keyframes for richer, controllable stories","Keyframed narratives give writers transparent AI control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measured benefits are attributed to the keyframing representation itself; if feeding the same traits, first-person perspectives, and selected evidence directly to the baseline produced the same gains, the central claim would not be supported.","fun_headline_variants_meta":{"raw":{"variants":["Keyframing AI stories: writers gain control, quality rises","Narrative keyframing: first-person tales steer AI prose","Study: narrative keyframing beats plain LLM prompting","AI writing tool uses keyframes for richer, controllable stories","Keyframed narratives give writers transparent AI control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2328,"prompt_tokens":844,"completion_tokens":1484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":1403}},"tokens_in":460,"tokens_out":1484,"duration_ms":9726,"temperature":1.0,"reasoning_tokens":1403,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:21:08.915090+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the technical evaluation with a flat-conditioning baseline that receives the auto-suggested traits, generated first-person perspectives, and selected evidence as additional prompt context but no keyframing interface; if that baseline matches the keyframing pipeline's scores on quality and characterization, the paper's claim that keyframing drives the improvement is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the animation keyframing principle from which the interaction metaphor is borrowed."},{"cited_title":"2003.Narrative Fiction: Contemporary Poetics","cited_arxiv_id":null,"evidence_quote":"provides the typology of textual indicators used to generate and extract evidence from first-person perspectives."},{"cited_title":"1995.The Art Of Dramatic Writing: Its Basis in the Creative Interpreta- tion of Human Motives","cited_arxiv_id":null,"evidence_quote":"supplies the three trait dimensions (physiology, psychology, sociology) that structure character keyframes."}],"review_version":1}