{"id":"b9c15a85-c18e-4d48-9f52-6ce7d77fe923","arxiv_id":"2501.06607","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A drawing agent that logs and visualizes human-AI co-creation through the co-creative sense-making framework, tested on ten sessions.","lead":"This paper describes a web-based drawing system where a human and an AI agent draw together on a shared canvas while the software automatically logs and charts the interaction. It introduces the co-creative sense-making framework as a way to turn collaboration into quantitative time-series data, demonstrated on ten five-minute sessions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CCSM coding weights in Table 2 are arbitrary and drive all Section 7 results; without robustness checks or construct validation, the claimed quantification of co-creation may be an artifact.","rationale":"The reader's weakest assumption correctly identifies the coding scheme in Table 2 as the linchpin of the measurement claims. I agree that this is the most load-bearing concern: the system's automatic logging is straightforward, but the interpretation of those logs as cognitive states (clamped/unclamped) is entirely mediated by the numeric codes. The paper itself acknowledges the arbitrariness of the polar values and, in Section 10, limits the case study to demonstration and notes that the user is the designer. This honesty, plus the existence of a publicly accessible platform, supports accepting the system presentation as a contribution. However, the claim that the results 'help validate CCSM' is not supported until the coding validity is established. As a result, the appropriate verdict remains conditional: the platform contribution can be accepted, but the framework's measurement claims require independent validation. No additional concern outweighs this one; the absence of released data/code is secondary given the paper's stated scope.","tokens_in":1015,"tokens_out":752,"duration_ms":51484,"concrete_test":"Recompute the Section 7 analyses from the raw logged interaction data (or reconstruct them) using several alternative monotonic codings of the four modes, e.g., [-1,0,0.5,1], [-2,-1,0,1], [-1,-0.5,0.5,1], while preserving the ordinal order. If the significant group differences in average coded value, CSM slope, and the qualitative trend classifications are not invariant across these codings, the results are artifacts of the arbitrary weights. As a complementary check, have independent raters code video of the sessions directly for clamped/unclamped cognition and measure agreement with the Table 2 mapping.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.2 introduces the coding convention (Table 2) that maps interaction modes to numbers: communicate = 1, manipulate interface = 0.5, wait = 0, execute = -1. Every derived result in Section 7 — the average coded value, the creative sense-making curve, its slope, and the MACD trend classifications — is a cumulative sum of these numbers. The paper states that 'the polar values of the continuum are arbitrary' and offers no theoretical or empirical justification for the specific spacing (e.g., why communication is twice interface manipulation, why waiting is 0 rather than -0.5). Since the abstract vs. representational comparison is essentially a weighted sum of raw mode counts, the significant p-values (e.g., p=.0003 for average coded value, p=.002 for slope) may simply reflect the chosen weights. Under a different but equally plausible monotonic coding (e.g., execute = -2, wait = -1, manipulate = 0.5, communicate = 1), the magnitude and even the direction of the derived metrics can change, and the reported differences could vanish or reverse. Consequently, the central claim that the system 'quantifies, models, and visualizes the co-creative process' and that the case study 'helps validate CCSM' depends on an unvalidated and arbitrary numerical mapping.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the AI Drawing Partner, a web-based co-creative drawing agent that also logs user and agent interactions according to the Co-Creative Sense-Making (CCSM) framework. The CCSM maps interaction modes to cognitive states (clamped/unclamped cognition) and assigns numeric codes (communicate=1, manipulate interface=0.5, wait=0, execute=-1). The system records raw interaction counts and computes cumulative sums to form creative sense-making (CSM) curves; slopes and stock-market-style MACD trend classifications are derived from these curves. A case study reports ten five-minute co-creative drawing sessions, five abstract and five representational, all conducted by the first author, who also designed the system. The paper finds significant between-group differences in average coded value, CSM slope, interaction counts, and other metrics, and interprets these as helping to validate CCSM. The primary claimed contribution is the AI Drawing Partner as a unique quantified co-creative AI system and research platform.","tokens_in":26197,"tokens_out":3093,"duration_ms":29368,"significance":"If the validation claims were supported, the paper would provide a useful open platform for studying co-creation and a domain-independent framework for quantifying interaction dynamics. The system is publicly accessible, automatically logs interaction data, and the companion analysis pipeline is a concrete step toward reproducible process-level analysis of human-AI co-creation. The adoption of the COFI framework to situate the system and the detailed system architecture are informative. However, the central evidential claim—that the case study validates CCSM by showing significant differences between abstract and representational sessions—is not supported by the data as analyzed. The numerical coding convention that drives every reported statistic is acknowledged in the paper to be arbitrary, and all sessions come from a single participant who is also the system designer. The paper is better characterized as a systems and demonstration contribution than as a validation of the CCSM framework.","major_comments":[{"comment":"The coding values in Table 2 are the foundation of every quantitative result in Section 7. The paper states that 'the polar values of the continuum are arbitrary' and offers no theoretical or empirical justification for the spacing among communicate=1, manipulate=0.5, wait=0, and execute=-1. The average coded value, CSM slope, and MACD trend classifications are all monotone functions of these weights. The significant differences reported in Section 7 (e.g., p=.0003 for average coded value, p=.002 for slope) therefore largely reflect the fact that abstract sessions contain more drawing (coded -1) and representational sessions contain more communication (coded +1), which is a property of the chosen coding scheme rather than an independent measurement of cognitive dynamics. I request a robustness analysis using alternative monotone codings (e.g., execute=-2, wait=-0.5, manipulate=0.5, communicate=1) or a non-parametric analysis based on raw interaction-mode counts, or a substantial weakening of the validation claim in the abstract and conclusions.","section":""},{"comment":"All ten sessions were conducted by the first author, who designed the AI Drawing Partner. The paper acknowledges in Section 10 that 'the user was also the designer of the AI Drawing Partner, and he was thoroughly familiar with the interface,' yet Section 7 treats the ten sessions as independent observations for statistical testing. With a single participant, these sessions are not independent replicates, and the reported p-values (e.g., p=.002 for slope, p=.0001 for turns) cannot support population-level claims about co-creation or about the validity of CCSM. At most, the case study demonstrates that the pipeline can detect differences in the designer's own behavior across two self-selected conditions. I recommend reframing the reported statistics as descriptive, and moving the validation claim to future work with external participants.","section":""},{"comment":"The claim that 'the results help validate the CCSM by showing significant differences' is not justified by the presented evidence. Because the abstract and representational sessions were deliberately chosen by the same individual who designed the system, the observed differences are expected from the experimental setup; the analysis then interprets those differences through the same coding convention that produced them. This creates a circularity concern that is not addressed by the paper. I suggest that the paper either (a) present the case study strictly as an illustrative demonstration of the analytics pipeline, with the validation claim removed, or (b) add a validation study with independent participants, pre-registered hypotheses, and robustness checks on the coding scheme.","section":""},{"comment":"The MACD trend analysis uses parameters (12-period EMA, 26-period EMA, 9-period signal) that are standard for financial data but are not justified for the interaction data sampled at 0.5-second intervals. Since the entire trend-classification visualization in Figure 9 depends on these periods, the paper should include a sensitivity analysis or at least a justification for why these specific windows are appropriate for co-creative interaction data. Without this, the trend sequences are one arbitrary choice among many, particularly given that the underlying coding values are already arbitrary.","section":""}],"minor_comments":[{"comment":"The word 'Maping' in the section title appears to be a typo; it should read 'Mapping.'","section":""},{"comment":"Reference [25] contains the typo 'Co-Creativve AI' in the title; please correct it.","section":""},{"comment":"The statement that 'Waiting is coded as 0 so waiting and non-action do not influence the direction of the trend' is unclear, since waiting can influence the slope when it appears between other coded events; please clarify whether the coding is applied per time step or per interaction event.","section":""},{"comment":"The sentence 'A line is calculated as the content between a pen down event and when the user raises their pen' defines a line as a stroke, but elsewhere the paper also refers to 'lines' as discrete algorithmic outputs; please make the terminology consistent.","section":""},{"comment":"The phrase 'the agent is actively engaged in sense-making' in the context of wait time seems to describe the user's cognitive state rather than the agent's; please revise for clarity.","section":""}],"recommendation":"major_revision","confidential_remarks":"The paper would be better framed as a systems and demonstration contribution. The current abstract and conclusions overclaim validation, and the statistical analysis in Section 7 is not appropriate for a single-participant case study. I would urge the editor to require either a substantial rewrite that removes the validation claim or the addition of a proper external validation study before considering publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid system paper and a plausible but unvalidated measurement framework. The AI Drawing Partner itself is real, described in enough detail to build on, and the CCSM data schema is a reasonable attempt to standardize co-creation metrics. But the case study does not validate CCSM. The significant differences between abstract and representational sessions are baked into the coding convention in Table 2.\n\nWhat's new: the integrated system that logs interaction modes and visualizes them as CSM curves is new, and applying MACD stock-market analysis to interaction time series is a fresh idea. The system is freely available, and the architecture description is credible. The COFI comparison is well done and the literature coverage is adequate. Credit where due: the paper is honest about the case study being demonstrative, and the limitations section explicitly says the user was the designer. That is more transparent than most.\n\nSoft spots: the stress-test is correct and it matters. The paper states the polar values are arbitrary, yet every result in Section 7 — average coded value, CSM slope, MACD trends — is a weighted sum of those arbitrary numbers. Drawing = -1, waiting = 0, interface = 0.5, communication = 1. A representational session with lots of communication will mechanically rise; an abstract session with lots of drawing will mechanically fall. The p-values restate the coding. Changing the weights, as the stress-test notes, can change magnitude and possibly direction. The paper needs robustness checks (e.g., alternative monotonic codings) or external construct validation before claiming CCSM measures sense-making.\n\nAlso soft: the abstract/representational split is post hoc, with n=10 sessions all by the first author. The paper acknowledges this in limitations, but then the conclusion still says the results \"help validate CCSM.\" That overstates. And no released data or code for the companion app is linked, which hurts reproducibility. Minor: multiple comparisons are not controlled; with roughly fifteen tests, a few p<.05 are expected.\n\nThe citation pattern looks fine; self-citations are to prior CSM work that this directly extends, which is legitimate.\n\nWho it's for: co-creative AI researchers who want a ready-made platform and a concrete proposal for a common metric. They get a useful system description and a clear target for validation work. The framework should not be adopted as validated until independent studies with naive users and released artifacts appear.\n\nRecommendation: this deserves serious peer review. The system is a real contribution and the CCSM idea is worth airing, but the measurement claims need to be reframed as a proposal, and the robustness problem needs to be addressed. Accept with major revision would be reasonable.","headline":"A genuinely useful system paper wrapped around a measurement framework whose validation is undercut by its own coding choices.","tokens_in":26664,"tokens_out":1723,"would_cite":false,"duration_ms":16510,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims a drawing agent can collaborate with a user and automatically quantify, model, and visualize the co-creative process, with case-study data validating the measurement framework.","keywords":["human-AI co-creation","co-creative drawing agent","co-creative sense-making framework","interaction dynamics","creative sense-making curve","enaction","quantified co-creative system","computational creativity"],"falsifier":"Record a set of co-creative sessions in which the user's moment-by-moment intentions are captured (e.g., by retrospective protocol analysis or think-aloud) and compare them against the automatically coded curve: if reported 'reflecting on the artwork' moments produce the same curve pattern as moments of merely waiting for the agent to finish, the cognitive interpretation of the wait category is falsified; if the coded phases instead align with reported intentions across many users, the framework's validity is supported.","tokens_in":25678,"feed_emoji":"🎨","tokens_out":8323,"duration_ms":68000,"temperature":0.7,"pith_summary":"The paper presents the AI Drawing Partner, a web-based agent that draws with a user in real time on a shared canvas while automatically logging every interaction. Those logs are structured by the co-creative sense-making (CCSM) framework, which assigns each action a code on a clamped-to-unclamped cognition scale: communicating counts as $1$, manipulating the interface as $0.5$, waiting as $0$, and executing a drawing action as $-1$. Summing the codes over time produces a creative sense-making curve that shows whether a session trends toward executing, regulating, or waiting, and the paper argues this turns qualitative theories of participatory sense-making into quantitative, comparable process data. A case study of ten five-minute sessions finds statistically significant differences between abstract and representational drawing (curve slope $p=.002$, communication $p=.003$, turns $p=.0001$), which the authors read as evidence that the framework captures real differences in co-creative experience. If the approach holds, co-creative AI researchers gain a common, domain-independent metric for comparing co-creation across systems and domains.","feed_headline":"AI drawing partner turns co-creation into a measurable curve","feed_subtitle":"A new framework codes each stroke, click, and pause, and the data cleanly separate abstract from representational sessions.","key_machinery":"The load-bearing mechanism is the creative sense-making coding convention: four interaction modes mapped to numeric values on a clamped-to-unclamped cognition continuum — communicate to the AI $= 1$, manipulate the interface $= 0.5$, wait $= 0$, execute a drawing action $= -1$. Continuously applied to the interaction log, these codes form a time series whose cumulative sum is the 'creative sense-making curve': rising segments mean the partner is regulating the interaction, falling segments mean fluid execution, and flat segments mean waiting. Linear regression on the curve yields its slope, and a moving-average convergence-divergence (MACD) analysis with exponential moving averages classifies each time step as regulate (buy), execute (sell), or wait (hold), producing the visualized trend sequences. This machinery converts an enactive theory of sense-making into a dataset the system records automatically — no human video coding, no inter-rater reliability — and every reported statistic, curve, and trend visualization is computed from it.","core_discovery":"The paper's central discovery is that a co-creative system can be both the collaborator and the measurement instrument: the AI Drawing Partner draws with the user and, in the same run, emits a complete quantitative record of the co-creative process. The record is generated by the co-creative sense-making framework, which draws on enactive cognitive science to define four data categories — cognitive dynamics, interaction dynamics, collaboration dynamics, and domain behaviors — and continuously codes user actions onto a continuum from clamped cognition (fluently executing, $-1$) to unclamped cognition (communicating, $1$). From this coded time series the system builds the creative sense-making curve, applies a linear regression for its slope, and uses stock-market-style moving-average analysis to label each time step as regulate, execute, or wait. In the demonstrative case study, the user's execute codes ($p=.039$), communication ($p=.003$), interface manipulation ($p=.009$), average coded value ($p=.0003$), CSM-curve slope ($p=.002$), lines drawn ($p=.005$), and turns ($p=.0001$) all differed significantly between five abstract and five representational sessions. The authors take these differences as validating CCSM: a quantitative method that coincides with the qualitative difference between the two creative strategies.","pith_inferences":["The coding scale's numerical spacing — $1$, $0.5$, $0$, $-1$ — is treated as interval data, yet the paper concedes the poles are arbitrary; a different monotone encoding would preserve the sign of slopes but could change which group comparisons reach significance, so cross-study comparability rests on the whole community adopting the identical convention.","The 'wait' code conflates at least three distinct experiences — user reflection, user waiting for the agent, and user watching the agent draw — and the paper itself flags the first two; separating them in the coding scheme would plausibly change flat segments of the CSM curve and therefore some trend classifications.","The case study's single user was the system's designer and thus already fluent with the interface, so the large effect sizes may overstate what naive users would show; a larger study with naive participants would test whether the significant differences replicate.","The framework invites a direct validity check the paper does not perform: pair each automatically detected trend phase with a retrospective protocol label of the user's intention, and test whether the curve's 'regulate' phases coincide with reported moments of reflection or evaluation."],"forward_implications":["If CCSM is valid, any co-creative system that adopts the coding schema produces interaction data comparable to any other, enabling within-domain and cross-domain comparison of co-creative experiences.","Automatic coding removes human coders from the loop, eliminating inter-rater reliability checks and reducing bias and error in the analysis of co-creative interaction.","The trend-sequence visualization splits a session into regulate, execute, and wait phases, allowing researchers to study co-creation in segments rather than as a single aggregate score.","Because the platform is public and the analysis pipeline is automated, the AI Drawing Partner can serve as an off-the-shelf experimental platform for studying human-AI drawing without building a new system.","Future work (per the paper) suggests the CSM curve can become a real-time model of user intent, letting the agent adapt its contributions to detected phases of ideation or refinement."],"supporting_citations":[{"why":"Supplies the participatory sense-making theory of interaction dynamics and interaction coupling that CCSM operationalizes.","marker":"[33]"},{"why":"Defines the creative sense-making coding framework and the CSM curve that the AI Drawing Partner implements.","marker":"[30]"},{"why":"Provides the COFI interaction-design dimensions used to situate the system and the survey finding that most co-creative systems lack AI-to-human communication.","marker":"[88]"},{"why":"Provides the Stable Diffusion text-to-image model behind the system's image-generation mode.","marker":"[91]"},{"why":"Provides the Sketch-RNN model used to generate object sketches in the shared canvas.","marker":"[49]"},{"why":"Supplies the improvisation offer/accept/elaborate vocabulary used for the collaboration-dynamics category.","marker":"[45]"},{"why":"The public AI Drawing Partner platform itself, where the quantified co-creative sessions are run.","marker":"[24]"},{"why":"The earlier co-creative drawing agent with object recognition whose interaction design the AI Drawing Partner builds on.","marker":"[31]"}],"fun_headline_variants":["AI drawing partner logs every stroke to model co-creation","Co-creative AI: draws with you, then quantifies the process","This AI co-creator measures creativity as you draw together","Drawing AI that also acts as a research platform for co-creation","AI partner that draws and tracks your creative synergy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything measured rests on the coding assumption that drawing equals clamped cognition ($-1$), waiting equals $0$, interface manipulation equals $0.5$, and communication equals unclamped cognition ($1$) — a mapping the paper itself admits has arbitrary polar values, so if these numbers do not track genuine sense-making, the curves, slopes, trend classifications, and all the significant differences in Section 7 are artifacts of the coding choice rather than measurements of co-creation.","fun_headline_variants_meta":{"raw":{"variants":["AI drawing partner logs every stroke to model co-creation","Co-creative AI: draws with you, then quantifies the process","This AI co-creator measures creativity as you draw together","Drawing AI that also acts as a research platform for co-creation","AI partner that draws and tracks your creative synergy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1499,"prompt_tokens":1065,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":681,"completion_tokens_details":{"reasoning_tokens":350}},"tokens_in":681,"tokens_out":434,"duration_ms":4825,"temperature":1.0,"reasoning_tokens":350,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:56:04.809496+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a set of co-creative sessions in which the user's moment-by-moment intentions are captured (e.g., by retrospective protocol analysis or think-aloud) and compare them against the automatically coded curve: if reported 'reflecting on the artwork' moments produce the same curve pattern as moments of merely waiting for the agent to finish, the cognitive interpretation of the wait category is falsified; if the coded phases instead align with reported intentions across many users, the framework's validity is supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the participatory sense-making theory of interaction dynamics and interaction coupling that CCSM operationalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the COFI interaction-design dimensions used to situate the system and the survey finding that most co-creative systems lack AI-to-human communication."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the improvisation offer/accept/elaborate vocabulary used for the collaboration-dynamics category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The earlier co-creative drawing agent with object recognition whose interaction design the AI Drawing Partner builds on."}],"review_version":1}