{"id":"3a2a6d7b-f7ee-4aad-8594-cb9c0dccb686","arxiv_id":"2607.26198","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across four interface paradigms, tool design shapes how data wrangling is approached but does not determine cleaning success.","lead":"A 40-person study compared data cleaning in Excel, Jupyter, ChatGPT, and OpenRefine, finding that tools shape how people approach cleaning but not how well they finish. The paper maps these differences as design trade-offs for building better wrangling tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unmeasured external AI use in non-ChatGPT conditions confounds the central causal claim that tool affordances steer wrangling strategies.","rationale":"The reader's weakest assumption bundled external help with deployment/config quirks (JupyterLite vs. JupyterLab, Excel web, noVNC). I agree that the hosting details matter, but the sharper load-bearing threat is the asymmetric, unmeasured use of external AI in the non-ChatGPT conditions. If participants in Excel/Jupyter/OpenRefine routinely consulted ChatGPT, then the 'treatment' is a combination of the assigned tool and an AI co-pilot, and the paper's causal language about tool affordances is not supported. The paper explicitly acknowledges the open-resource design, so this is not a hidden flaw; it is a recognized limitation. However, because the main positive claim is qualitative and rests on strategy coding, the limitation is more than a boundary condition—it directly threatens the attribution. The proposed test—stratifying the existing qualitative analysis by measured external AI use—would settle whether the trade-offs (e.g., opportunistic vs. systematic cleaning) survive when external AI is absent or minimal. If they do, the claim is robust; if not, the paper should be revised to describe 'primary tool plus ambient AI' effects. The reader's conditional verdict already allows for revisions, so I recommend UNCHANGED rather than a stronger penalty; the concern strengthens the rationale for the requested revisions but does not by itself push the verdict to REJECT, since the study is explicitly exploratory and the qualitative observations are internally consistent. I also note that the single-coder TDoPS coding is a genuine reliability concern, but I see external AI contamination as the more load-bearing issue because it undermines the causal interpretation of whichever codes are produced.","tokens_in":21614,"tokens_out":6550,"duration_ms":70603,"concrete_test":"Re-analyze the existing screen recordings and think-aloud data from the 40 participants: code minutes spent in external AI sites (ChatGPT/Claude, etc.) per participant in the Excel, Jupyter, and OpenRefine conditions. Then reproduce the Sec. V-C analysis of opportunistic vs. systematic cleaning and the Table I trade-offs within the subset of participants with no or below-median external AI use. If the qualitative pattern is unchanged, the confound is not load-bearing; if it shifts or disappears, the paper must be revised to weaken the causal claim and reframe results as effects of 'primary tool + ambient AI support.' This check uses data already collected and requires no new experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central affirmative claim—that tool affordances 'steer user strategies' (Abstract, Sec. VII)—requires that observed strategy differences (e.g., data-centered vs. abstraction-centered cleaning) be attributable to the assigned tool's interface paradigm. But Sec. V-B reports that 'most participants needed external help (e.g., AI or docs) to finish,' and Sec. VII states that participants in every condition were allowed to consult external resources, with 'those who were in non-AI conditions usually turn[ing] to AI (such as ChatGPT and Claude) when they needed help writing formulas or expressions.' Since ChatGPT participants were instructed to stay in the ChatGPT window while Excel/Jupyter/OpenRefine participants could freely use ChatGPT as a co-pilot, the effective treatment is not 'tool X' but 'tool X plus (often) an AI assistant.' The paper lists this as a limitation but never quantifies external AI use nor tests whether the coded strategy patterns hold among participants with little or no external AI help. If the 'opportunistic cleaning' or 'systematic transformations' patterns are driven by differential AI uptake rather than by the grid/facet/notebook interface, the central claim is an artifact of the study's open-resource design rather than a property of the tools. This is a correctness risk for the main positive result, not just a boundary condition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a between-subjects observational study (N=40) in which participants performed two data-cleaning tasks in one of four tools—Jupyter, Excel, ChatGPT, and OpenRefine—while thinking aloud under screen capture. The authors evaluate outcomes with Jaccard similarity against author-constructed 'silver tables' and analyze process qualitatively through the Technical Dimensions of Programming Systems (TDoPS) framework. They conclude that, within this study, tool affordances steer wrangling strategies but do not determine outcomes: no tool consistently outperformed others, and the main differences are trade-offs between data-centered and abstraction-centered interfaces. Data-centered tools are said to encourage opportunistic cleaning, abstraction-centered tools to support systematic transformation at higher cognitive cost, and ChatGPT to reduce effort while obscuring intermediate states and risking automation bias. The paper explicitly forgoes statistical inference, is exploratory in tone, and provides open materials, analysis code, and replays.","tokens_in":21916,"tokens_out":5253,"duration_ms":57920,"significance":"If its central qualitative claim survives the confounding concerns, the paper would be a useful empirical contribution to HCI and data-wrangling research: it demonstrates a systematic way to compare interface paradigms via TDoPS, supplies rich observation-based accounts of how feedback immediacy, error handling, notational structure, and reuse mechanisms shape behavior, and responsibly avoids overclaiming performance differences. The strongest features are the openly available study materials and analysis code, the use of screen recordings and think-aloud data, the explicit acknowledgment of the exploratory status, and several concrete, falsifiable observations that future work can test. The main risk is that the central causal attribution—interface paradigm steers strategy—is not yet isolated from external AI use and deployment differences. This is a fixable weakness rather than a fundamental flaw, provided the authors can re-analyze their recordings or soften the claim accordingly.","major_comments":[{"comment":"Section V-B states that 'most participants needed external help (e.g., AI or docs) to finish,' and Section VII reports that non-AI participants 'usually turned to AI (such as ChatGPT and Claude)' when writing formulas or expressions. Because the paper's central claim is that the assigned tool's interface paradigm steers wrangling strategy, uncontrolled external AI use is a direct confound: the effective treatment in three of four arms is 'tool X plus an AI assistant.' The limitation is acknowledged but not quantified or conditioned on. I request a post hoc analysis of the screen recordings and think-aloud data: (i) count and characterize external-AI use by condition; (ii) compare the coded strategy patterns (opportunistic vs. systematic cleaning, inspection behavior, error handling) for participants with little or no external AI help against those with substantial help; and (iii) state w","section":"§V-B, §VII, §III"},{"comment":"The silver tables are author-constructed reference solutions, and the clustering in Fig. 4 is used to support claims such as 'tools alone do not drive successful outcomes' and the observation that a ChatGPT-heavy cluster appears in Task 1. The paper correctly disclaims statistical inference, but the cluster-level reading is used in substantive arguments. The cluster labels ('partial/no cleaning', 'aggressively delete columns') were assigned after inspecting outputs with the same team's coding; no inter-rater reliability or sensitivity analysis is reported. With a modest sample and Jaccard similarity on raw rows (plus manual column canonicalization for Task 2), the clusters may be sensitive to distance measure, preprocessing, and the particular silver-table set. Please report robustness of the clustering (e.g., alternative distance/similarity measures, subsets of silver tables), provide p","section":"§IV, Appendix B"},{"comment":"The paper does not describe the assignment mechanism (random vs. balanced by experience) and recruitment used tool-specific advertisements targeting users familiar with each tool. Fig. 7 suggests that self-reported experience differs visibly across conditions. In addition, the compared conditions are not 'the paradigm' but specific deployments: Excel for the web, ChatGPT-5.2 through a Windows VM via noVNC, OpenRefine 3.10.0, and JupyterLite rather than hosted JupyterLab. Because the qualitative comparisons attribute observed differences to interface paradigm, baseline differences in expertise and the differing deployment contexts are confounds. At minimum, state the assignment procedure, report condition-level experience descriptively, and discuss how the deployment differences could mimic or mask paradigm effects. Ideally, incorporate experience into the quantitative description and int","section":"§III, Fig. 7, Appendix: Experimental Configuration Details"}],"minor_comments":[{"comment":"Participant IDs are formatted inconsistently (e.g., P AI7 vs. PAI8, POR5 vs. P OR5, PXL1 vs. P XL1). Please standardize the spacing convention.","section":"Throughout §V"},{"comment":"The phrase '5.2±0.8 on a 7-point scale' lacks a clear referent: what exactly was rated, and in which post-task survey item? Please specify.","section":"§V-E"},{"comment":"The sentence 'For Task 2, we canonicalized the columns via a manual alignment, including inserting columns of nulls for those that had been deleted and enforcing a consistent ordering' is grammatically awkward; consider rephrasing for clarity.","section":"§IV"},{"comment":"The figure is dense, especially for Task 2. Larger fonts, separate panels per task, or explicit per-tool row/column annotations would improve legibility.","section":"Fig. 4"},{"comment":"Because the thematic analysis is deductive with codes drawn from TDoPS, the finding that TDoPS dimensions describe observed behavior is partly by construction. Please frame the section accordingly (e.g., as an application of the framework rather than an independent validation of it).","section":"§V introduction"}],"recommendation":"major_revision","confidential_remarks":"The external-AI-use confound is the main correctness risk and is fixable because the study captured screen recordings and think-aloud audio; the revision should quantify external AI use or substantially soften the causal language in the abstract and Section VII. The paper is otherwise a reasonable fit for an HCI venue, and the open materials are a strength. I would not reject, but I would not accept without the re-analysis or a reframing that explicitly treats the study as comparing deployed configurations plus ambient AI support rather than isolated paradigm-level treatments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read this with the reader's take in front of me, and I largely agree with the conditional verdict. The paper is worth engaging: it produces a new empirical result — a structured comparison of notebook, spreadsheet, conversational AI, and visual wrangler tools for data cleaning, using TDoPS as a scaffold. The qualitative observations are concrete and well-illustrated with participant quotes and screen-capture evidence. The quantitative section is properly cautious: no statistical inference, silver tables explicitly framed as references rather than gold standards. The authors repeatedly hedge with 'within the context of our study,' which is appropriate for an exploratory N=40 study. I do believe the central claim — tool affordances steer strategies but don't determine outcomes — is not just defined into existence; it's supported by specific observed trade-offs such as data-centered vs abstraction-centered cleaning, error visibility vs attentional capture, structured reuse vs ad hoc repetition. That's a genuine contribution.\n\nThe biggest soft spot is the one the stress-test note names: external AI use is unmeasured. The authors acknowledge in Sections VI and VII that participants in all conditions could and often did consult ChatGPT or other AI tools, and that non-AI conditions 'usually turned to AI' for formulas. That means the 'tool' condition is really 'tool plus optional AI assistant,' and the paper never quantifies how often or how much participants relied on external AI, nor does it test whether the strategy patterns hold among participants who didn't. That doesn't sink the paper — the authors explicitly interpret findings as 'primary tool environments rather than isolated treatments' — but it does mean the causal-sounding 'steer' language in the abstract overstates what can be concluded. I'd want the authors to either soften the framing or provide some sensitivity analysis (e.g., compare coded strategies for participants with low vs high external AI use). Minor additional issues: the qualitative coding used a single primary coder without inter-rater reliability (the appendix gives procedural criteria but not agreement metrics); the silver-table baselines are internally constructed, though the appendix shows reasonable variants; and the specific configurations (JupyterLite, Excel web, noVNC VM) are convenient but not necessarily representative. None of these are fatal; they're typical of a first exploratory study.\n\nWho is this for? Researchers in HCI, data wrangling, and end-user programming. Practitioners might also find the trade-off table useful. It deserves a serious referee — I would not desk-reject it. The authors have done the work to make the materials available (OSF, GitHub), and the paper is honest about its limits. My recommendation: send it to peer review with a request for the external-AI sensitivity analysis and a modest tightening of the causal language. If the authors can't provide the sensitivity analysis, they should at least explicitly reframe the contribution as descriptive rather than causal.","headline":"A worthwhile exploratory study with a new TDoPS-based comparison; the unmeasured AI co-pilot confound means the central claim is more suggestive than causal.","tokens_in":22344,"tokens_out":2696,"would_cite":true,"duration_ms":25231,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tool affordances steer data-wrangling strategies but do not determine outcomes, a 40-person study across four interface paradigms finds.","keywords":["data wrangling","interface paradigms","user study","technical dimensions","spreadsheets","notebooks","conversational AI","visual wranglers"],"falsifier":"An experiment that assigns the same participants to the same two paradigms with hosting and versions matched (e.g., desktop Excel versus hosted JupyterLab) and measures both process (operation sequences, planning versus opportunistic actions) and outcome would settle whether paradigm-level differences persist; if operation sequences become indistinguishable across paradigms once hosting and external help are controlled, the central claim fails.","tokens_in":21528,"feed_emoji":"🧹","tokens_out":3751,"duration_ms":35926,"temperature":0.7,"pith_summary":"This paper tries to establish that the interface paradigm of a data-wrangling tool—spreadsheet, notebook, conversational AI, or visual wrangler—shapes how people approach cleaning tasks, not just how well they finish. Across 40 participants using Jupyter, Excel, ChatGPT, or OpenRefine, no tool consistently produced better results, and outputs within each tool varied widely. The study's key finding is a set of trade-offs: data-centered interfaces like Excel and OpenRefine invite opportunistic, visible-error-driven cleaning, while abstraction-centered interfaces like Jupyter push toward systematic transformations at higher cognitive cost. The authors argue that tool design structures the process and subjective experience of wrangling, and that recognizing these trade-offs can make tool design more intentional.","feed_headline":"Data tools steer cleaning strategy but not output quality","feed_subtitle":"A 40-person study finds trade-offs between spreadsheet, notebook, AI, and wrangler interfaces—and no consistent winner.","key_machinery":"The Technical Dimensions of Programming Systems (TDoPS) framework, a set of seven clusters of dimensions for comparing programming systems—interaction, errors, conceptual structure, notation, complexity, customizability, and adoptability. The study uses it as a coding scheme for deductive thematic analysis of think-aloud and screen-capture data, mapping each tool onto trade-off axes such as visibility versus abstraction, predictability versus simplicity, big versus small steps, and structured reuse versus ad hoc repetition. The framework does the work of turning observed behaviors into comparable design tensions.","core_discovery":"Within the context of this study, tool affordances steer user strategies but do not determine outcomes: no single tool offers a consistent advantage, and results do not converge within tools. The key tension is between data-centered and abstraction-centered interfaces. Data-centered interfaces (Excel, OpenRefine) encourage opportunistic cleaning driven by visible data issues rather than systematic, planned transformations, but they come with a cognitive burden; abstraction-centered interfaces (Jupyter) support systematic cleaning but require more procedural effort. AI interfaces (ChatGPT) reduce reported effort and frustration but obscure intermediate states, inviting automation bias and mis","pith_inferences":["If tool-induced sensemaking is real, a testable extension is that the same cleaning task given to the same user in two paradigms should produce measurably different sequences of operations (e.g., opportunistic versus systematic), which could be verified with interaction logs.","The paper's outcome measure (Jaccard similarity to silver tables) mainly captures row and column structure, not semantic correctness; future work could use multi-set semantic matching to see whether hidden quality differences lie beneath the 'no consistent advantage' claim.","The browser-hosted configurations and permitted external AI use introduce confounds; a replication that isolates the interface paradigm from hosting environment (e.g., desktop Excel versus Excel for the web, hosted Jupyter versus JupyterLite) would strengthen the attribution to paradigm.","The study hints that AI assistance may flatten strategy differences between tools; a direct comparison of the same tools with and without ambient AI help would test whether AI acts as a great equalizer of wrangling process."],"forward_implications":["If tool paradigms shape process but not outcome quality, then evaluating wrangling tools solely by final data quality misses most of what design affects; user experience and strategy should be explicit design targets.","The trade-off between opportunistic and systematic cleaning suggests designers could combine data visibility with structured transformations—for example, spreadsheet-like direct manipulation with notebook-like abstractions—to get the best of both.","The finding that ChatGPT users reported lower frustration but produced no better outputs, and that participants in all conditions turned to AI for help, implies AI is best positioned as an ambient co-pilot inside existing tools rather than a primary interface.","The pattern of tool-induced sensemaking—users redefine their task in terms of what the tool makes easy—implies that adding a prominent feature (e.g., OpenRefine's faceting) can steer attention toward that operation, for better or worse.","Since task constraints narrow the space of viable strategies, tool differences matter most in open-ended wrangling; benchmarking and design guidance should be task-dependent."],"fun_headline_variants":["Data tools steer cleaning strategy, not results","No winning wrangling tool—only trade-offs","Spreadsheets vs. notebooks vs. AI: all trade-offs","AI cleaning feels easy but hides missteps","Wrangling strategy shifts with tool, not quality"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that observed differences in strategy and experience are caused by the interface paradigm of the assigned tool, rather than by the specific hosting configuration (web-based Excel, JupyterLite, VM-based ChatGPT, OpenRefine version) or by participants' external use of AI and other resources.","fun_headline_variants_meta":{"raw":{"variants":["Data tools steer cleaning strategy, not results","No winning wrangling tool—only trade-offs","Spreadsheets vs. notebooks vs. AI: all trade-offs","AI cleaning feels easy but hides missteps","Wrangling strategy shifts with tool, not quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1220,"prompt_tokens":716,"completion_tokens":504,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":431}},"tokens_in":460,"tokens_out":504,"duration_ms":5296,"temperature":1.0,"reasoning_tokens":431,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:29:46.167674+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An experiment that assigns the same participants to the same two paradigms with hosting and versions matched (e.g., desktop Excel versus hosted JupyterLab) and measures both process (operation sequences, planning versus opportunistic actions) and outcome would settle whether paradigm-level differences persist; if operation sequences become indistinguishable across paradigms once hosting and external help are controlled, the central claim fails.","supporting_citations":[],"review_version":1}