{"id":"2e6cac95-d524-4f79-840a-a2de1958b10e","arxiv_id":"2605.21363","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"CoTrace traces AI goal-shaping contributions in 638 logs showing 11-26% direct involvement with higher indirect and lower-level effects, while user exposure to the analyses shifts perceived contributions by nearly 2 points on a 5-point scale.","lead":"This paper introduces CoTrace, a framework that decomposes explicit goals into verifiable requirements and traces direct and indirect AI contributions across dialogue turns in human-AI collaborations. Smart generalists should read it because it reveals how users often misjudge their own role versus the AI's in shaping goals, which affects reliance calibration and work evaluation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"CoTrace attribution percentages rest on unvalidated manual decomposition of goals into requirements, risking systematic omission of subtle influences or coder bias.","rationale":"The reader's weakest assumption matches the load-bearing step exactly. Because the full text is now available yet still centers on human-coded decomposition without reported reliability metrics, the concern remains the primary threat to the quantitative claim. The proposed test directly measures whether the attribution process is reproducible enough to support the reported percentages; low agreement would require the claim to be qualified or the framework to be augmented with automated checks.","tokens_in":1700,"tokens_out":362,"duration_ms":23489,"concrete_test":"Select a stratified random sample of 50 logs; have two independent annotators apply the CoTrace decomposition and tracing procedure; compute Cohen's kappa on both requirement identification and contribution attribution labels. If kappa < 0.65 on either, re-run the full 638-log analysis with reconciled labels and report the change in the 11-26% range.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim (models contribute 11-26% to goal-shaping) is produced by applying CoTrace to 638 logs: explicit goals are decomposed into verifiable requirements, then each requirement is traced to direct or indirect contributions across turns. For the 11-26% figure (and the contrast with lower-level requirements) to be reliable, this decomposition must be both exhaustive and free of systematic coder bias or missed indirect effects. The abstract provides no inter-annotator agreement statistics, no validation against an external gold standard, and no sensitivity analysis on how alternative decompositions would alter the percentages. If coders consistently under-attribute subtle model influences or over-attribute user intent, the headline contribution numbers shift.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces CoTrace, a goal-level attribution framework that decomposes explicit goals into verifiable requirements and traces both direct contributions and indirect influences across dialogue turns in human-AI collaboration. Applied to 638 real-world logs, it reports that models account for 11-26% of goal-shaping contribution while contributing more to lower-level concrete requirements and various indirect effects. Controlled simulations demonstrate that interaction design choices affect model goal-shaping behavior. A user study finds that exposing participants to goal-level analyses shifts their perceived contributions by nearly 2 points on a 5-point scale, indicating miscalibration in users' understanding of AI-assisted work.","tokens_in":1872,"tokens_out":596,"duration_ms":29848,"significance":"If the CoTrace attribution method proves reliable, the work could meaningfully advance understanding of how LLMs shape goals in collaboration, with implications for user calibration of reliance and for designing better human-AI interfaces. The use of real collaboration logs, simulations testing design choices, and a user study on perception shifts provides a multi-method approach that strengthens potential impact in HCI and AI ethics. The focus on process-level tracing rather than final artifacts is a clear strength.","major_comments":[{"comment":"Methods section on CoTrace: The central quantitative claims (11-26% goal-shaping contribution and contrasts with lower-level requirements) rest on manual decomposition of goals into verifiable requirements followed by tracing across turns, yet no inter-annotator agreement statistics, validation against an external gold standard, or sensitivity analysis to alternative decompositions are reported. This directly affects reliability of the attribution percentages.","section":"Methods (CoTrace application)"},{"comment":"Results on 638 logs: The contribution figures and data exclusion rules are presented without confidence intervals, details on coder training, or robustness checks, making it impossible to rule out post-hoc choices or systematic coder bias that could shift the headline 11-26% range.","section":"Empirical Analysis / Results"},{"comment":"User study: The nearly 2-point shift on the 5-point scale is reported without sample size, exact statistical test, p-value, or pre-registration details, which is load-bearing for the claim of systematic miscalibration in perceived contributions.","section":"User Study"}],"minor_comments":[{"comment":"Abstract: Replace the vague 'nearly 2 points' with the precise mean difference and scale anchors for clarity.","section":"Abstract"},{"comment":"Figure captions: Ensure diagrams distinguishing direct vs. indirect contributions include explicit legends and examples from the logs.","section":"Figures"},{"comment":"Related work: Add citations to prior work on contribution attribution in collaborative writing or goal-setting systems to better situate CoTrace.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits well within HCI/AI collaboration venues; however, the lack of validation details on the core measurement raises reproducibility concerns that should be addressed before acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback, which identifies key areas where additional reporting and validation will improve the transparency and robustness of our claims. We address each major comment below and commit to revisions that directly respond to the concerns raised.","responses":[{"response":"We agree that these details are important for establishing reliability. In the revised manuscript we will add inter-annotator agreement statistics (Cohen's kappa) for both the goal decomposition and tracing steps, along with a sensitivity analysis that applies alternative decompositions and reports the resulting range of contribution percentages. We will also explicitly discuss the absence of an external gold standard, noting the interpretive character of goal decomposition, and describe our annotation protocol and training procedures in greater detail.","revision_made":"yes","referee_comment":"[Methods (CoTrace application)] Methods section on CoTrace: The central quantitative claims (11-26% goal-shaping contribution and contrasts with lower-level requirements) rest on manual decomposition of goals into verifiable requirements followed by tracing across turns, yet no inter-annotator agreement statistics, validation against an external gold standard, or sensitivity analysis to alternative decompositions are reported. This directly affects reliability of the attribution percentages."},{"response":"We will revise the Results section to include 95% confidence intervals for all reported contribution percentages. We will also add a dedicated subsection describing coder training, qualification criteria, and the exclusion rules applied to the 638 logs. Finally, we will present robustness checks that vary the exclusion criteria and re-compute the 11-26% range to demonstrate that the headline findings are not sensitive to these choices.","revision_made":"yes","referee_comment":"[Empirical Analysis / Results] Results on 638 logs: The contribution figures and data exclusion rules are presented without confidence intervals, details on coder training, or robustness checks, making it impossible to rule out post-hoc choices or systematic coder bias that could shift the headline 11-26% range."},{"response":"We will expand the User Study section to report the exact sample size, the statistical test used (paired t-test), the associated p-value, and the effect size. We will also clarify the pre-registration status: the analysis plan was fixed prior to data collection even though formal pre-registration was not completed; we will note this limitation transparently while providing the planned analysis details that support the reported perception shift.","revision_made":"yes","referee_comment":"[User Study] User study: The nearly 2-point shift on the 5-point scale is reported without sample size, exact statistical test, p-value, or pre-registration details, which is load-bearing for the claim of systematic miscalibration in perceived contributions."}],"tokens_in":1457,"tokens_out":585,"duration_ms":29078,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is that the paper supplies a concrete framework called CoTrace for breaking explicit goals into verifiable requirements and then tracing direct and indirect contributions across dialogue turns. They apply it to 638 real logs and report that models account for 11-26% of goal-shaping while doing more on lower-level concrete requirements, plus indirect effects. The user study shows that feeding people the goal-level breakdown moves their self-ratings by nearly two points on a five-point scale, which points to a real calibration issue in how users see their own work with AI.","headline":"CoTrace gives a workable process for tracing goal contributions in human-AI logs, with numbers showing modest model influence at the high level and a perception shift in the user study, but the manual decomposition step is the part that needs more checks.","tokens_in":2357,"tokens_out":206,"would_cite":false,"duration_ms":26835,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We introduce a goal-level attribution framework, CoTrace, that decomposes explicit goals into verifiable requirements and traces both direct contributions and indirect influences across dialogue turns."},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"role-level contribution of speaker p to requirement r through role ρ is M(p,ρ,r) = ∑_{a∈A_p} 1[role(a)=ρ] I(a→r)"}],"headline":"CoTrace goal-requirement decomposition and influence tracing operates in empirical HCI attribution; no overlap with RS cost-forcing or distinction-derived structure.","alignment":"orthogonal","rationale":"The paper's core machinery (decomposing goals into verifiable requirements, labeling direct/indirect influence via LLM judges, aggregating speaker×role contribution matrices M(p,ρ,r)) is a process-tracing pipeline for dialogue logs. RS derives J(x)=½(x+x⁻¹)−1, φ, 8-tick periodicity, D=3, and constants from a single distinction via functional equations and absolute-floor closure. No shared primitives, cost functions, ratio symmetry, or periodicity appear; the domains (collaboration measurement vs. parameter-free physics emergence) are disjoint.","tokens_in":57340,"confidence":"high","tokens_out":350,"duration_ms":12083,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A framework called CoTrace traces AI contributions to goal formation in human collaborations and finds models shape only 11-26 percent of high-level goals.","keywords":["human-AI collaboration","goal attribution","LLM contributions","interaction design","user perception","contribution tracing","dialogue analysis"],"falsifier":"Applying the same decomposition and tracing process to a fresh set of collaboration logs and obtaining goal-shaping percentages outside the 11-26% range, or finding no shift in user perceptions after exposure to the analysis, would challenge the central results.","tokens_in":2607,"feed_emoji":"🤖","tokens_out":627,"duration_ms":37861,"temperature":0.7,"pith_summary":"The paper introduces CoTrace to break down explicit goals into verifiable requirements and track both direct inputs and indirect influences across conversation turns. Analysis of 638 real collaboration logs shows AI plays a smaller role in shaping overall goals but adds more concrete lower-level requirements and various indirect effects. Controlled tests reveal that changes in interaction design alter these patterns, while a user study finds that presenting the goal-level breakdown shifts how much credit participants assign to themselves or the AI by nearly two points on a five-point scale.","feed_headline":"AI shapes 11-26% of goals in human collaborations","feed_subtitle":"New tracing method finds larger role in concrete requirements and indirect effects plus shifts in how users assign credit","key_machinery":"CoTrace, the goal-level attribution framework that decomposes explicit goals into verifiable requirements and traces direct and indirect contributions across dialogue turns.","core_discovery":"CoTrace decomposes explicit goals into verifiable requirements and traces both direct contributions and indirect influences across dialogue turns. When applied to 638 real-world collaboration logs, models account for 11-26% of goal-shaping contribution yet contribute substantially more to introducing lower-level concrete requirements and various indirect influences. Interaction design choices affect model goal-shaping behavior in controlled simulations, and exposing users to the resulting analyses shifts their perceived contributions by nearly 2 points on a 5-point scale.","pith_inferences":["Similar tracing methods could help users in other domains calibrate how much they rely on AI during planning or creative tasks.","The approach might be adapted to evaluate contribution patterns in group settings without AI, such as team project logs.","Designers could embed lightweight versions of the decomposition step into real-time collaboration interfaces to surface contributions as work unfolds."],"forward_implications":["Interaction design choices can be tuned to increase or decrease the extent of model goal-shaping.","Users hold systematically miscalibrated views of their own contributions in AI-assisted work.","Providing goal-level attribution data can correct those miscalibrations by nearly two points on a five-point scale."],"fun_headline_variants":["CoTrace finds AI at 11-26% of goal-level contributions","Models contribute more to concrete requirements in logs","Goal analysis exposure shifts user ratings by 2 points","Interaction choices affect AI goal shaping behavior"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Decomposing explicit goals into verifiable requirements captures contributions completely and without systematic bias or omission of subtle influences.","fun_headline_variants_meta":{"raw":{"variants":["CoTrace finds AI at 11-26% of goal-level contributions","Models contribute more to concrete requirements in logs","Goal analysis exposure shifts user ratings by 2 points","Interaction choices affect AI goal shaping behavior"]},"model":"grok-4.3","cost_usd":0.00905,"raw_usage":{"total_tokens":3974,"prompt_tokens":654,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":90503000,"prompt_tokens_details":{"text_tokens":654,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3260,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":654,"tokens_out":60,"duration_ms":51480,"temperature":1.0,"reasoning_tokens":3260,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T04:40:04.206257+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Applying the same decomposition and tracing process to a fresh set of collaboration logs and obtaining goal-shaping percentages outside the 11-26% range, or finding no shift in user perceptions after exposure to the analysis, would challenge the central results.","supporting_citations":[],"review_version":1}