{"id":"4e9b4034-f96b-4fea-a9cd-c43c1e533af1","arxiv_id":"2412.16866","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A formative design study shows how trauma-informed computing could reshape qualitative coding tools, and argues for safety-as-enablement.","lead":"This paper applies trauma-informed computing principles to the design of qualitative coding software, proposing a prototype called TIQA that tracks and nudges researchers about their exposure to upsetting content. In interviews, 15 researchers imagined how such tools could support self-care, and the authors argue safety should be reframed as enabling users rather than just reducing exposure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unvalidated ExposureModeler proxy (counts + time) is the weakest link: if it does not track felt distress, TIQA's core self-management loop and the 'self-reflective surface' role for trauma are built on sand, and no calibration evidence is reported.","rationale":"The reader correctly identified the exposure-metric validity as the weakest assumption. I agree, and I do not see a more load-bearing concern. The paper is a design provocation, so small sample size, lack of released code, and lack of efficacy testing are acceptable for the stated exploratory contribution; the authors are also transparent about these limitations. The exposure proxy, however, is load-bearing because it is the one place where the empirical claim (machine assistance can help mitigate trauma) rests on a quantitative measurement that is never checked. Participants themselves flagged the issue, and the paper does not refute it—it only reports it. A neutral reader should therefore treat the 'self-reflective surface for traumatic responses' role as an interpretation offered by participants about a hypothetical feature, not as evidence that the feature works. This does not change the verdict: the contribution is a design space exploration, and the concern is acknowledged and does not invalidate the conceptual reframing. The paper should probably tighten wording in future revision to distinguish 'a system that may measure exposure' from 'a system that measures exposure,' but for an exploratory provocation, ACCEPT remains appropriate. My concrete test would settle the validity question and make the claim much stronger if it passes.","tokens_in":31788,"tokens_out":6047,"duration_ms":58738,"concrete_test":"Run a within-subjects validation study in which 20–30 analysts code segments from the same synthetic corpora while periodically rating their subjective distress (e.g., SAM or a single-item 'how distressing right now' scale); after the session, compute the correlation between the cumulative number of ExposureModeler-matched segments/time and the cumulative self-reported distress, and compare with a 'graphicness' rating of the segments. If the count-based measure explains less than 20% of the variance in self-reported distress (e.g., R² < 0.2) and graphicness ratings do substantially better, the ExposureModeler proxy is not valid for its purpose, and the paper's claims about self-managing exposure should be explicitly relabeled as speculative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central mechanism for self-managing traumatic exposure is ExposureModeler (Table 1), which measures 'prior exposure' as the number of times a code was used plus session time, and 'upcoming exposure' as the count of semantically matched segments. This assumes count-of-occurrence is a valid proxy for traumatic impact. Participants directly challenged this assumption: P05 (Section 4.2.2) stated that impact depends on graphicness, not quantity of instances, and that a single graphic description could be more impactful than repeated mentions of 'assault.' The paper reports no calibration or validation of the metric against any self-report or physiological measure of distress. If the proxy is invalid, the feedback loop in Figure 3 (measure exposures → predict reactions → nudge breaks) is not just noisy but potentially misleading: an analyst may be nudged away from work with low graphic but repeated mentions, or not warned before a single graphic segment. The derived role 'self-reflective surface for an analyst's traumatic responses' (Table 3) is consequently contingent on an unvalidated measurement. The safety-as-enablement reframing (Section 5.1) is conceptual and would survive, but the empirical contribution about machine assistance in trauma mitigation is weakened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper explores how trauma-informed computing (TIC) principles can be operationalized in software for qualitative coding, presenting a prototype system called TIQA that combines user-defined code models, semantic search, and exposure tracking. Through a formative study with 15 researchers who used TIQA on synthetic intimate-partner-violence and social-media datasets, the authors identify potential roles for machine assistance: a self-reflective surface for analysts, an initiator of collaboration and a culture of care, and a fair allocator of team responsibility, while participants rejected the role of enforcer. The paper also advances a conceptual reframing, safety-as-enablement, and argues for evaluating the trauma-informedness of design processes.","tokens_in":31985,"tokens_out":10309,"duration_ms":93066,"significance":"This is a useful and timely contribution at the intersection of CSCW, HCI, and trauma-informed computing. The paper provides a concrete design exploration (TIQA) that demonstrates how high-level TIC principles can be translated into lower-level design decisions, and it grounds the proposed roles in participant feedback from a scenario-based study with an appropriately experienced sample. The safety-as-enablement reframing is a valuable conceptual contribution that extends prior TIC scholarship and offers design guidance beyond qualitative coding. The authors are careful to frame the study as formative and to acknowledge limitations such as synthetic data, short sessions, and an unvalidated exposure metric. The inclusion of participant critiques (e.g., P05's challenge to the count-based metric) strengthens the paper's credibility and provides a basis for future work.","major_comments":[],"minor_comments":[{"comment":"Section 3.3 states that the ReactionPredicter module was not implemented in full, yet Section 4.1 says participants were asked to 'explore the reaction predictions and nudges.' The interview protocol in Appendices A and C asks participants to imagine the feature, so please explicitly state in Section 4.1 that the prediction/nudge functionality was described as a concept rather than presented as a functional component, to avoid ambiguity about what participants actually encountered.","section":"Section 3.3 and Section 4.1"},{"comment":"The ExposureModeler proxy (count of code occurrences plus session time) is not validated against any self-report or physiological measure of distress. While the paper acknowledges this concern in Section 4.2.2 and Section 5.3, the abstract and Section 3.2.2 use the term 'measure' without qualification. Please add an explicit statement in the Limitations that this metric is an unvalidated design artifact and that the empirical findings reflect participants' perceptions of a design provocation rather than evidence of effective measurement.","section":"Section 3.3 and Table 1"},{"comment":"Interview protocol question 3 describes the exposure metric as a function of (a) time spent, (b) number of annotations, and (c) 'how difficult this concept was for you to read,' but the implementation described in Table 1 and Section 3.3 only includes time and count. Please reconcile the protocol with the actual implementation, or clarify that (c) was part of the interview script to elicit discussion rather than a system feature.","section":"Appendix A.3 and C.3"},{"comment":"The row for 'Measurement of user's prior and upcoming traumatic exposures' lists 'ML-assisted content warnings are a value-add' as participant feedback, but the corresponding quote from P06 refers to the general idea of using ML for trauma tracking rather than content warnings. Consider aligning the table summary more closely with the quoted participant statements.","section":"Table 3"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid formative design study with a clear conceptual contribution. The main concerns are transparency issues around the unimplemented ReactionPredicter and the unvalidated exposure metric; these can be addressed with clarifications. The paper fits well within the scope of CSCW/HCI venues interested in trauma-informed design. No concerns about citation practices or novelty were identified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read for you on arXiv:2412.16866. The paper does what it says: it takes trauma-informed computing principles and actually builds a prototype (TIQA) for qualitative coding, then runs a 15-person provocation study with people who code sensitive content. The genuinely new bit is the conceptual reframing of safety as enablement rather than exposure reduction. That shift is useful beyond this specific tool, and the paper argues for it with a decent table of tradeoffs (Table 4). The empirical work is careful: two realistic scenarios, synthetic data written by someone with domain experience, think-aloud protocol, and an honest positionality statement. The authors don't oversell—they call it formative and exploratory.\n\nThe weaknesses are real but mostly acknowledged. The ExposureModeler measures prior and upcoming exposure as code counts plus session time. That is a crude proxy, and the participants themselves called it out (P05's point about graphicness vs. quantity is spot on). The ReactionPredicter module wasn't implemented. There's no released code or data, though the appendices include the interview protocols and the synthetic corpora, which is something. None of this is hidden; Section 5.3 lists limitations.\n\nThe stress-test note worries that the self-reflective surface role is built on sand if the proxy is invalid. I think that overstates it. The paper's main contribution is the roles and the reframing, which are drawn from participants' reactions to the design provocation, not from a validated measurement. Even if the metric is noisy, the study still shows that analysts can imagine using such a tool for self-reflection, and the safety-as-enablement argument doesn't depend on the metric being accurate. That said, the exposure tracking feature is the least convincing part of the prototype, and a careful revision should either validate the proxy in a longer study or present it more explicitly as a placeholder.\n\nOverall: this deserves a serious referee. It's a solid contribution to the CSCW/HCI literature on mixed-initiative qualitative analysis and trauma-informed design. The citation pattern is appropriate, building on Chen et al. and the TIC literature without overclaiming. I'd recommend acceptance with revisions, mainly around the exposure metric and some tightening of the claims about machine assistance.","headline":"Solid formative design study; the safety-as-enablement reframing is the real contribution, and the unvalidated exposure metric is a weakness but not a fatal one.","tokens_in":32555,"tokens_out":1847,"would_cite":true,"duration_ms":44786,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that trauma-informed computing can be translated into concrete design decisions for qualitative coding software, producing a system that lets analysts define their own traumatic concepts, track their exposure, and manage…","keywords":["trauma-informed computing","qualitative coding","mixed-initiative systems","researcher well-being","semantic search","code embeddings","safety-as-enablement","design science research"],"falsifier":"A longitudinal comparison of TIQA's exposure counts and nudge predictions against analysts' own break decisions and self-reported distress would settle the proxy question: if a single highly graphic passage produces as strong a reaction as dozens of mild mentions, or if count-based predictions do not anticipate when analysts actually take breaks, the system's core feedback loop is not measuring what it claims to measure.","tokens_in":31547,"feed_emoji":"🪞","tokens_out":6961,"duration_ms":59911,"temperature":0.7,"pith_summary":"This paper tries to establish that the principles of trauma-informed computing can be translated into concrete software decisions for qualitative research tools, and that the resulting tool can reduce researchers' traumatic exposure without taking control away from them. It presents TIQA, a qualitative coding system in which an analyst defines the concepts that disturb them, a semantic-search model finds and counts similar passages, and the analyst can track prior and upcoming exposure and signal when they need a break. In interviews, 15 researchers who study intimate partner violence or online hate used TIQA on synthetic datasets; their reactions support the view that machine assistance should act as a self-reflective surface and an initiator of a culture of care, not as an enforcer of efficient coding or mandatory breaks. The paper also argues for a conceptual shift from 'safety as exposure reduction' to 'safety as enablement,' and for evaluating the trauma-informedness of design processes rather than only measuring user outcomes.","feed_headline":"Trauma-aware coding tool wins analysts as a mirror, not an enforcer","feed_subtitle":"A 15-analyst study finds machine help works best as a self-reflective surface and a prompt for team care.","key_machinery":"The carrying object is TIQA's personalized code-embedding loop, which is reused for both annotation and trauma tracking. A code is modeled as the average sentence-embedding (a vector encoding a text segment's meaning) of the passages the analyst has annotated with it; SemanticSearch returns segments whose cosine similarity to that embedding exceeds a user-set threshold; and ExposureModeler uses those matches to estimate prior exposure (count of annotated matches plus session time) and upcoming exposure (predicted matches in the remaining document). A ReactionPredicter module—not fully implemented in the formative study—would train a classifier on those three features plus the analyst's explicit 'I'm taking a break' signals to issue future self-care nudges. The six trauma-informed computing principles (safety, trust, enablement, peer support, collaboration, intersectionality) serve as the design lens, and the paper's conceptual pivot is the interpretation of safety as enablement: giving users the tools to manage their own experience rather than automatically hiding content.","core_discovery":"The central claim is that a qualitative coding tool can be deliberately designed from trauma-informed computing principles to let each analyst define their own personally traumatic concepts, and that this design changes what machine assistance is for. TIQA operationalizes the definition by treating a user-defined code as an embedding of the passages annotated with it, using semantic search to suggest other matching passages, and then reusing those matches to measure how many traumatic instances the analyst has already seen and how many remain. The authors' formative study with 15 researchers indicates that analysts welcome this as a value-add precisely because it does not automate the core work of interpretation: they imagined the tool as a self-reflective surface for understanding their own coding and stress reactions, an initiator of peer support and team conversations about care, and a fair allocator of documents—provided it stays accountable to human supervisors and protects individual privacy. The paper further claims that safety in such systems is better understood as enablement than as shielding, and that this reframing resolves tensions between the safety principle and the other trauma-informed computing principles.","pith_inferences":["Going beyond the paper, the count-based exposure model might be improved by weighting matches with a graphicness or intensity rating derived from the analyst's own annotations, and then testing whether weighted counts predict self-reported distress better than raw counts.","Going beyond the paper, the safety-as-enablement framing suggests a general design pattern for content moderation and journalism tools: instead of automatically obscuring content, systems could offer users configurable warnings, privacy-preserving aggregates of their own exposure, and optional break prompts.","Going beyond the paper, sharing a 'trusted peer's code embedding' could be the seed of a privacy-preserving sensitivity-sharing protocol, in which teams exchange exposure models without exposing the individual passages or reactions that formed them."],"forward_implications":["Qualitative coding software can be built so that warnings are personal: each analyst defines their own traumatic concepts, and the system learns them from the analyst's own annotations.","Researchers will accept machine assistance for trauma mitigation when it supports self-reflection and collaboration, but will reject it if it acts as a productivity enforcer or mandates breaks.","Exposure measurements can inform self-care planning (e.g., scheduling heavy analysis before a recovery activity) and team workload allocation, provided privacy safeguards prevent supervisors from profiling individuals.","Nudges toward self-care are most plausible as team-culture initiators, not as hard constraints, because analysts expect they would otherwise bypass or ignore them.","Trauma-informed design processes should document tradeoffs among safety, trust, enablement, peer support, collaboration, and intersectionality, and be evaluated for their trauma-informedness even when outcome measures of trauma remain unsettled."],"supporting_citations":[{"why":"Supplies the trauma-informed computing framework and its six principles that the design operationalizes.","marker":"[23]"},{"why":"Demonstrates retrospective TIC design analysis that this work extends to a prospective design inquiry.","marker":"[86]"},{"why":"Provides the design science research process (problem scoping to evaluation) used to structure the inquiry.","marker":"[57]"},{"why":"Supplies the sentence-embedding technique used to model codes and find semantically similar segments.","marker":"[66]"},{"why":"Defines trauma and traumatic exposure, grounding the problem the system addresses.","marker":"[48]"},{"why":"Defines vicarious and secondary trauma from hearing others' trauma, the core harm TIQA targets.","marker":"[17]"},{"why":"Surveys content-warning efficacy debates and trauma-informed social media, motivating self-management over automatic filtering.","marker":"[71]"},{"why":"Frames safety through feminist design and the notion of a 'trust pause,' underpinning the safety-as-enablement argument.","marker":"[76]"},{"why":"Defines human-machine complementarity, which frames the desired versus undesired roles for TIQA.","marker":"[85]"}],"fun_headline_variants":["Qual coding tool reframes trauma safety as enablement, not shielding","15 analysts test TIQA: machine help is a mirror for self-reflection","Trauma-informed coding: let analysts define their own triggers","TIQA: measuring analyst trauma exposure in qualitative coding","From shielding to enablement: trauma-aware tool for qualitative coders"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design assumes that an analyst's traumatic exposure can be approximated by counting how many text segments match a user-defined traumatic code (plus session length), rather than by how graphic or personally significant any single passage is.","fun_headline_variants_meta":{"raw":{"variants":["Qual coding tool reframes trauma safety as enablement, not shielding","15 analysts test TIQA: machine help is a mirror for self-reflection","Trauma-informed coding: let analysts define their own triggers","TIQA: measuring analyst trauma exposure in qualitative coding","From shielding to enablement: trauma-aware tool for qualitative coders"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001174,"raw_usage":{"total_tokens":4870,"prompt_tokens":975,"completion_tokens":3895,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":3808}},"tokens_in":591,"tokens_out":3895,"duration_ms":25646,"temperature":1.0,"reasoning_tokens":3808,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T06:00:52.158636+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal comparison of TIQA's exposure counts and nudge predictions against analysts' own break decisions and self-reported distress would settle the proxy question: if a single highly graphic passage produces as strong a reaction as dozens of mild mentions, or if count-based predictions do not anticipate when analysts actually take breaks, the system's core feedback loop is not measuring what it claims to measure.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the trauma-informed computing framework and its six principles that the design operationalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates retrospective TIC design analysis that this work extends to a prospective design inquiry."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the design science research process (problem scoping to evaluation) used to structure the inquiry."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines trauma and traumatic exposure, grounding the problem the system addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines vicarious and secondary trauma from hearing others' trauma, the core harm TIQA targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Surveys content-warning efficacy debates and trauma-informed social media, motivating self-management over automatic filtering."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames safety through feminist design and the notion of a 'trust pause,' underpinning the safety-as-enablement argument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines human-machine complementarity, which frames the desired versus undesired roles for TIQA."}],"review_version":1}