{"id":"7bb8e736-3d2b-461a-951e-2eef9b9a82bd","arxiv_id":"2412.15363","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A short educational workshop increased self-reported transparency literacy and advocacy willingness in a small, self-selected sample of professionals.","lead":"The authors built a free two-hour workshop teaching professionals about algorithmic transparency and how to advocate for it, then tested it with 27 people in news media and tech startups. It is an early, openly shared experiment in whether education can create 'transparency advocates' who push organizations to adopt more transparent AI practices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-rated understanding, not objective knowledge, is the only quantitative evidence for the literacy claim; the 'didn't know what I didn't know' effect makes this measure unreliable, so the central claim is not established.","rationale":"The reader's weakest assumption was that participant self-reports are an unbiased measure of workshop effects. I agree with that core concern and sharpen it: the quantitative literacy measure is not merely self-reported but is a self-assessment of understanding, which is known to be sensitive to calibration shifts. The paper's own qualitative finding that participants 'didn't know what they didn't know' illustrates exactly this problem: increased awareness of ignorance can lower self-ratings, while social desirability and demand characteristics can raise them, making the observed 5.00-to-7.79 jump uninterpretable as a measure of learning. The advocacy outcome suffers from the same issue, with no control group and only unrecorded, self-reported actions; P2's own statement that they 'always would've advocated anyway' directly undercuts causal attribution. The paper is honest and provides valuable open materials, but the central claim requires an objective outcome measure or a counterfactual. Since the reader already rated the paper CONDITIONAL, my concern reinforces that verdict rather than moving it.","tokens_in":13267,"tokens_out":3250,"duration_ms":32418,"concrete_test":"Run a waitlist-controlled evaluation of the workshop with a matched control group, replacing the self-rated Q5 with an objective knowledge test (e.g., 20 items drawn directly from the workshop modules: define algorithmic transparency, identify stakeholders, match tools like model cards or dashboards to use cases). Compare pre/post gains on this objective test between treatment and waitlist control, and also compare self-rated vs. objective gains to quantify the bias from self-perception. If objective test gains do not significantly exceed control gains, the literacy claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that the workshop increased algorithmic transparency literacy and willingness to advocate rests on self-report measures that cannot distinguish actual learning from changes in self-perception. The survey's literacy item (Q5) asks participants to rate their own understanding on a 1-10 scale; it is not a knowledge test. This is especially problematic because a key reported outcome is that participants realized they 'didn't know what they didn't know' (P2, P3, P4 in Results, 'Uncovering knowledge gaps'). Such an awareness shift should bias self-ratings downward after the workshop, yet Table 3 shows General Understanding rising from 5.00 to 7.79. The direction of bias is ambiguous, and no control group or objective pre/post measure exists to resolve it. Willingness items (Q7/Q8) are likewise self-reported Likert scales, and the only behavioral evidence is self-reported conversational, implementational, and influential actions from unrecorded interviews. Since P2 explicitly said 'I always would've advocated for transparency anyway,' the observed post-workshop actions cannot be attributed to the workshop without a counterfactual. Thus the claim that 'the workshops were effective in teaching algorithmic transparency and increasing participants' willingness to advocate' is not established by the data as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on a workshop-based educational intervention designed to create \"transparency advocates\" who can drive bottom-up organizational change toward algorithmic transparency. The authors delivered a two-hour workshop, developed over several years and based on their prior stakeholder-first transparency playbook, to two groups of professionals: 15 news/media professionals and 12 technology startup professionals. Data include pre/post workshop surveys (n=15 matched responses) and semi-structured follow-up interviews with 7 participants, which were not recorded but captured via detailed notes. The paper claims that the workshop increased participants' algorithmic transparency literacy and their willingness to advocate for transparency in their professional lives, and it proposes a three-level taxonomy of advocacy (conversational, implementational, influential) plus domain-specific differences in advocacy barriers. The authors openly acknowledge the small, self-selected sample and the absence of statistical testing, and they make all workshop materials publicly available.","tokens_in":13477,"tokens_out":3148,"duration_ms":30849,"significance":"If the central claims were established, this work would make a valuable contribution to responsible AI education and AI governance by offering a concrete, replicable, bottom-up mechanism for translating XAI research into organizational practice. The paper's strengths include a freely available, open-source workshop curriculum, a two-domain comparison, and a honest acknowledgment of key limitations (small sample, self-selection, descriptive statistics only). The proposed taxonomy of advocacy levels is a useful conceptual contribution, and the reported real-world actions (e.g., a participant speaking up at an organization-wide AI strategy meeting) are motivating examples. However, the evidentiary basis for the causal claim that the workshop increased literacy and advocacy willingness is fragile, resting primarily on self-report measures and unrecorded interviews; the paper's significance therefore hinges on how these evidential limitations are treated in the framing of the conclusions.","major_comments":[{"comment":"The claim that the workshop increased algorithmic transparency literacy is not established by the quantitative evidence. The survey's literacy measure (Q5) is a self-rating of understanding on a 1-10 scale, not an objective knowledge test. The observed increase from 5.00 to 7.79 could reflect increased confidence or a shift in self-perception rather than actual knowledge gain. This ambiguity is especially acute because the qualitative results (\"Uncovering knowledge gaps\") indicate that participants realized they \"didn't know what they didn't know,\" which should, if anything, bias post-workshop self-ratings downward; the paper does not address this tension. The authors should either temper the literacy claim to \"perceived understanding\" or supplement the self-report with an objective measure (e.g., a short knowledge quiz) in future iterations, and they should explicitly discuss the direction of self-report bias in the current study.","section":"Methods (Data Collection and Analysis); Table 3"},{"comment":"The attribution of participants' advocacy actions to the workshop is not sufficiently supported. P2 is quoted as saying \"I always would've advocated for transparency anyway,\" and the paper acknowledges that actions \"appeared to be motivated, at least in part, by the workshop.\" This \"at least in part\" rests entirely on participants' own judgments, which are vulnerable to social desirability bias and hindsight bias. Without a counterfactual, a control group, or independent verification of the reported actions, the paper cannot rule out the possibility that participants were already inclined to advocate and would have acted identically without the workshop. The authors should either substantially soften the causal language in the abstract and conclusion, or provide a more rigorous analysis of the specific workshop elements that plausibly changed behavior (e.g., the role-playing activity's direct influence on P4's strategy-meeting intervention).","section":"Results (\"Taking action\"); Discussion (\"Levels of Advocacy\")"},{"comment":"The reliability of the qualitative findings is compromised by the decision not to record interviews. The paper states that \"we chose not to record the interviews\" and instead \"took detailed notes,\" but it does not describe how the notes were verified (e.g., by member checking, participant review, or second-coder agreement). Because the thematic analysis was conducted by the authors on their own notes, there is a risk of selective note-taking and confirmatory interpretation. The paper should either provide evidence of note reliability or explicitly acknowledge this as a threat to the trustworthiness of the central qualitative claims, and it should explain how the 33 codes and 6 themes were audited beyond \"two separate working sessions.\"","section":"Methods (Data Collection and Analysis); Results (Thematic Analysis Findings)"}],"minor_comments":[{"comment":"There is a typo: \"we found vast differences in the attitudes towards algorithmic transparency in new and media vs. technology startups\" should read \"news and media.\"","section":"Discussion (\"The Importance of Domain-of-use\")"},{"comment":"The name \"Myerson\" appears in the text, but the reference is to \"Meyerson (2003)\"; please use the correct spelling consistently throughout.","section":"Introduction; References"},{"comment":"The reference for Covert et al. begins with \"DBLP:journals/corr/abs-2004-00668\" which appears to be a leftover identifier; this should be cleaned up.","section":"References"},{"comment":"The figure is referred to as \"Appendix Figure 2\" in the Methods section but is simply \"Figure 2\" in the appendix; please number it consistently.","section":"Appendix (Figure 2)"},{"comment":"The paper states that the survey was \"adapted from previous work by Lewis and Stoyanovich (2021)\" but does not give details on which items were changed or validated for this context; adding this information would improve replicability.","section":"Methods (Pre- and post-workshop surveys)"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of cs.CY and likely AIES, but the central claims need to be brought in line with the evidence. The authors have been transparent about the small sample and descriptive statistics, which is commendable. My main concern is that the abstract and introductory \"Summary of findings\" make causal statements (\"the workshops were effective\") that the data cannot support, even though the body of the paper is more nuanced. If the authors substantially soften the causal framing and reframe the contribution as an exploratory pilot study with a detailed description of the intervention and its perceived impact, the paper could be suitable after revision. I would encourage the editors to ask the authors to address the self-report bias issue explicitly, since the \"didn't know what I didn't know\" effect cuts directly against the direction of the observed change in self-rated understanding."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a well-intentioned, open, small-scale pilot study, but the headline claim—that the workshop teaches algorithmic transparency and boosts advocacy—is only weakly supported by the evidence. The stress-test note is on point: the key quantitative measure is self-rated understanding, not objective knowledge.\n\nWhat's genuinely useful: the three-level taxonomy of advocacy (conversational, implementational, influential) is a clean way to categorize bottom-up change and will likely be cited. The two-domain comparison, while rough, raises an interesting hypothesis about domain-specific barriers. And the authors deserve credit for publishing their workshop materials and giving a concrete example of a participant speaking up at an AI strategy meeting.\n\nThe soft spots are the same ones the authors acknowledge, but the framing in the abstract is still too strong. The sample is 27 self-selected professionals, 15 matched surveys, 7 unrecorded interviews. The survey uses self-reports of understanding and willingness. The 'didn't know what I didn't know' effect complicates the interpretation of the increase in self-rated understanding: participants may be more confident without being more knowledgeable, or more aware of gaps but still rating themselves higher for other reasons. No control group means you can't attribute the reported actions to the workshop—P2 explicitly said they would have advocated anyway. The domain comparison is also confounded by delivery format and recruitment differences.\n\nFor all that, this is an honest pilot. The conclusion has the right caveats; the problem is that the abstract and results section don't carry them. With a serious referee pushing for softened language, or a follow-up with an objective knowledge test and a waitlist control, this would be a useful contribution.\n\nMy take: worth peer review, but as a practice/experience paper, not as strong empirical evidence. People working on AI literacy or responsible AI education will find the taxonomy and open materials worth their time.\n\nRecommendation: send it out to review. The authors have done the field a service by making everything open; a referee can help them align the claims with the data.","headline":"Honest small pilot on teaching transparency advocacy; useful taxonomy, but effectiveness claim overstates self-report evidence.","tokens_in":13950,"tokens_out":2545,"would_cite":true,"duration_ms":23503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-hour educational workshop can turn professionals into algorithmic transparency advocates, with participants reporting real advocacy actions such as speaking up at an AI strategy meeting.","keywords":["algorithmic transparency","Explainable AI","responsible AI education","transparency advocacy","tempered radicals","AI literacy","organizational change"],"falsifier":"A randomized evaluation in which employees within the same organizations are assigned to the workshop or to a placebo training, with advocacy actions measured from verifiable records such as meeting minutes, adoption of model cards or datasheets, emails, or project artifacts, would settle whether the workshop causes advocacy rather than merely accompanying it.","tokens_in":13070,"feed_emoji":"🎓","tokens_out":4273,"duration_ms":37753,"temperature":0.7,"pith_summary":"This paper tests whether a short educational workshop can create \"transparency advocates\"—employees who push their organizations toward greater algorithmic transparency from the ground up. Across two workshops with 27 professionals in news and media and technology startups, participants reported higher understanding of algorithmic transparency and greater willingness to advocate for it. Four participants described taking real advocacy actions in the days after the workshop, including speaking up at an organization-wide AI strategy meeting. The paper argues that advocacy is not a single behavior but occurs at conversational, implementational, and influential levels, and that domain context shapes both willingness and ability to advocate.","feed_headline":"A 2-hour workshop can create algorithmic transparency advocates","feed_subtitle":"Field interviews and surveys show participants speaking up at AI strategy meetings and changing workflows.","key_machinery":"The load-bearing mechanism is the two-hour workshop, built around five modules covering transparency definitions, tools (model cards, datasheets, explainer dashboards, Shapley values), the stakeholder-first Transparency Playbook, a role-playing breakout activity where participants argue for and against disclosure at a fictional startup, and common objections to transparency with rebuttals. The role-play is central because it rehearses the tensions participants will meet as advocates and equips them with counterarguments. A pre/post survey adapted from prior responsible data-science teaching measures six constructs, and semi-structured interviews capture reported actions.","core_discovery":"The paper claims that education can translate algorithmic transparency research into practice by equipping motivated individuals inside organizations with knowledge, tools, and argumentative strategies to advocate for transparency. The authors delivered a two-hour, open-source workshop built on a stakeholder-first transparency playbook, and their qualitative interviews and pre/post surveys indicate gains in both transparency literacy and willingness to advocate. The most notable reported outcome is a participant who raised transparency concerns at an organization-wide AI strategy meeting days after the workshop, directly applying the workshop's lessons. The paper also claims that advocacy has three levels—conversational, implementational, and influential—and that news and media professionals tend to be more willing but less equipped to act, while startup professionals have more technical means but less organizational room to prioritize transparency.","pith_inferences":["If the effects replicate with larger samples, workshop-based advocate cultivation could complement regulation by creating internal demand for transparency that outlasts individual workshops.","A natural extension is a longitudinal study tracking whether reported advocacy translates into auditable organizational artifacts, such as published model cards or disclosure policies, months after training.","The level taxonomy suggests a testable prediction: training that matches advocacy level to role, such as engineers to implementational actions and managers to influential ones, will produce more sustained change than generic training.","Domain differences imply that startup-focused interventions may need to bundle transparency with business value, such as risk reduction or investor due diligence, rather than relying on ethical appeals alone."],"forward_implications":["Organizations can cultivate bottom-up pressure for algorithmic transparency through short, low-cost workshops, even in the absence of strong regulation.","Advocacy is multi-level: education should prepare people for conversational, implementational, and influential actions, since each requires different skills and authority.","Professional domain shapes advocacy, so training and support must be tailored: news and media professionals need tools, while startup professionals need resources and prioritization.","A single advocate can bring transparency onto a high-stakes organizational agenda, and the reported AI strategy meeting episode suggests a possible ripple effect on company practices.","Open-source workshop materials make the intervention replicable, allowing other organizations and educators to test and adapt it."],"supporting_citations":[{"why":"Supplies the \"tempered radicals\" concept that transparency advocates are modeled on.","marker":"(Meyerson 2003)"},{"why":"Provides the stakeholder-first Transparency Playbook that forms the workshop's core content.","marker":"(Bell, Nov, and Stoyanovich 2023)"},{"why":"Defines the stakeholder-first responsible AI literacy approach the workshop builds on.","marker":"(Domínguez Figaredo and Stoyanovich 2023)"},{"why":"The pre/post survey instrument is adapted from this prior responsible data-science course evaluation.","marker":"(Lewis and Stoyanovich 2021)"},{"why":"Supplies the six-stage thematic analysis method used to code interview notes.","marker":"(Braun and Clarke 2006)"},{"why":"Documents organizational barriers such as profit incentives and diffuse ethics ownership, motivating the need for advocates.","marker":"(Metcalf, Moss et al. 2019)"},{"why":"Shows practitioners already feel responsibility for ethical values, supporting the feasibility of bottom-up advocacy.","marker":"(Rakova et al. 2021)"}],"fun_headline_variants":["Workshop turns employees into AI transparency advocates","Two-hour workshop turns professionals into transparency advocates","Workshop empowers staff to advocate for AI transparency","Education builds transparency advocates from the ground up"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions rest on participants' own unverified reports, in surveys and unrecorded interviews, that their understanding and advocacy increased, without a control group or external confirmation of the reported actions.","fun_headline_variants_meta":{"raw":{"variants":["Workshop turns employees into AI transparency advocates","Two-hour workshop turns professionals into transparency advocates","Workshop empowers staff to advocate for AI transparency","Education builds transparency advocates from the ground up"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000538,"raw_usage":{"total_tokens":2559,"prompt_tokens":896,"completion_tokens":1663,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1607}},"tokens_in":512,"tokens_out":1663,"duration_ms":12295,"temperature":1.0,"reasoning_tokens":1607,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:28:44.279585+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized evaluation in which employees within the same organizations are assigned to the workshop or to a placebo training, with advocacy actions measured from verifiable records such as meeting minutes, adoption of model cards or datasheets, emails, or project artifacts, would settle whether the workshop causes advocacy rather than merely accompanying it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the \"tempered radicals\" concept that transparency advocates are modeled on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the stakeholder-first Transparency Playbook that forms the workshop's core content."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the stakeholder-first responsible AI literacy approach the workshop builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The pre/post survey instrument is adapted from this prior responsible data-science course evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents organizational barriers such as profit incentives and diffuse ethics ownership, motivating the need for advocates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows practitioners already feel responsibility for ethical values, supporting the feasibility of bottom-up advocacy."}],"review_version":1}