{"id":"a891348f-20cf-4c66-91ab-cd3784afd2a7","arxiv_id":"2505.04291","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An analysis of the AI Incident Database shows the dataset is better understood as a record of societal accountability and response patterns after AI harms than as a tool for preventing technical failures.","lead":"This paper analyzes 962 incidents from the AI Incident Database and argues the database's real value is documenting how society responds to AI harms, not preventing technical failures. It finds that having identifiable responsible parties does not reliably produce accountability, and that deepfake incidents with anonymous actors trigger the strongest societal and legislative responses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The comparative findings treat absence of an AIID response tag as absence of a response, but tags depend on media submission and editorial practice; category-specific false negatives could explain the apparent accountability gaps.","rationale":"The paper's strongest contribution is reframing the AIID as a record of societal responses rather than implementation failure. I agree with the reader that the media-sourced nature of the database is the central vulnerability, and I want to make the mechanism precise: the comparative claims require not merely that incidents be representative, but that the absence of a response tag be independent of whether a response actually occurred. The authors do not establish this, and their own numbers show why it matters—97% of tagged responses fail the official definition, meaning the 'response' variable is a noisy editorial construct. The paper is careful to disclaim representativeness in Section 3.2.4, but that caveat does not cover the false-negative asymmetry: 'no responses found in the AIID' is repeatedly interpreted as 'no responses occurred' (e.g., third-party LLM incidents), and even when framed as 'found,' the comparison is only meaningful if missingness is category-invariant. The proposed external verification directly tests this. If false negatives are balanced, the central claim survives with its caveats; if not, the quantitative patterns in Section 5 are unsupported, though the qualitative case studies would still stand. Because the reader already conditioned acceptance on addressing data-quality concerns, my read does not change the verdict: the paper should remain conditional until this test is run.","tokens_in":17534,"tokens_out":7162,"duration_ms":74660,"concrete_test":"Stratified external verification: draw a random sample of 20–30 incidents without response tags from each of three categories—(1) Big Tech Developer=Deployer with user harm, (2) third-party LLM deployment with known developer, (3) unknown developer/deployer deepfakes. For each, search Google News, company press pages, regulatory filings, court dockets, and official social media accounts for any public response by the alleged developer/deployer within 90 days of the incident date, using a pre-specified definition matching the AIID's. Compute the false-negative rate (incidents with an external response but no AIID response tag) per category.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5's central claim—identifiable responsible parties do not increase accountability, and unknown-developer/deepfake incidents trigger more societal response—rests on comparing incidents with and without AIID 'response' tags (Figure 7, Appendix E, Sections 4.3 and 5). A missing tag does not mean no response occurred. Tags are assigned only to submitted media reports that editors recognize as responses; a corporate statement, regulatory filing, court action, or press release that no one submitted, or that an editor did not flag, is invisible to the dataset. The paper's own statistics show the label's fragility: 97% of tagged responses did not meet the official definition of a developer/deployer response (Section 4.2), forcing the authors to reclassify them into societal-actor responses, indirect acknowledgements, and official responses. That reclassification is itself a manual judgment without reported inter-rater reliability. The absence side is even noisier. If false-negative rates differ by category—for example, if follow-up reporting on Tesla crashes or corporate statements is thinner than coverage of deepfake legislative reactions—then the observed 0.17 vs 0.09 response proportion favoring unknown parties, and the 'no responses in third-party LLM cases' finding (Section 4.3.2), are documentation artifacts rather than accountability patterns. Section 3.2.4 acknowledges sampling bias generally but does not test the assumption that non-tagging is orthogonal to actual response behavior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the AI Incident Database (AIID) to argue that the database's primary value lies not in technical failure analysis but in documenting how developers, deployers, harmed groups, and wider society respond to AI harms. The authors perform a three-tier mixed-methods analysis of 962 incidents and 4,743 reports, focusing on 48 incidents with 'response' tags. They find that identifiable responsible parties do not necessarily generate more documented responses, that incidents with unknown developers/deployers (mostly deepfakes) attract more societal and legislative responses, and that the content of responses depends on who was harmed and whether deployment was direct or third-party. The paper proposes that AI incident databases can serve as resources for studying accountability and social learning around AI harms.","tokens_in":17753,"tokens_out":8164,"duration_ms":77128,"significance":"If the claims hold, the paper makes a valuable contribution by reframing AI incident databases as records of societal accountability and social learning rather than only as engineering failure logs. The mixed-methods design, the qualitative re-examination of response tags, and the detailed case study of Incident 597 are genuine strengths, and the authors are appropriately candid about the exploratory nature of the analysis and the database's sampling biases. The main findings, however, rest on a small and potentially biased response-tagged subset (48 incidents; only 5 official responses by the authors' own reclassification), and the paper's central comparative claims depend on treating the absence of a response tag as the absence of a response. This assumption is not validated and is load-bearing for the conclusion that unknown responsible parties lead to greater societal response. The paper therefore offers a plausible and interesting reframing, but the empirical support for its specific comparative patterns is fragile and requires additional analysis.","major_comments":[{"comment":"The central comparative claims, including the finding that incidents with unknown developers/deployers have a higher response proportion (0.17 vs 0.09 in Figure 7), treat absence of an AIID response tag as absence of a response. However, response tagging began in 2023 (Section 3.2.2), and the manuscript does not establish whether tags were applied retrospectively to the large number of pre-2023 incidents in the corpus. If tags are only assigned to recently submitted reports, then older incident categories (e.g., Tesla crashes, government AI applications) would be systematically under-tagged, and the observed response-rate differences could be an artifact of submission timing and editorial practice rather than a genuine accountability pattern. The authors should test the missing-tag assumption, for example by restricting the comparison to incidents with reports submitted after the tagging initiative, or by auditing a random sample of untagged incidents for evidence of unrecorded responses.","section":"Section 4.2, Figure 7, Section 5"},{"comment":"The reclassification of the 163 response-tagged reports (97% of which do not meet the official response definition) into 'societal actor responses', 'indirect acknowledgements', and 'official responses' is a manual coding exercise with no reported codebook, inter-rater reliability, or adjudication procedure. Because the subsequent findings—such as the rarity of official responses and the prevalence of societal responses in unknown-developer cases—depend on this reclassification, the coding rubric and reliability statistics (e.g., Cohen's kappa for a subset scored by both authors) should be reported. Without this, the reader cannot assess whether the reclassification is stable or idiosyncratic to the authors.","section":"Section 4.2"},{"comment":"The claim that 'there were no responses in the cases of LLMs deployed by third parties' is based on a small, inductively defined category and on the absence of response tags. Given the tagging and submission biases discussed above, and the fact that corporate statements, regulatory filings, or court actions may exist without being captured by AIID media reports, this absence cannot be interpreted as evidence that no responses occurred. The authors should report the number of incidents in this category and the number of reports examined, and should present a sensitivity analysis or a manual check for undetected responses before treating this as a substantive finding.","section":"Section 4.3.2"},{"comment":"The statement that 'the presence of identifiable responsible parties does not necessarily lead to increased accountability' is phrased as a general claim about accountability, but the operational measure is the presence of a response tag in the AIID, which the paper itself shows is mostly not an official response (97% of tagged reports fail the official definition). The paper should either qualify the claim to refer to 'documented responses in the AIID' or provide a clear argument for why the response-tag proxy, after the authors' reclassification, is a valid measure of accountability. As written, the abstract's second claim overstates what the data can show.","section":"Section 5"}],"minor_comments":[{"comment":"The sentence 'Only 5 out of the 64 reports are official and proactive responses from the developer or deployer' conflicts with the earlier statement that there are 163 response-tagged reports; if 97% of 163 do not meet the definition, the correct denominator should be 163, and '64' appears to be a typo or an unexplained subset. Please clarify.","section":"Section 4.2"},{"comment":"The phrase 'delating changes to how Detroit police uses facial recognition' appears to contain a typo; it should likely be 'detailing changes' or 'delineating changes'.","section":"Section 4.3.2"},{"comment":"The harmed-group labelling is described as 'checked for consistency by cross-labelling,' but no quantitative inter-rater agreement is reported. For reproducibility, please provide a measure such as Cohen's kappa on a shared subset.","section":"Section 3.2.1"},{"comment":"The heatmap's two columns are not explicitly defined in the caption or legend. The text should state clearly that the two columns represent, for each developer/deployer category, the share of all incidents and the share of response-tagged incidents (or whatever the intended comparison is).","section":"Appendix E, Figure 7"},{"comment":"The reference to 'WIRED's Artificial Intelligence Database' links to WIRED's general AI coverage rather than to a structured incident database; this is misleading and should be corrected or replaced with an actual database reference.","section":"Section 2.1.2"},{"comment":"The statement that 'since the response initiative began, only 5.8% of new incidents have responses' is not reconciled with the 48/962 (5%) figure elsewhere; please clarify the denominator and whether responses are tagged retrospectively.","section":"Section 2.1.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the core idea is worth publishing, but the empirical foundation for the comparative claims needs substantial work. The central missing-tag-as-absence assumption and the unvalidated manual reclassification are load-bearing; the authors should be asked to address these with additional analyses (e.g., audit of untagged incidents, restriction to the post-2023 window, inter-rater reliability) before the manuscript can be accepted. The paper would also benefit from a more careful distinction between 'documented response' and 'accountability' in its central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this paper makes a real empirical move: it takes the AIID's 'response' tags, re-reads the underlying reports, and finds that 97% of the tagged responses are not official developer/deployer responses under the database's own definition. That is a useful, original result in itself, and it underlies the paper's central argument that the AIID is better read as a record of societal accountability than as a technical failure log. Second, the comparative findings about who gets responses and who doesn't are real but soft; they rest on 48 response-tagged incidents, 5 of which meet the official definition, and on the assumption that a missing tag means no response happened.\n\nThe paper's framing is well argued and its qualitative work is careful. The case study of the New Jersey deepfake incident gives concrete texture, and the paper is honest about the database's sampling bias throughout. I agree with the reader that the central claim—that identifiable responsible parties do not reliably produce accountability—is plausible and supported by the reclassification exercise.\n\nThe soft spots are real but not fatal. The stress-test note lands: the absence of a response tag is a function of media submission and editorial practice, and the paper never tests whether false negatives in tagging differ by category. So the finding that unknown-developer/deepfake incidents show higher response proportions could be a documentation artifact. The manual reclassification of the 163 tagged reports has no reported inter-rater reliability, and there is an internal inconsistency in the response counts (163 tagged reports earlier, then '5 out of 64 reports' in Section 4.2). These are fixable. The authors should reframe the comparative patterns as hypotheses, add a sensitivity analysis that treats missing tags as missing data, and report coding reliability.\n\nWho is this for? AI governance scholars, incident database curators, and regulators who want to know what can and cannot be learned from this kind of data. It deserves a serious referee, and I would send it to review with a request for revision rather than desk-reject.","headline":"A solid exploratory re-read of the AIID that deserves serious refereeing, but the comparative accountability claims should be treated as hypotheses given how noisy the response tags are.","tokens_in":18322,"tokens_out":1955,"would_cite":true,"duration_ms":18292,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The AI Incident Database's primary value is as a record of societal responses to AI harms, not as a tool for preventing implementation failures.","keywords":["AI Incident Database","algorithmic accountability","social learning","AI harms","incident response","deepfakes","media reporting","responsibility"],"falsifier":"Check whether response rates track media attention rather than actual accountability: for a sample of incidents, compare the AIID's response tags against primary-source records such as court filings, regulatory findings, and company statements. If anonymous deepfake incidents show no response advantage once media attention is controlled for, or if most of their 'responses' are simply more news articles rather than actions by institutions, then the pattern that unknown responsible parties stimulate accountability would collapse into an artifact of media selection.","tokens_in":17274,"feed_emoji":"⚖️","tokens_out":4753,"duration_ms":44326,"temperature":0.7,"pith_summary":"This paper argues that the AI Incident Database is best understood not as an aviation-style failure log for engineers, but as a record of how society reacts when AI causes harm. Analyzing 962 incidents and 4,743 reports, the authors find that having an identifiable responsible party does not reliably lead to accountability: many incidents involving major technology companies drew few or no formal responses, while anonymous deepfake incidents stimulated the most public outcry and legislative action. They also find that the likelihood and substance of a response depend on context, including who was harmed and whether deployment was direct or third-party. The paper concludes that the database's real value lies in documenting patterns of harm, institutional response, and social learning around AI failures.","feed_headline":"AI incident logs are a map of who answers, not what failed","feed_subtitle":"An analysis of 962 incidents finds identifiable culprits rarely respond; anonymous deepfake harms draw the most action","key_machinery":"The central object is the AIID's 'response'-tagged reports: 163 reports spread over 48 incidents that the database editors marked as public official responses from an entity allegedly responsible for developing or deploying the AI system. The paper treats these tags not as reliable corporate disclosures but as traces of societal reaction, and supplements them with qualitative reading of all 638 reports attached to response incidents. The method is a three-tier analysis: first, all 962 incidents are categorised by developer/deployer relationship and harmed group; second, the response-tagged incidents are examined qualitatively; third, inductive 'typical incident' categories, such as Big Tech user harm, third-party LLM deployment, government applications, and deepfakes, are compared to find which contextual factors make a substantive response more or less likely. This machinery lets the authors make meaningful comparisons between similar incidents with and without responses, despite acknowledging that the database is not representative of all AI harms.","core_discovery":"The paper's central claim, stated in Section 5, is that the primary value of the AI Incident Database is not learning to avoid implementation failure, but learning about the state of incidents and the responses of different actors in the wake of AI harm. Through a three-tier mixed-methods analysis of 962 incidents and 4,743 reports, the authors show that the presence of identifiable responsible parties does not necessarily lead to increased accountability. Incidents where a major technology company is both developer and deployer, such as Tesla crashes and social media harms, are proportionally less likely to receive responses than incidents where the responsible parties are unknown, which are mostly deepfake cases that generated substantial societal reaction and calls for legislation. When substantive responses do occur, they are shaped by context: organisational victims tend to receive more formal investigations than individual users, and regulatory or legal pressure often accounts for the difference. The authors also find that 97% of reports tagged as responses in the database do not meet the definition of an official developer or deployer response, and that the few official responses that exist are superficial. They conclude that the AIID serves as a record of societal accountability and social learning, and that both controversy-rich and controversy-absent incidents offer insight into how society negotiates responsibility for AI harms.","pith_inferences":["If the response patterns hold, regulators could use incident databases as early-warning sensors for accountability gaps, flagging categories where harms recur with no meaningful response.","The same three-tier method could be applied to other incident collections, such as AIAAIC or future EU AI Act complaint logs, to test whether new regulation shifts response patterns over time.","The finding that organisational victims receive more substantive responses suggests a testable hypothesis: as AI procurement shifts toward institutions, accountability will concentrate where economic leverage exists, potentially leaving individual consumers underserved.","A direct extension would be to code the language of corporate responses for 'learning signals'—specific technical changes, apologies, or commitments—and test whether any response type correlates with reduced recurrence of similar incidents."],"forward_implications":["AI incident databases can be read as social records: they show how society negotiates responsibility after AI harms, even though they cannot support aviation-style technical failure-avoidance learning.","Knowing the responsible party is not enough: incidents with named Big Tech developers and deployers are less likely than anonymous deepfake incidents to receive a substantive response.","The quality of response is context-dependent: organisational victims often get formal investigations while individual users get blog posts or silence, unless regulation or courts force a stronger response.","The near-total absence of official developer/deployer responses means that voluntary corporate response reporting has not produced the technical learning loop the AIID originally sought.","Absence of controversy is itself diagnostic: incidents where governments harmed the public with AI, such as wrongful arrests, drew little response and reveal accountability blind spots."],"supporting_citations":[{"why":"Establishes the AIID's origin and aviation-inspired mission of collective learning from failures, which the paper reinterprets.","marker":"[McGregor, 2021]"},{"why":"Supplies the formal definition of an AI Incident Response that the paper uses to evaluate tagged responses.","marker":"[Schwartz and Cunningham, 2022]"},{"why":"Provides the social-learning lens for reading public reactions to autonomous vehicle incidents as a window onto haphazard social learning.","marker":"[Stilgoe, 2023]"},{"why":"Supports the treatment of public controversies documented in media as empirical occasions showing relations among actors.","marker":"[Marres and Moats, 2015]"},{"why":"Establishes that incident reporting data typically reflects reporting behaviour more than underlying occurrence, grounding the paper's caution about sampling bias.","marker":"[Macrae, 2016]"},{"why":"Represents the criticism that media reports lack sufficient information for implementation-failure learning, which the paper accepts and pivots away from.","marker":"[Durso et al., 2022]"},{"why":"Provides the interpersonal-ethics framing that responsibility centres on responsiveness to harms, not only on prevention.","marker":"[Strawson et al., 2003]"},{"why":"Supplies the notion of an 'adequate response' and contextual factors affecting accountability, which the paper applies to AI incidents.","marker":"[Raji et al., 2022]"},{"why":"Guides the interpretation of absent controversy, showing how media framing can freeze or suppress public debate around AI implementation.","marker":"[Dandurand et al., 2023]"},{"why":"Contributes the 'many hands' discussion of diffused responsibility, which the paper extends to social media obscuring deepfake perpetrators.","marker":"[Nissenbaum, 1996]"}],"fun_headline_variants":["Who answers for AI harms? Not the builders","Deepfake chaos sparks more response than Tesla crashes","AI incident logs show blame goes unanswered","AI incident data reveals an accountability vacuum"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central patterns depend on the AIID's media-sourced reports and response tags being a sufficiently faithful record of real societal and institutional responses that comparisons across categories—developer type, harmed group, response presence—are meaningful rather than artifacts of what got reported and tagged.","fun_headline_variants_meta":{"raw":{"variants":["Who answers for AI harms? Not the builders","Deepfake chaos sparks more response than Tesla crashes","AI incident logs show blame goes unanswered","AI incident data reveals an accountability vacuum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000565,"raw_usage":{"total_tokens":2757,"prompt_tokens":1101,"completion_tokens":1656,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":1601}},"tokens_in":717,"tokens_out":1656,"duration_ms":12589,"temperature":1.0,"reasoning_tokens":1601,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:32:15.313142+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether response rates track media attention rather than actual accountability: for a sample of incidents, compare the AIID's response tags against primary-source records such as court filings, regulatory findings, and company statements. If anonymous deepfake incidents show no response advantage once media attention is controlled for, or if most of their 'responses' are simply more news articles rather than actions by institutions, then the pattern that unknown responsible parties stimulate accountability would collapse into an artifact of media selection.","supporting_citations":[{"cited_title":"Preventing repeated real world ai failures by cataloging incidents: The ai incident database","cited_arxiv_id":null,"evidence_quote":"Establishes the AIID's origin and aviation-inspired mission of collective learning from failures, which the paper reinterprets."},{"cited_title":"Enhancing ai incident reports through developer response","cited_arxiv_id":null,"evidence_quote":"Supplies the formal definition of an AI Incident Response that the paper uses to evaluate tagged responses."},{"cited_title":"Machine learning, social learning and the governance of self-driving cars","cited_arxiv_id":null,"evidence_quote":"Provides the social-learning lens for reading public reactions to autonomous vehicle incidents as a window onto haphazard social learning."},{"cited_title":"The problem with incident reporting","cited_arxiv_id":null,"evidence_quote":"Establishes that incident reporting data typically reflects reporting behaviour more than underlying occurrence, grounding the paper's caution about sampling bias."},{"cited_title":"Analyzing failures in artificial intelligent learning systems (fails)","cited_arxiv_id":null,"evidence_quote":"Represents the criticism that media reports lack sufficient information for implementation-failure learning, which the paper accepts and pivots away from."},{"cited_title":"Freedom and resentment","cited_arxiv_id":null,"evidence_quote":"Provides the interpersonal-ethics framing that responsibility centres on responsiveness to harms, not only on prevention."},{"cited_title":"Outsider oversight: Designing a third party audit ecosystem for ai governance","cited_arxiv_id":null,"evidence_quote":"Supplies the notion of an 'adequate response' and contextual factors affecting accountability, which the paper applies to AI incidents."},{"cited_title":"Freezing out: Legacy media's shaping of ai as a cold controversy","cited_arxiv_id":null,"evidence_quote":"Guides the interpretation of absent controversy, showing how media framing can freeze or suppress public debate around AI implementation."},{"cited_title":"Accountability in a computerized society","cited_arxiv_id":null,"evidence_quote":"Contributes the 'many hands' discussion of diffused responsibility, which the paper extends to social media obscuring deepfake perpetrators."}],"review_version":1}