{"id":"88f94330-54fb-49cd-aa94-6b066b45fbce","arxiv_id":"2607.04601","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"AIGC bans raise question volume ~13% while cutting the share of questions with accepted answers inside eight hours by ~3.3 percentage points, effects confined to non-STEM Stack Exchange communities.","lead":"Banning AI-generated answers on Stack Exchange raises the number of questions posted but lowers the share that get a satisfactory answer quickly, and only in non-STEM communities. Platform operators deciding how to compete with LLMs need this trade-off between engagement and resolution speed.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The efficiency decline may partly reflect post-ban question composition shifts rather than pure loss of AI productivity, weakening the double-edged claim.","rationale":"The reader correctly flags matching/exchangeability as a residual observational risk and rates the paper CONDITIONAL with high confidence. That concern is real but secondary: the paper already shows balance on AIGC rates/discussions, parallel trends (Figure 2), and recovers the pattern with Callaway–Sant’Anna, counterfactual, and two-stage estimators. The more load-bearing soft spot for the strongest claim is the causal attribution of the efficiency decline itself. Because the authors document (and partially endorse) post-ban question hardening exactly where efficiency falls, the –0.033/–0.067 coefficients cannot be cleanly read as “contribution efficiency falls because AIGC cost advantages were removed.” A composition-controlled re-estimate would settle whether the double-edged trade-off survives. This does not overturn the seeking result or the non-STEM heterogeneity, so the verdict remains CONDITIONAL rather than REJECT; it simply sharpens the condition that must be met for the efficiency half of the claim.","tokens_in":21120,"tokens_out":560,"duration_ms":4944,"concrete_test":"Re-estimate the WithAcceptedAsw DiD (Eq. 1 / Table 2 Col. 3 and non-STEM Panel B) after residualizing or controlling for the three post-ban question-content measures (length, subjectivity, socialness) at the community-week level, or reweight questions to pre-ban content distributions. If the AIGC_Ban coefficient on efficiency shrinks by >50% or loses significance, the pure productivity-loss interpretation of the double-edged effect is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim treats the drop in WithAcceptedAsw (–0.033 overall; –0.067 non-STEM) as reduced contribution efficiency from removing AIGC’s cost-reducing advantage. Yet Section 6.3 and Table 8 show that, precisely in non-STEM communities, questions become longer, more subjective, and higher-socialness after the ban, while answers also lengthen. The paper itself notes that these adaptations “may explain the observed efficiency decline.” If harder questions mechanically lower the eight-hour acceptance rate, the efficiency coefficient confounds a pure productivity loss with a demand-side composition effect. Parallel trends and staggered estimators address selection into ban but do not isolate whether the same questions would have been answered more slowly without AIGC. The double-edged framing (more seeking + less efficiency) therefore rests on an untested attribution of the efficiency drop.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper estimates the causal effects of staggered AIGC bans on Stack Exchange communities after ChatGPT’s release, using kernel propensity-score matching and two-way fixed-effects DiD (Eq. 1) on community-week data. Main results (Table 2) show a ~13% rise in question volume, no change in answers per question, and a 3.3 pp drop in the share of questions receiving an accepted answer within eight hours; effects are concentrated in non-STEM communities (Table 3) and are recovered by Callaway–Sant’Anna, counterfactual, and two-stage DiD estimators. Mechanism analyses split communities by human-rated AIGC reliability and a socialness lexicon (Tables 6–7) and document post-ban shifts toward longer, more subjective, higher-socialness questions and answers (Table 8). The authors interpret the pattern as a socio-technical trade-off: bans raise seeking where humans hold comparative advantage but reduce contribution efficiency where LLMs previously lowered production costs.","tokens_in":21337,"tokens_out":1175,"duration_ms":10839,"significance":"If the estimates hold, the paper supplies one of the first quasi-experimental accounts of platform governance under generative-AI competition, documenting a clear engagement–efficiency trade-off that is heterogeneous by domain and by two socio-technical dimensions (reliability and social interactivity). The design is unusually thorough for the literature: matching balance checks (including AIGC-related covariates), parallel-trends plots, multiple staggered-DiD estimators, placebo and look-ahead matching, detector validation (Figure 1), and external mechanism measures. These features make the reduced-form results a useful benchmark for platform managers and for subsequent work on AI substitution in knowledge communities. The socio-technical comparative-advantage framing also productively extends classic machine-substitution arguments beyond task routineness.","major_comments":[{"comment":"Section 6.3 and Table 8 show that, precisely in non-STEM communities where efficiency falls, questions become longer, more subjective, and higher-socialness after the ban, and answers lengthen correspondingly. The paper itself notes that these adaptations “may explain the observed efficiency decline.” The central double-edged claim (Table 2 Col. 3; Table 3 Panel B) attributes the drop in WithAcceptedAsw to loss of AIGC’s cost-reducing advantage. Without isolating composition (e.g., reweighting by pre-ban question characteristics, conditioning on length/subjectivity/socialness, or an event-study of efficiency for fixed question types), the efficiency coefficient confounds pure productivity loss with demand-side shifts toward harder questions. This attribution is load-bearing for the “double-edged” framing and for the practical claim that bans reduce contribution efficiency.","section":"§6.3, Table 8; Tables 2–3"},{"comment":"The efficiency measure is the share of questions with an accepted answer within eight hours (Table 1; §4.4). Acceptance is an asker choice that may itself respond to the ban (e.g., higher standards for “human” answers, delayed acceptance of more complex posts). Alternative windows and controls are reported in Appendix C.1, but the manuscript does not show that the decline survives when efficiency is measured by time-to-first-answer, answer arrival rates, or non-acceptance-based resolution. Clarifying whether the result is robust to pure speed metrics would strengthen the contribution-efficiency interpretation.","section":"§4.4, Table 1; Appendix C.1"}],"minor_comments":[{"comment":"Figure 2 notes pre-period differences for question volume before week −8; a short discussion of why those early deviations do not threaten identification (or a restricted-window robustness check) would help readers.","section":"Figure 2, §5.4.1"},{"comment":"The AIGC detector threshold (90% likelihood) and the median splits for reliability/socialness are free parameters; sensitivity tables for alternative cutoffs would be useful.","section":"§4.2; Tables 6–7"},{"comment":"Title and abstract use “Generative AI” / “AIGC” somewhat interchangeably; a single consistent term after first definition would improve clarity.","section":"Title, Abstract"},{"comment":"Footnote 8 acknowledges that LLM reasoning has improved since the sample; a brief caveat in the discussion about external validity to later models would be appropriate.","section":"§5.3, §7.3"}],"recommendation":"major_revision","confidential_remarks":"The empirical package is careful and the non-STEM heterogeneity is interesting; the main risk is over-claiming a pure efficiency loss when the paper’s own content analysis points to composition. If the authors can either partial out question difficulty or reframe the efficiency result as a joint productivity-plus-composition effect, the paper would be a solid contribution. Scope fits cs.CY / IS journals well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first careful quasi-experiment on AIGC bans rather than ChatGPT’s release. The headline result is clean: after the ban, questions rise ~13% while the share of questions getting an accepted answer inside eight hours falls ~3.3 points, with no change in answers per question; everything is confined to non-STEM communities. Matching + staggered DiD + Callaway–Sant’Anna + matrix-completion counterfactuals + two-stage DiD + placebo + look-ahead matching all point the same way, and Figure 1 shows the ban actually cut detected AI answers. That package is better than most of the recent Stack Overflow papers.\n\nWhat is new is the efficiency margin, the non-STEM split, and the two external mechanism tests (human-rated LLM reliability on pre-ChatGPT accepted answers; Diveica socialness lexicon). The socio-technical framing is not just window dressing; the splits line up with the theory and the content-adaptation results in Table 8 are consistent with it.\n\nThe soft spot the stress-test flags is real but already half-acknowledged by the authors. In non-STEM communities questions become longer, more subjective, and higher-socialness after the ban, and answers lengthen too. The paper itself says these adaptations “may explain the observed efficiency decline.” So the –0.033 (or –0.067) coefficient mixes pure loss of AI productivity with a demand-side composition shift. Parallel trends and the staggered estimators do not isolate the pure productivity channel. That weakens the sharpest version of the “double-edged” claim, but it does not overturn the reduced-form facts or the heterogeneity. Residual selection into ban and imperfect AI detection remain, yet none look load-bearing given the robustness battery.\n\nCitation pattern is appropriate; methods are standard and transparent. This is useful for anyone working on platform governance, online knowledge systems, or AI substitution in socio-technical settings. I would send it to peer review; a good referee can push them to quantify how much of the efficiency drop survives after controlling for question complexity. Worth reading and, for the right paper, worth citing.","headline":"Solid first causal look at AIGC bans (not just LLM release) that cleanly shows a seeking–efficiency trade-off concentrated in non-STEM communities; the efficiency drop is partly confounded by harder post-ban questions, but the paper already flags this and the rest of the design holds up.","tokens_in":21939,"tokens_out":548,"would_cite":true,"duration_ms":5182,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Banning AI-generated answers on Stack Exchange raises question volume by about 13% but lowers the share of questions answered on time, and only in non-STEM communities.","keywords":["Generative AI","Online Q&A Community","AIGC","Large Language Models","Knowledge Exchange","Platform Competition","Difference-in-Differences","Stack Exchange"],"falsifier":"Re-estimate the same DiD (or Callaway–Sant’Anna) specification after adding later-banning communities as controls or after instrumenting ban adoption with pre-period AIGC-discussion intensity; if the positive question and negative efficiency coefficients disappear or reverse sign, the central claim fails.","tokens_in":22021,"feed_emoji":"⚖️","tokens_out":928,"duration_ms":7611,"temperature":0.7,"pith_summary":"This paper asks what happens when online Q&A communities ban AI-generated content after ChatGPT arrives. Using staggered bans across Stack Exchange sites and a matched difference-in-differences design, the authors show a double-edged result: more questions get posted, yet a smaller share of them receive an accepted answer inside the usual eight-hour window, while answers per question stay flat. The trade-off appears only in non-STEM communities. The authors argue that the ban repositions human-centered platforms along two socio-technical dimensions—informational reliability and social interactivity—so seekers return where humans hold a comparative advantage, while contributors lose the speed gains that reliable AI once supplied where social expectations are low. Both askers and answerers then adapt by writing longer, more subjective, and more socially flavored posts. The practical upshot is that bans can restore engagement but at an efficiency cost, and that mixed policies keyed to reliability and social demand may be wiser than blanket rules.","feed_headline":"AI answer bans raise questions 13% but slow replies","feed_subtitle":"Only non-STEM Stack Exchange sites show the trade-off; reliability and social demand explain why","key_machinery":"A matched staggered difference-in-differences design that exploits community-level AIGC bans, combined with a socio-technical comparative-advantage frame that treats informational reliability and social interactivity as the two dimensions along which human platforms and LLMs compete.","core_discovery":"Banning AIGC on Stack Exchange produces a double-edged effect confined to non-STEM communities: knowledge seeking rises (roughly 13% more questions) while contribution efficiency falls (about 3.3 percentage points fewer questions receiving an accepted answer within eight hours), with no change in answers per question. The same pattern holds under alternative staggered DiD estimators. Mechanism tests show questions rise where AI reliability is low and social interactivity is high, while efficiency falls where AI is reliable and social demand is low; content becomes longer, more subjective, and more social after the ban.","pith_inferences":["As reasoning models improve STEM reliability, the same efficiency penalty may eventually appear in STEM communities that ban AIGC.","The same reliability-and-social-interactivity frame can be used to predict which Reddit or enterprise knowledge bases will gain or lose from generative-AI bans.","If platforms publish real-time AIGC-detection rates, researchers can test whether enforcement intensity, rather than the ban announcement itself, drives the observed trade-off."],"forward_implications":["Platform managers who ban AIGC should expect more questions but slower resolution, especially outside STEM.","Efficiency losses concentrate where LLMs are already reliable and social interaction is not the main draw.","Question and answer text will become longer, more subjective, and more socially flavored after a ban.","A mixed policy that labels some questions ‘open for AIGC’ and others ‘human-only’ can preserve speed without sacrificing authenticity.","Complementary incentives or expert routing will be needed to keep experienced contributors answering the harder posts that arrive after a ban."],"fun_headline_variants":["AI bans lift non-STEM questions 13% but cut answer efficiency","Stack AIGC bans: more seeking, slower replies outside STEM","Double-edged ban raises Qs 13% yet slows answers in non-STEM","AIGC bans boost questions where AI weak, slow replies where strong","Non-STEM Stack bans: higher volume, lower 8-hour answer rates"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That matching treated and control communities on eight pre-ban observables plus AIGC-related balance checks makes ban adoption as-good-as-random and keeps pre-ban trends parallel.","fun_headline_variants_meta":{"raw":{"variants":["AI bans lift non-STEM questions 13% but cut answer efficiency","Stack AIGC bans: more seeking, slower replies outside STEM","Double-edged ban raises Qs 13% yet slows answers in non-STEM","AIGC bans boost questions where AI weak, slow replies where strong","Non-STEM Stack bans: higher volume, lower 8-hour answer rates"]},"model":"grok-4.5","effort":"low","cost_usd":0.006532,"raw_usage":{"total_tokens":1744,"prompt_tokens":887,"num_sources_used":0,"completion_tokens":103,"cost_in_usd_ticks":65320000,"prompt_tokens_details":{"text_tokens":887,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":754,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":887,"tokens_out":103,"duration_ms":6887,"temperature":1.0,"reasoning_tokens":754,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T16:35:50.975643+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-estimate the same DiD (or Callaway–Sant’Anna) specification after adding later-banning communities as controls or after instrumenting ban adoption with pre-period AIGC-discussion intensity; if the positive question and negative efficiency coefficients disappear or reverse sign, the central claim fails.","supporting_citations":[],"review_version":1}