{"id":"80e07f3d-72ff-49b7-8fb1-f22804d30a04","arxiv_id":"2607.10820","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A flow-balance identity yields a closed-form geometric-amplification estimator of the global treatment effect of sharing features under network flywheel interference, with consistency under homogeneity and valid A/A inference.","lead":"The paper gives a log-based estimator that corrects A/B tests of social sharing features for multi-round flywheel cascades that break standard no-interference assumptions. Platforms can use it with existing attribution logs to measure true global engagement lift before shipping share redesigns.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Homogeneous q (A2) is the load-bearing gap for consistency of the geometric plug-in, but sims already stress it and the real A/B is only one platform.","rationale":"The reader's weakest_assumption correctly isolates A2 as the single condition on which the closed-form geometric estimator's consistency rests. The flow-balance identity (Prop. 2) and the representation IS = μ^d/(1-q_eff) are exact only when a common effective rate exists; without A2 the plug-in bq is an average that need not equal the correct amplification factor for the GTE contrast. The paper's own sims already break A2 on the sender side and still show low bias, and the A/A CLT (Thm 2) does not require A2, so the concern is not fatal. It does, however, keep the main claim conditional on further heterogeneity checks and multi-platform evidence. No stronger internal inconsistency appears: proofs under the stated assumptions are coherent, code is released, and the real A/A p-value uniformity is a useful diagnostic. Verdict therefore stays CONDITIONAL; no upgrade to ACCEPT and no downgrade to REJECT.","tokens_in":22957,"tokens_out":674,"duration_ms":9247,"concrete_test":"In the released Hawkes simulator, replace the homogeneous-q construction with content- and receiver-dependent q_ik = θ_k · φ_j (draw θ_k ~ U(0.1,0.9), φ_j ~ U(0.05,0.4)), keep A1 bounds, and recompute bias/MSE of [GTE vs DM/DM-FO over 200 replications at \rho0∈{0.3,0.6,0.8}. If absolute bias of [GTE exceeds 15% of true GTE (or exceeds DM-FO) at any \rho0, the consistency claim under realistic heterogeneity fails and the CONDITIONAL verdict should tighten.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 1 proves [GTE \to p GTE only under A2: q^a_ik ≡ q_a for every (i,k) in each regime. The estimator plugs a single bq_a = bY^s_a / cW^s_a into the geometric factor 1/(1-bq_a). If downstream rates vary systematically with sender type, content, or treatment (e.g., treated users preferentially share high-virality content, or high-degree users have higher q), the common-q representation of GTE is misspecified and the plug-in need not recover the true global contrast. The paper itself flags A2 as stronger than necessary (§4.3) and relies on aggregation plus sims that already violate A2 via sender-side δ_i perturbations. That is supportive but not a proof that bias vanishes under realistic receiver/content heterogeneity. The real-platform Table 2 significance therefore rests on an untested extrapolation of that robustness.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies A/B evaluation of sharing features under multi-round social interference (the “flywheel”). It models share-induced views via a multivariate Hawkes process, derives a sender–receiver flow-balance identity, and interprets share-induced engagement as geometric amplification with effective downstream rate q_eff. This yields a log-based plug-in estimator [GTE = bY^d_T/(1−bq_T) − bY^d_C/(1−bq_C) for the global treatment effect on share-induced engagement, using attribution logs under Bernoulli randomization. Under Assumptions A1–A2 the estimator is consistent (Theorem 1); under A/A tests a finite-population CLT and plug-in SE give asymptotic Type I control (Theorem 2). Simulations (n=50k, multiple graphs/parameters) show large bias reduction versus DM, DM-FO, and GCR; a Poisson-heuristic extension targets reactivation; a real-platform A/A is roughly uniform and one A/B finds a significant effect where baselines do not.","tokens_in":23295,"tokens_out":953,"duration_ms":30028,"significance":"If the method is reliable beyond the homogeneous-q regime, it fills a clear gap: interference from multi-round sharing is practically important and under-studied relative to marketplace and neighborhood interference. Strengths include a transparent flow-balance derivation that does not require fitting Hawkes kernels, closed-form propagation adjustment from standard attribution logs, machine-checkable-style proofs in Appendix A, public simulation code, extensive robustness sweeps (topology, density, parameter laws, spectral radius), and real deployment with A/A pipeline validation. The geometric representation is simple enough for production experimentation stacks. The main scientific value is a practical GTE estimator tailored to share cascades rather than generic exposure mappings.","major_comments":[{"comment":"Theorem 1 establishes consistency of [GTE only under Assumption A2 (homogeneous downstream rate q^a_ik ≡ q_a for all (i,k) within each regime). The estimator plugs a single bq_a = bY^s_a/cW^s_a into the geometric factor 1/(1−bq_a). Section 4.3 correctly flags A2 as stronger than necessary and appeals to aggregation plus simulations that already violate A2 via sender-side δ_i. That is supportive but not a proof that the common-q representation of GTE remains approximately unbiased under systematic receiver/content/treatment heterogeneity (e.g., treated senders preferentially share high-virality content, or high-degree receivers have higher q). For the central practical claim, please either (i) prove consistency/approximate unbiasedness under a weaker average-q or L2-heterogeneity condition, or (ii) add a targeted bias analysis/simulation design where q_ik varies with treatment, content, a","section":null},{"comment":"Theorem 2 and the SE in Eq. (11) are justified only under A/A tests (identical kernels across arms). Table 1 reports near-nominal 95% coverage in A/B simulations, and Table 2 reports A/B p-values (e.g., 0.005 for the proposed method) that drive the launch decision narrative in §8. The manuscript does not state conditions under which bse remains valid when treatment changes dd, ds and thus the joint law of (Y^d, Y^s, W^s). Please either extend the CLT/SE theory to local alternatives / A/B under A1–A2, or clearly label Table 2 A/B inference as heuristic, report bootstrap or design-based alternatives, and temper the claim that the feature is “statistically significant” solely on the A/A-derived SE.","section":null},{"comment":"Section 6’s reactivation estimator [GTE_ra = exp(−bλ_C)−exp(−bλ_T) is presented as a Poisson approximation with no consistency theorem, and Figure 2 shows remaining bias (though lower MSE than EW/HEW). The abstract and contributions list this as a framework extension on equal footing with the IS estimator. Please either supply conditions under which the Poisson plug-in is consistent for GTEra, or reframe §6 as an exploratory heuristic and avoid implying the same guarantees as Theorem 1 for the reactivation metric.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is a log-based plug-in for the global treatment effect on share-induced engagement under multi-round sharing: [GTE = bY^d_T/(1-bq_T) - bY^d_C/(1-bq_C), from a sender–receiver flow-balance identity and a geometric amplification factor. That is the actual novelty. Network interference, Hawkes cascades, and cluster randomization are all prior art; the contribution is the model-free closed form tailored to attribution logs and the A/A inference pipeline that goes with it.\n\nWhat they do well is concrete. The identity and the consistency argument under A1–A2 are clean (Appendix A). Sims at n=50k with 500 replications, multiple topologies, parameter distributions, and stronger receiver-side propagation show clear bias reduction versus DM, DM-FO, and GCR; A/A coverage sits near nominal. The real-platform A/A p-values are roughly uniform, and the A/B case (Table 2) is the right kind of evidence: baselines insignificant, theirs significant, and post-launch behavior consistent. Code is public. Citation pattern is honest about marketplace interference and exposure-mapping work.\n\nThe soft spot is A2 (homogeneous downstream q). Theorem 1 needs a single q_a per regime; if q varies systematically with sender type, content, or treatment, the geometric factor is misspecified. The paper flags this itself (§4.3). Sims already break A2 via sender-side δ_i and still look good, which is supportive but not a proof under realistic receiver/content heterogeneity. The reactivation extension is explicitly heuristic (Poisson). Neither of these sinks the main estimator; they bound how far you should trust the theory outside the stated regime.\n\nThis is for people who run or design A/B systems on social/content platforms and for methodologists working on interference with logs. It deserves a serious referee. I would engage with it and cite the estimator when the setting matches.","headline":"Clean, deployable GTE estimator for sharing flywheels; A2 is a real but already-stressed soft spot, not a collapse.","tokens_in":23817,"tokens_out":497,"would_cite":true,"duration_ms":7110,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Share-feature A/B tests can recover the platform-wide flywheel effect with a closed-form adjustment from attribution logs, not cluster randomization.","keywords":["A/B tests","experimental design","causal inference","interference","social network","global treatment effect","Hawkes process","flywheel effect"],"falsifier":"On a platform where per-user or per-content downstream share rates can be measured separately under global treatment and global control, check whether the single-rate plug-in GTE matches the true full-deployment difference; large systematic gaps would falsify the homogeneous-rate claim that carries consistency.","tokens_in":23865,"feed_emoji":"🔄","tokens_out":946,"duration_ms":19776,"temperature":0.7,"pith_summary":"Online platforms use sharing features to reactivate users and amplify engagement through multi-round social cascades, the so-called flywheel effect. Classical A/B tests miss this impact because one user’s treatment changes what others see and share, violating the no-interference assumption. This paper models sharing as a multivariate Hawkes process and shows that share-induced engagement obeys a sender–receiver flow-balance identity. From that identity it builds a plug-in estimator that treats multi-round diffusion as geometric amplification and corrects for it using ordinary attribution logs. Under mild conditions the estimator is consistent for the global treatment effect, an A/A procedure controls Type I error, and simulations plus a large-platform deployment show lower bias than difference-in-means or first-order adjustments—sometimes flipping an insignificant launch decision into a significant one.","feed_headline":"Flywheel A/B tests fixed with a geometric log-based correction","feed_subtitle":"Attribution logs recover multi-round share impact that difference-in-means and first-order metrics miss.","key_machinery":"The flow-balance identity equating total sender-side offspring share-views to total receiver-side share-induced views; it yields the geometric representation IS = (average discovery-driven offspring)/(1−q_eff) and the closed-form propagation-adjusted GTE estimator.","core_discovery":"The global treatment effect of a sharing feature on share-induced engagement equals the difference, between full treatment and full control, of discovery-driven share-view volume scaled by the geometric multiplier 1/(1−q_eff), where q_eff is the effective downstream sharing rate. This quantity is identified from Bernoulli A/B data by the plug-in estimator that replaces discovery-driven offspring counts and the two group-level rates with their sample analogues from attribution logs, and the estimator is consistent when downstream rates are homogeneous within regime.","pith_inferences":["The same flow-balance plus geometric correction may apply to other self-exciting product loops (referral bonuses, invite chains, collaborative play) whenever logs separate discovery-origin events from socially induced ones.","If platforms store generation depth or multi-hop attribution, one could test whether a generation-stratified rate estimator shrinks residual bias when homogeneity fails.","Heterogeneous or near-critical cascades (spectral radius close to one) are the regime where the method’s advantage over first-order adjustments should grow most, matching the paper’s stronger-propagation simulations."],"forward_implications":["Product teams can evaluate multi-round sharing features with ordinary Bernoulli randomization and existing attribution logs, without redesigning experiments as graph clusters.","Difference-in-means and first-order cascade metrics systematically understate flywheel impact and can miss launch-worthy effects that the propagation-adjusted estimator detects.","A valid A/A pipeline check exists: under identical treatment and control the estimator is asymptotically normal with a plug-in standard error that keeps Type I error near nominal.","The same discovery-driven volume and downstream rates give a Poisson-based estimator for the global effect on user reactivation probability.","When the estimator is significant and baselines are not, platforms have a quantitative basis to launch and to validate post-launch against pre-launch metrics."],"fun_headline_variants":["Geometric logs fix flywheel bias in share A/B tests","Attribution recovers multi-round share engagement via 1/(1-q)","Flow-balance estimator corrects interference in sharing features","Plug-in geometric adjust yields consistent global share effects","Discovery volumes scaled by q_eff identify full flywheel impact"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Downstream sharing rates are the same for every user–content pair within a treatment regime; if those rates systematically differ by who is treated or who receives the content, the single geometric factor is wrong and bias returns.","fun_headline_variants_meta":{"raw":{"variants":["Geometric logs fix flywheel bias in share A/B tests","Attribution recovers multi-round share engagement via 1/(1-q)","Flow-balance estimator corrects interference in sharing features","Plug-in geometric adjust yields consistent global share effects","Discovery volumes scaled by q_eff identify full flywheel impact"]},"model":"grok-4.5","effort":"low","cost_usd":0.006638,"raw_usage":{"total_tokens":1729,"prompt_tokens":838,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":66380000,"prompt_tokens_details":{"text_tokens":838,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":824,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":838,"tokens_out":67,"duration_ms":9564,"temperature":1.0,"reasoning_tokens":824,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T09:00:09.246630+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a platform where per-user or per-content downstream share rates can be measured separately under global treatment and global control, check whether the single-rate plug-in GTE matches the true full-deployment difference; large systematic gaps would falsify the homogeneous-rate claim that carries consistency.","supporting_citations":[],"review_version":1}