{"id":"1eb643bc-881d-434b-b5d8-31c419ad15aa","arxiv_id":"2506.21598","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A cluster-randomized A/B test design for paid search that builds query-product SERP interference networks, projects them to product graphs, and partitions them to reduce SUTVA bias.","lead":"Search advertisers cannot easily A/B test new bidding algorithms because products shown together on a search results page interfere with each other. This paper builds a network from search query reports, clusters products that appear together, and randomizes experiments by cluster instead of by product.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Censored query-report data means the SERP graph and the 36% leakage are computed from won auctions only; the central 'good estimate of actual lift' claim stands or falls on this unvalidated graph.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the observational query-report data are censored to won auctions, so the projected product graph and all downstream cluster/leakage quantities are computed from a potentially biased view of true SERP interference. This is not a stylistic objection; it directly targets the condition needed for the central claim that the cluster-randomized design gives a good estimate of actual lift. If the interference graph misses edges that involve lost auctions, METIS can place strongly interfering products in different clusters, and the reported 36% between-cluster leakage becomes an unreliable lower-bound-style quantity rather than a validation. The paper acknowledges censorship only as a limitation and defers better visibility to future work, but the current validation (a single non-concurrent product-split comparison with 44% vs. 24% lift, plus a vague 'consistent' DID result) cannot rule out the possibility that the cluster estimate is biased in the same direction as the product-split estimate, merely to a smaller degree. The proposed test using internal bid-request logs is feasible because the advertiser's own bidding system already participates in auctions it loses; comparing clusters built from those logs against the censored query-report graph would directly quantify the distortion. If the test shows little distortion, the concern is resolved and the method's promise is supported; if not, the conclusion should be softened to a design proposal with an illustrative case study, which is exactly the conditional posture the reader already recommended. Therefore no verdict change is needed.","tokens_in":8040,"tokens_out":6712,"duration_ms":94667,"concrete_test":"Rebuild the bipartite graph for a three-month window from internal bid-request logs that record every auction the advertiser participated in, including losses, and regenerate product clusters using the paper's W_uni (Eq. 1) and METIS. Compare the clusters and true between-cluster leakage against the censored query-report version. If the uncensored leakage exceeds the reported 36% by more than, say, 10 percentage points, or if the adjusted Rand index between the two clusterings is below 0.8, then censorship materially changes the interference graph and the near-unbiasedness claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 claims the cluster-randomized design 'gives a good estimate of the actual lift,' but every component of that design is built from the ad publisher's search query report, which Section 6 concedes is 'censored, that is, we only have data when we win the auction.' The reported data (Table 1) contains rows only for products whose ad was actually shown; if a product was eligible for a query but lost the auction, no row exists. Equation (1) therefore constructs W_uni co-occurrence edges only among products that jointly won impressions. The interference we care about for a bidding change occurs at the auction stage, before winners are known, so losing auctions can carry substantial spillover and are exactly the cases the graph cannot see. Because METIS clusters are chosen to minimize observed edge cut, missing edges involving lost auctions are disproportionately likely to cross clusters, so the reported 36% between-cluster leakage is not a reliable estimate of true spillover—it may substantially understate it. The one supporting number, a prior product-split estimate of 44% lift versus 24% actual lift (Section 3), is non-concurrent and reported without confidence intervals, and the online DID result is only described as 'consistent' with post-rollout lift. Without a formal bias bound or an uncensored graph comparison, the graph could place real competitors in different clusters, and the cluster-DID estimator would inherit exactly the SUTVA violation the design is meant to remove.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses interference in search-advertising A/B tests when the unit of randomization is a product. It proposes constructing a bipartite graph between search queries and products from the advertiser's daily search query reports, weighting a one-mode projection by Eq. (1), partitioning the product graph with METIS into balanced clusters, and randomizing clusters to treatment and control for a difference-in-differences evaluation of a new bidding model. The authors report that the cluster-randomized online test produced a lift 'consistent with' the post-rollout lift, whereas an earlier product-split test overstated lift (~44% vs ~24%). They also describe a SageMaker-based system architecture for deploying the design. The paper is candid that the query-report data are censored to won auctions and that a formal meta-experiment was not run.","tokens_in":8303,"tokens_out":4330,"duration_ms":47905,"significance":"If the central claim were established, the paper would provide a practical and scalable solution to a real industrial problem, with a computationally efficient weighting function and a production architecture. The manuscript is honest about its main limitations (censored data, absence of a meta-experiment), and it reports a real deployment rather than a simulation. However, the evidence for the headline claim—that the cluster-randomized design 'gives a good estimate of the actual lift'—is thin: the supporting quantitative comparison is non-concurrent and lacks uncertainty quantification, the censoring problem is acknowledged but not mitigated, and the DID results are reported only qualitatively. The methodology is coherent enough to warrant revision, but not yet sufficient to support the causal claim as stated.","major_comments":[{"comment":"Section 6 explicitly concedes that 'such observational data is censored, that is, we only have data when we win the auction.' This concession applies to the entire construction in Section 3: the bipartite graph, the weights in Eq. (1), and the reported 36% leakage statistic are computed from impressions on won auctions only. Since interference for a bidding change operates at the auction stage before winners are determined, losing auctions are exactly the cases most likely to carry spillover, and missing edges are not missing at random. The paper provides no argument or bound showing that clusters built from censored data control true interference; the claim that the design 'gives a good estimate of the actual lift' is therefore not established. Please provide either a formal sensitivity analysis under a model of censoring, a comparison against an uncensored SERP data source (e.g., the seoClarity data mentioned in Section 6), or a synthetic-data study demonstrating that the censoring does not materially change the estimated spillover structure.","section":"Section 6 and Section 3, Eq. (1)"},{"comment":"The only quantitative support for bias reduction is the comparison between a prior product-split experiment (~44% lift) and the actual post-rollout lift (~24%). The authors themselves note in Section 3 that the proper validation is a meta-experiment randomizing over two designs (citing Saveski et al.), and they did not run it. Moreover, the 44% and 24% figures come from different time periods and no confidence intervals are reported. As a consequence, the claimed 'overstatement' is not statistically characterized, and this comparison cannot validate the cluster-DID estimator. Please report confidence intervals for both lifts, or replace the comparison with a formal bias bound connecting the leakage L to the bias of the cluster-randomized DID estimator.","section":"Section 3, 'Magnitude of Spillover' and Section 6"},{"comment":"The DID analysis is described only as showing 'a lift in click-through-rate for the treatment group which was consistent with the lift observed post roll out.' No point estimate, standard error, confidence interval, or p-value is reported. Without these, 'consistent' is not assessable; a sufficiently wide confidence interval could make even the 44% product-split estimate consistent with the rollout. Please report the cluster-robust DID estimate and its confidence interval, and specify what 'consistent' means in terms of a pre-specified equivalence margin or similar criterion.","section":"Section 4, 'Measurement'"},{"comment":"The choice of k=10,000 clusters and the associated 36% edge-weight leakage is selected from an elbow plot, but no sensitivity analysis is reported. Because both the number of clusters and the leakage tolerance are free design choices, and because the leakage is computed from the censored graph, the paper should show that the main conclusion is robust to reasonable variations in k and L. Without this, it is unclear whether the reported result is an artifact of a particular tuning choice rather than a property of the interference structure.","section":"Section 3, 'Optimal number of clusters'"}],"minor_comments":[{"comment":"Equation (1) is difficult to parse as printed: the summation appears as '𝑛∑︀𝑖=1' without the summed expression being clearly aligned, and the roles of n, 𝑊bi, 𝑓𝑠𝑞𝑖, and the indicator function are only given in the itemized list. Please rewrite the equation with standard notation and define all quantities before first use.","section":"Section 3, Eq. (1)"},{"comment":"The term 'network dismantling algorithms' is used to describe METIS graph partitioning, but these are distinct families of methods; please adjust the terminology to avoid confusion.","section":"Section 3, 'Graph Partitioning Methodology'"},{"comment":"There is a grammatical error: 'The serves as a blueprint' should read 'This serves as a blueprint.'","section":"Section 5"},{"comment":"There is a typo: 'we can than use' should be 'we can then use.'","section":"Section 6"},{"comment":"The abstract mentions a tripartite network, but the paper only develops the bipartite search-query/product network; please either remove the tripartite mention or clarify that it is planned future work.","section":"Abstract and Section 3"},{"comment":"The auction format is described as 'second price Vickrey–Clarke–Groves (VCG)', but VCG is not generally a second-price auction; please rephrase to avoid an inaccurate characterization.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style industry contribution with a real deployment, and the authors are commendably explicit about the main limitations: censored data and the absence of a meta-experiment. The core weakness is evidential rather than methodological: the claim of bias reduction rests on a non-concurrent 44%-vs-24% comparison without confidence intervals and on a qualitative 'consistent' DID result. I would encourage the editor to request a revision that adds uncertainty quantification and a censoring sensitivity analysis, even if the meta-experiment cannot be run. The system-architecture section is of secondary interest and could be shortened to make room for the missing statistical evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is using the ad publisher's own query-impression reports to build a SERP interference network, with a log-frequency-downweighted projection that scales to 200M+ products. That is a real contribution, and the paper is honest about its central limitation: the reports are censored to won auctions. The design itself is clearly described, and the choice of METIS over connected components and PIC is sensible given the imbalance-versus-leakage tradeoff. The system architecture on SageMaker is a useful blueprint, and the authors do not oversell what they have: they openly say no meta-experiment was run and that the censorship is a downside.\n\nThe soft spot is the validation. The abstract and conclusion claim the design \"gives a good estimate of the actual lift,\" but the supporting evidence is thin. The 44% versus 24% comparison in Section 3 is non-concurrent, has no confidence intervals, and comes from a different product-split experiment. The online DID result is only described as \"consistent\" with post-rollout lift, which is not the same as showing the estimate is unbiased. More importantly, the graph is built only from queries where the advertiser won the auction. Interference from losing auctions—where a competitor's bid affects whether your ad shows—is invisible, and those missing edges are exactly the ones that could carry substantial spillover. The reported 36% between-cluster leakage may therefore understate the true leakage, and since METIS clusters are chosen to minimize observed edge cut, missing edges are disproportionately likely to cross clusters. The paper does not address this, and it undermines the central claim. A formal bias bound or a concurrent comparison with an uncensored graph would be needed to support it.\n\nThat said, the paper is not a throwaway. As a design proposal with an illustrative case study, it is worth engaging with. The graph construction is novel and could be built on by other search marketing teams. The authors are clear thinkers who acknowledge the main weakness. The right fix is to soften the claims and, if possible, run the meta-experiment they mention or provide some sensitivity analysis around the censoring.\n\nSend it to peer review. It deserves referee time, but the editor should expect heavy revision on the validation section.","headline":"Practical cluster-randomized design for SERP interference built from the ad publisher's own query reports; the construction is new and credible, but the validation of the 'good estimate of actual lift' claim rests on one non-concurrent comparison and censored graph data.","tokens_in":8843,"tokens_out":1425,"would_cite":true,"duration_ms":17482,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cluster-randomized experiments built from SERP interference networks estimate a bidding algorithm's true lift, where product-split tests overstate it.","keywords":["SERP interference network","sponsored search","A/B testing","cluster randomized experiment","unipartite graph projection","bidding algorithm","censored observational data","paid search advertising"],"falsifier":"The clearest falsifier is a meta-experiment that randomizes between a Bernoulli product-split design and the cluster design on the same marketplace and compares the two estimated total average treatment effects; if they are statistically indistinguishable, the cluster design provides no bias correction over product splits.","tokens_in":7791,"feed_emoji":"📊","tokens_out":8504,"duration_ms":93714,"temperature":0.7,"pith_summary":"Search advertisers cannot randomize A/B tests on customers because search users are anonymous to them, and randomizing on products instead violates the no-interference assumption when two products in different groups appear in the same auction. This paper claims that clustering products that co-appear on search results pages, and randomizing whole clusters, restores a trustworthy comparison. The authors build the clusters from the advertiser's own search query reports, using a cheap weighted graph projection and a scalable graph partitioning algorithm. In their application, the cluster-randomized estimate of a new bidding algorithm's effect was consistent with the lift observed after full rollout, whereas an earlier product-split test had overstated the lift. If the design works as claimed, search marketing teams can run faster and more reliable A/B tests without geo-cloning campaigns.","feed_headline":"Cluster tests capture true bid lift where product splits overstate","feed_subtitle":"SERP co-appearance clusters matched post-rollout lift in online tests; product splits showed inflated gains.","key_machinery":"The central object is the SERP interference network: an undirected bipartite graph between search queries and products, built from the ad publisher's query reports over a year of impressions. The paper's computational device is a weighted one-mode projection with edge weight\n$$W_{\\rm uni}(a,b)=\\sum_i \\frac{1}{\\log_e(f_{sq_i})}\\frac{\\min(W_{\\rm bi}(sq_i,a),W_{\\rm bi}(sq_i,b))}{\\max(W_{\\rm bi}(sq_i,a),W_{\\rm bi}(sq_i,b))}\\mathbb{I}[sq_i,a,b],$$\nwhere $f_{sq_i}$ is the number of distinct products query $i$ drives impressions to and the indicator is one exactly when the query triggered impressions for both products. This weight suppresses contributions from broad upper-funnel queries while scoring products that repeatedly co-appear on the same results page as strongly interfering. The projection is partitioned with a multilevel graph partitioning algorithm that runs in $O(|E|)$ time and balances cluster sizes while minimizing the weighted edge cut between clusters, and the resulting clusters become the randomization units.","core_discovery":"The paper's central claim is that a product-cluster randomized control trial formed from a SERP interference network yields a near-unbiased estimate of a bidding algorithm's causal effect. The interference network is a bipartite graph whose edges connect search queries to the products that the ad publisher showed for them; a weighted one-mode projection places an edge between any two products that co-appear for a query, with weight down-weighted for generic queries. Partitioning that product graph into balanced clusters (10,000 clusters in the reported run) and randomizing clusters, rather than individual products, makes SUTVA approximately hold because spillover is concentrated inside clusters and both treatment and control bids are absent in the same auction. The reported online A/B test, analyzed with difference-in-differences and cluster-robust standard errors, showed treatment-group lift consistent with the lift seen after the new bidding model was rolled out, while a previous naive product-split test showed inflated lift of roughly 44% versus actual approximately 24%.","pith_inferences":["The most direct test of the design would be the meta-experiment the paper mentions but could not run: randomizing one arm by Bernoulli product split and another by cluster split, then testing whether the two estimated treatment effects differ significantly.","Because the query reports only record auctions the advertiser won, the true interference graph likely has more edges than the projected one; adding SERP-level observation data would probably raise the measured between-cluster leakage and could change cluster boundaries.","The log-frequency down-weighting in the projection is one defensible choice; weighting by query position, ad slot, or auction presence might sharpen the interference model, but that extension would need its own validation.","If clusters can be refreshed as query reports stream in, the approach would generalize from one-off experiments to continuous experimentation, conditional on the cluster-stability analysis the paper lists as future work."],"forward_implications":["Search marketers can run online A/B tests on bidding algorithms without cloning whole ad accounts or waiting out long switchback adjustment periods, because randomization happens on product clusters rather than individual products.","The earlier naive product-split experiment overestimated the bidding model's impact (roughly 44% lift versus roughly 24% realized), so the cluster design is what brought the online estimate in line with post-rollout performance.","The leakage-versus-cluster-count curve gives an operational rule for choosing the number of clusters: pick the largest $k$ before between-cluster edge weight climbs, balancing statistical power against bias.","Cluster-robust standard errors, applied after difference-in-differences estimation, prevent within-cluster error correlation from producing misleadingly small p-values."],"supporting_citations":[{"why":"Provides the graph cluster randomization concept that motivates randomizing clusters of products instead of individual units.","marker":"[9]"},{"why":"Supplies the auction-experiment precedent of randomizing keyword clusters to avoid interference between auctions.","marker":"[11]"},{"why":"Presents the weighted one-mode projection method whose computational cost this paper avoids with its cheaper weighting function.","marker":"[20]"},{"why":"Supplies the multilevel graph partitioning algorithm used to create balanced product clusters that minimize cross-cluster edge weight.","marker":"[26]"},{"why":"Describes the meta-experiment over randomized designs that would empirically quantify the bias reduction from clustering.","marker":"[27]"},{"why":"Supplies the difference-in-differences estimator used to measure the treatment effect in the online experiment.","marker":"[28]"},{"why":"Supplies the cluster-robust standard error approach used after estimation to account for within-cluster correlation.","marker":"[29]"}],"fun_headline_variants":["Cluster-based ad tests fix interference, reveal true bid lift","SERP co-appearance clusters beat product splits for ad lift","True ad lift from cluster randomization, not product splits","Clustering on SERP co-appearance yields unbiased bid lift","Randomize on SERP clusters for reliable search ad lift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design assumes the advertiser's search query reports faithfully represent the interference structure, even though they only show auctions the advertiser won; if losing auctions carry substantial spillover, the projected product graph and clusters are biased.","fun_headline_variants_meta":{"raw":{"variants":["Cluster-based ad tests fix interference, reveal true bid lift","SERP co-appearance clusters beat product splits for ad lift","True ad lift from cluster randomization, not product splits","Clustering on SERP co-appearance yields unbiased bid lift","Randomize on SERP clusters for reliable search ad lift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001004,"raw_usage":{"total_tokens":4242,"prompt_tokens":937,"completion_tokens":3305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":3223}},"tokens_in":553,"tokens_out":3305,"duration_ms":27066,"temperature":1.0,"reasoning_tokens":3223,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:44:32.632739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The clearest falsifier is a meta-experiment that randomizes between a Bernoulli product-split design and the cluster design on the same marketplace and compares the two estimated total average treatment effects; if they are statistically indistinguishable, the cluster design provides no bias correction over product splits.","supporting_citations":[{"cited_title":"Ugander, B","cited_arxiv_id":null,"evidence_quote":"Provides the graph cluster randomization concept that motivates randomizing clusters of products instead of individual units."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the auction-experiment precedent of randomizing keyword clusters to avoid interference between auctions."},{"cited_title":"Stram, P","cited_arxiv_id":null,"evidence_quote":"Presents the weighted one-mode projection method whose computational cost this paper avoids with its cheaper weighting function."},{"cited_title":"Karypis, V","cited_arxiv_id":null,"evidence_quote":"Supplies the multilevel graph partitioning algorithm used to create balanced product clusters that minimize cross-cluster edge weight."},{"cited_title":"Saveski, J","cited_arxiv_id":null,"evidence_quote":"Describes the meta-experiment over randomized designs that would empirically quantify the bias reduction from clustering."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the difference-in-differences estimator used to measure the treatment effect in the online experiment."},{"cited_title":"White, Asymptotic Theory for Econometricians, Academic Press, 2014","cited_arxiv_id":null,"evidence_quote":"Supplies the cluster-robust standard error approach used after estimation to account for within-cluster correlation."}],"review_version":1}