{"id":"952978c5-6c2b-4031-9968-9167405d567d","arxiv_id":"2605.09726","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":8.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"No uniformly consistent specification test exists for exposure mapping models of interference, as worst-case Type I and Type II error rates must sum to one for any test.","lead":"This paper proves that specification tests for causal interference models based on exposure mappings cannot achieve both low false positive and low false negative rates in the worst case. Any such test has worst-case errors summing to one, making uniformly consistent testing impossible without further restrictions on alternatives.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flags that the result hinges on the separation condition and fixed mappings; the abstract supplies no counter-example or loophole that would make the impossibility fail to hold under those conditions. The explicit construction of a consistent test once the alternative is narrowed further corroborates rather than undermines the general negative result.","tokens_in":1755,"tokens_out":257,"duration_ms":15538,"concrete_test":"Re-derive the minimax lower bound (presumably via a least-favorable pair of distributions) for the specific no-interference vs. network-linear-in-means case; confirm that the lower bound equals 1 only when the alternative class is left unrestricted and drops below 1 once the linear-in-means restriction is imposed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states a minimax result: for any test, worst-case Type I + Type II error equals 1 over maximally separated alternatives (with bounded outcomes and fixed exposure mappings), for every sample size and every exposure mapping. This is consistent with the provided positive result under a restricted alternative class. No internal contradiction, hidden assumption, or derivation gap is visible from the given material that would invalidate the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that specification testing for interference models based on exposure mappings is impossible in a minimax sense: for any test, the worst-case Type I error rate plus the worst-case Type II error rate equals 1 for all exposure mappings, all sample sizes, uniformly bounded outcomes, and alternatives maximally separated from the null. This rules out uniformly consistent tests and is achieved by a naive test that discards data and rejects at random. The authors also construct a uniformly consistent test under the restricted alternative of distinguishing no-interference from a network-linear-in-means model.","tokens_in":1820,"tokens_out":335,"duration_ms":40330,"significance":"If the result holds, it is significant because it provides a decision-theoretic explanation for the poor power of existing tests and shows that informative specification testing requires additional restrictions on the alternative beyond the exposure mapping itself. The generality across all exposure mappings and sample sizes, together with the sharpness of the bound (achieved by a data-ignoring procedure), strengthens the negative result. The positive illustration under a restricted alternative is a constructive strength that offers practical guidance for when consistent testing is feasible.","major_comments":[],"minor_comments":[{"comment":"The abstract would benefit from a brief parenthetical clarification of how 'maximally separated' is formalized (e.g., via total variation or outcome distribution distance) to help readers immediately gauge the scope of the impossibility.","section":"Abstract"},{"comment":"Consider adding one sentence in the introduction contrasting the general impossibility with the restricted positive result to improve readability for readers focused on applications.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary of the manuscript, recognition of its significance, and recommendation for minor revision. No specific major comments were provided in the report, so we have no points requiring point-by-point rebuttal or revision at this stage.","responses":[],"tokens_in":1293,"tokens_out":69,"duration_ms":11221,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that specification testing for exposure mapping models is impossible in a strong sense: for any test, the worst-case Type I and Type II errors sum to one across all sample sizes and all exposure mappings when outcomes are bounded and alternatives are maximally separated from the null. This matches the error rate of just ignoring the data and flipping a coin, so no uniformly consistent test exists without further restrictions.\n\nWhat is new is the clean minimax argument establishing this bound for the entire class of exposure mappings. The paper also supplies a positive counterpart by constructing a uniformly consistent test that separates a no-interference null from a network linear-in-means alternative. That example is useful because it demonstrates exactly where the impossibility can be escaped.\n\nThe work is clear on why existing tests have low power and why that is not an accident of poor design. The decision-theoretic framing keeps the argument direct and avoids extra assumptions.\n\nThe main soft spot is the reliance on maximally separated alternatives. If actual model violations in data are not that far from the null, the bound may not constrain practical tests as tightly, and some power could remain. The result is also stated for fixed exposure mappings, which leaves open whether data-dependent mappings change the picture. These conditions are stated explicitly, so they are not hidden, but they do limit how broadly the impossibility applies.\n\nThis paper is for researchers working on causal inference under interference, particularly those designing or evaluating specification tests in network experiments. A reader who needs to know the fundamental limits on testing exposure mappings will find it directly useful. It deserves serious peer review because the central negative result is sharp, the positive illustration balances it, and the topic matters for applied work in the area. I would send it out.","headline":"The paper proves no test of exposure mapping models can beat random guessing in the worst case over separated alternatives, but shows consistency is possible once alternatives are narrowed.","tokens_in":2269,"tokens_out":426,"would_cite":true,"duration_ms":18848,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Specification tests for exposure mapping models of interference cannot be uniformly consistent.","keywords":["specification testing","exposure mappings","interference","causal inference","hypothesis testing","uniform consistency","error rates"],"falsifier":"Existence of a specification test where the sum of worst-case Type I error and worst-case Type II error is strictly less than one under the stated conditions of bounded outcomes and maximally separated alternatives.","tokens_in":2669,"feed_emoji":"🚫","tokens_out":539,"duration_ms":24256,"temperature":0.7,"pith_summary":"The paper proves that any test of an interference model based on exposure mappings has worst-case Type I and Type II error rates that sum to one. This bound is the same as that of a test which ignores the data entirely and rejects the null at random. The result holds for every exposure mapping, every sample size, bounded outcomes, and alternatives that are as far from the null as possible. A sympathetic reader would conclude that useful specification tests must impose further restrictions on the possible alternatives beyond those in the exposure mapping itself.","feed_headline":"Exposure mapping model tests cannot be consistent","feed_subtitle":"Any specification test has worst-case Type I and Type II errors summing to one for all exposure mappings and sample sizes.","key_machinery":"The worst-case error rate sum of one for specification tests under exposure mappings, which prevents uniform consistency.","core_discovery":"The central claim is that the worst-case Type I and Type II error rates must sum to one for any specification test of exposure mapping models. This rules out the existence of a uniformly consistent test. The result applies to all exposure mappings, all sample sizes, uniformly bounded outcomes, and alternatives maximally separated from the null.","pith_inferences":["Researchers need to specify concrete alternative models to make testing informative.","This impossibility may suggest similar limits in other causal models defined by summary statistics or mappings.","Future work could explore the minimal restrictions needed on alternatives to allow consistent testing."],"forward_implications":["Existing specification tests suffer from poor power against some model violations.","Tests can only attain power by restricting the set of alternatives considered.","The paper provides an example of a consistent test when restricting to no-interference versus a network linear-in-means model.","Any test will leave some relevant departures from the model undetectable."],"fun_headline_variants":["Exposure mapping tests impossible","No consistent test for exposure mappings","Specification testing fails for exposure models","Exposure model tests cannot be consistent"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Alternatives to the null model are maximally separated from it.","fun_headline_variants_meta":{"raw":{"variants":["Exposure mapping tests impossible","No consistent test for exposure mappings","Specification testing fails for exposure models","Exposure model tests cannot be consistent"]},"model":"grok-4.3","cost_usd":0.003327,"raw_usage":{"total_tokens":1780,"prompt_tokens":683,"num_sources_used":0,"completion_tokens":43,"cost_in_usd_ticks":33274500,"prompt_tokens_details":{"text_tokens":683,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1054,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":683,"tokens_out":43,"duration_ms":8223,"temperature":1.0,"reasoning_tokens":1054,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T22:32:28.134818+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Existence of a specification test where the sum of worst-case Type I error and worst-case Type II error is strictly less than one under the stated conditions of bounded outcomes and maximally separated alternatives.","supporting_citations":[],"review_version":2}