{"id":"ab8eeeea-270a-446a-81de-c0ffbac17cef","arxiv_id":"2507.03438","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Model independence is defined by two conditions, no target model and a well-defined background, and deep-learning anomaly detection can meet them on a spectrum.","lead":"This philosophy-of-physics paper proposes a precise definition of model independence as a matter of degree, requiring no target model and a well-defined background. It argues that deep-learning anomaly searches in high-energy physics can satisfy these conditions and thus broaden the search for new physics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own Section 4 taxonomy is undercut by its Section 3.3.1: AEN pipelines are optimized on BSM benchmark signals, so by the paper's contrast with \"model agnostic\" they have many target models rather than none; the target/benchmark distinction needs an operational criterion.","rationale":"The reader's weakest assumption concerned generalization to untested BSM signals. That is a real and empirically observable risk, and the paper honestly acknowledges it. But the more load-bearing concern is definitional and comes earlier in the argument: the paper's own Section 3.3.1 item (v) shows that AEN pipelines are optimized and thresholded using specifically chosen BSM benchmark scenarios. The paper's Section 4 then tries to classify these as compatible with \"no target model,\" without an operational rule for when a signal used in the pipeline counts as a target. By the paper's own contrast with simplified models and model agnosticism, a benchmark set of BSM scenarios looks exactly like \"many target models,\" not \"no target model.\" This is not an ad hominem or a demand for impossible standards; it is an internal tension between the taxonomy and the described practice. The proposed two-condition definition could be made precise by defining target models in terms of the role a signal plays in the final anomaly decision, but that step is missing. Because the philosophical contribution is otherwise careful and valuable, the right response is not rejection but a conditional acceptance: the central classification claim should be revised or explicitly qualified until the target/benchmark distinction is given a testable criterion.","tokens_in":17225,"tokens_out":3962,"duration_ms":49550,"concrete_test":"Use a public anomaly-detection benchmark, e.g., the Dark Machines or LHC Olympics datasets, and train the same autoencoder twice with identical background but different BSM benchmark sets used for hyperparameter and threshold selection: one benchmark set containing a single-resonance signal and one containing a two-decay signal. Then apply both trained detectors to a held-out set containing both signal types. If the held-out signal rankings change with the chosen benchmark set, the benchmark set is functioning as a target model, and condition 1 of the paper's definition is not satisfied in practice. If the rankings are stable, the concern would be substantially weakened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that autoencoder-based anomaly searches satisfy condition 1 (\"has no target model\") depends entirely on the paper's distinction between a target model and a BSM benchmark. The paper uses this distinction to place AENs beyond model agnosticism: model-agnostic searches, e.g., simplified models, have many target models, while AENs have none. But Section 3.3.1(v) concedes that the network and its loss function are optimized by testing on BSM scenarios, and Section 3.3 says a threshold is set \"that captures many BSM signals but does not capture SM processes or noise.\" Thus the benchmark set is not a peripheral validation step; it constrains architecture, latent-space dimension, loss function, and anomaly threshold, all of which are constitutive parts of the search method. If a benchmark signal is a target model, then AENs have multiple target models and fall into the paper's own model-agnostic category. If a benchmark signal is not a target model, the paper gives no criterion for distinguishing the two, so any model-agnostic search could relabel its target as a benchmark and cross the threshold; the definition becomes unfalsifiable. The empirical record the paper itself cites points the same way: in LHC Olympics Black Box 3, unsupervised methods failed on a two-decay signal, and Dark Machines reported signal-dependent performance across BSM scenarios. That is exactly what would happen if the benchmark set were doing target-model work. The definition may be salvageable, but the paper does not state what would count as a target model versus a benchmark in a way that can be checked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that 'model independence' is a meaningful, graded concept in high energy physics, rather than an empty ideal or a synonym for 'model agnosticism.' It proposes two necessary conditions for a search method to count as model-independent: (1) it has no target model of new physics, and (2) it has a well-defined background model against which deviations are defined. The paper places deep-learning anomaly detection, especially autoencoder networks, at the model-independent end of this spectrum, claiming that these methods go beyond model-agnostic approaches such as simplified-model searches. It reviews relevant DL technology and evidence from the LHC Olympics and Dark Machines challenges, discusses epistemic issues including architecture choices, simulation dependence, and interpretability, and concludes that the compromises are not fatal to the aim of model independence.","tokens_in":17511,"tokens_out":5613,"duration_ms":64657,"significance":"The paper is a valuable conceptual intervention in the philosophy of HEP methodology. It offers a precise, falsifiable proposal: model independence is defined by two conditions and is situated on a spectrum, which could clarify ongoing terminology disputes and give a principled basis for claims that current LHC anomaly searches are model-independent. The author is unusually candid about contrary evidence, including the failure of many LHC Olympics methods on Black Box 3 and the signal-dependent performance reported by Dark Machines. The paper contains no circular derivations, fits no parameters, and its central definition does not depend on the author's other work. Its main significance lies in the target/benchmark distinction, which, if made operational, would give a useful threshold for classifying searches; as it stands, that distinction is the load-bearing point that needs further work.","major_comments":[{"comment":"The central distinction between a 'target model' and a BSM 'benchmark' is never operationalized. Section 3.3 (fourth step) states that the anomaly threshold is set so that it 'captures many BSM signals but does not capture SM processes or noise,' and §3.3.1(v) concedes that '[o]ne optimizes a network and the loss function by testing on given signal identification tasks.' These benchmarks constrain the architecture, latent-space dimension, loss function, and threshold, all of which are constitutive of the search method. Hence either a benchmark signal is a target model, in which case the AEN has many target models and falls into the paper's own 'model-agnostic' category, or it is not, in which case the paper needs a criterion that prevents any model-agnostic search from relabelling its targets as benchmarks and thereby crossing the threshold. Without such a criterion, condition 1 ('has no target model') is unfalsifiable. The author notices the worry in §3.3.1(v) ('one could harbor reservations that the AEN is more model agnostic than model-independent'), but §4 does not resolve it. Section 4 needs either an operational distinction between a target model and a benchmark, or a weakened condition such as 'no privileged target model,' with the gradability of model independence then doing the work.","section":"§4, cf. §3.3 and §3.3.1(v)"},{"comment":"The practical conclusion that AENs can flag genuinely unexpected new-physics signals rests on an extrapolation from tested BSM benchmarks, but the paper's own cited evidence points in the opposite direction for at least some cases. LHC Olympics Black Box 3 was not correctly predicted by the submitted methods before unblinding (Kasieczka et al., 2021), and Dark Machines found that even the best algorithms had essentially no improvement for some injected signals (Aarrestad et al., 2022). Section 4 asserts that 'it has been demonstrated in various anomaly detection studies and competitions that a network tested to have high anomaly scores on various BSM scenarios, is able to flag further kinds of new signals,' but it supplies no citation or quantitative detail for that claim; the two competitions described in §3.3 show signal-dependent performance rather than a general-capability result. To make this load-bearing claim, the author should either cite the specific demonstrations or reformulate the conclusion as a conjecture about DL generalization, explicitly treating the LHC Olympics and Dark Machines results as evidence about the limits of that generalization.","section":"§4, §3.3"}],"minor_comments":[{"comment":"The phrase 'flag further kinds of new signals as anonymous' should read 'as anomalous.'","section":"§4"},{"comment":"There are several typographical and formatting errors: 'Novemeber' should be 'November'; the Morrison reference lists 'Cumbridge University Press' instead of 'Cambridge University Press'; and the Baldi reference is malformed ('Baldi, P., S. P. . W. D. (2014)').","section":"Front matter, References"},{"comment":"The wording 'A search method is model-independent if it: 1. has no target model; 2. has a well-defined background' is inconsistent with calling these 'necessary conditions'; necessary conditions should be expressed with 'only if,' and if the intent is to give a definition, the text should say 'if and only if.'","section":"§2.2"},{"comment":"The figure captions omit the information needed to read the plots: Figure 2 does not identify which curve corresponds to which network, and Figure 3 does not name the axes or the classifiers being compared.","section":"Figures 2 and 3"},{"comment":"Several references are incomplete or fragmentary, for example Plehn et al. (2022) lacks a journal or arXiv identifier, and the reference to Kukačka et al. (2017) gives no publication venue.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is clearly within scope for physics.hist-ph and makes a genuinely useful contribution to the philosophy of HEP methodology. The concern raised here is not about missing evidence or novelty; it is that the §4 taxonomy needs an operational target/benchmark criterion to be stable. If the author can supply that criterion, or explicitly weaken the 'no target model' condition to a graded claim, the paper would be a solid accept; without it, the central distinction risks collapsing into relabelling."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read. The interesting move is the two-condition definition: no target model plus a well-defined background, and the explicit claim that autoencoder anomaly searches clear the bar while simplified-model searches do not. That is a genuine refinement of McCoy and Massimi and Dillon et al. The literature review is accurate and unusually candid: it reports LHC Olympics failures, Dark Machines signal-dependence, and the fact that unsupervised methods missed a two-decay signal. Credit where it is due.\n\nThe soft spot is exactly the one flagged by the stress-test: the target-model/benchmark distinction is doing heavy lifting and no operational criterion is given. The method sets architecture, latent dimension, and threshold using BSM benchmarks; Section 3.3.1(v) admits this. If 'target model' means anything that shapes the search statistic, then autoencoder networks have many target models and fall into the paper's own model-agnostic bucket. If it does not, the paper needs to say what it does mean. The final generalization caveat is honest, but it does not solve the boundary problem, because the threshold is part of the method and benchmark signals are not merely after-the-fact validation. This is not fatal, but it is the central conceptual gap. A revised paper could define target model as a signal template used at inference time, or require that the benchmark set be diverse and not tuned to maximize sensitivity to any one scenario. The paper just does not do that work.\n\nMinor point: the title says Deep Learning and Model Independence, but the two-condition definition applies just as well to precision measurements and SMEFT. That is not a flaw, just a broader frame than advertised. The self-citation to Bechtle et al. is background-only; no circularity problem.\n\nVerdict: solid philosophy-of-physics contribution, honest about evidence, but the definition needs a sharper target/benchmark criterion before it lands as more than a useful schema. Send to a referee with expertise in HEP anomaly detection and ask them to press on that distinction. It deserves review, not desk rejection. I would bring it to reading group; I do not expect to cite it in my own work, but if I were writing on model independence it would be in the conversation.","headline":"A genuinely useful two-condition definition of model independence, but the target/benchmark line needs sharpening before the deep-learning claim lands.","tokens_in":18043,"tokens_out":2649,"would_cite":false,"duration_ms":32565,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Autoencoder anomaly searches at the LHC can count as genuinely model independent.","keywords":["model independence","deep learning","autoencoder","anomaly detection","high energy physics","beyond the Standard Model","unsupervised learning","philosophy of science"],"falsifier":"Run a blinded anomaly search on a collider dataset in which a non-resonant or multi-pronged new-physics signal has been injected, using an autoencoder that has strong benchmark scores on several resonant BSM scenarios; if the network's anomaly scores show no significant excess where the injected signal is known to be, the paper's generalization assumption is refuted in that regime.","tokens_in":17009,"feed_emoji":"🔬","tokens_out":5034,"duration_ms":55677,"temperature":0.7,"pith_summary":"High-energy physics has spent decades hunting for new particles with searches tuned to specific beyond-Standard-Model theories, and none has paid off. This paper argues that the field's turn to 'model-independent' methods is real and definable: a search counts as model independent when it has no target model and has a well-defined background against which deviations register. On that definition, deep-learning anomaly detection with autoencoders — networks trained only to reconstruct Standard Model events — qualifies as genuinely model independent, not merely model agnostic. The stakes are methodological: if the definition holds, independence from models is a graded property with a clear threshold, and current LHC anomaly searches occupy the model-independent end of the spectrum.","feed_headline":"Deep-learning anomaly searches can claim real model independence","feed_subtitle":"The paper defines model independence as no target model plus a well-defined background, then shows autoencoders meet the bar.","key_machinery":"The load-bearing object is the distinction between target model and background model, coupled with the two-condition definition of model independence built on it. The autoencoder, and its variational variant, is the concrete mechanism: it learns to reconstruct Standard Model events in a compressed latent space, then scores new events by reconstruction error or likelihood, so any large deviation from the learned background flags an anomaly without ever specifying a BSM signal. Target models are demoted to benchmarking tools, used after training to check that known BSM scenarios produce high anomaly scores. The definition separates the search itself, which involves only a background model, from the validation of the search, which involves benchmark models, and that separation carries the argument.","core_discovery":"The paper's central claim is a definitional thesis with an empirical application. It proposes that model independence is not all-or-nothing but a spectrum, and that a search method is model-independent exactly when it has no target model and has a well-defined background against which deviations can be seen. Autoencoder-based anomaly detection satisfies both conditions: the network is trained on Standard Model background, and there is no BSM signal hypothesis in the search itself. The author argues this is stronger than model agnosticism, which still assumes a set of target models; DL earns the label further because the network constructs its own features, so it may flag patterns physicists have not anticipated. Target-model assumptions reappear only in benchmarking, optimization, and post-hoc interpretation, which the author treats as legitimate external roles rather than as part of the search's model dependence.","pith_inferences":["Read as a proposed definition rather than a report, the paper invites a quantitative refinement: a 'model-independence breadth' measure based on the diversity of benchmark signals an anomaly detector flags, which would let physicists compare networks across the spectrum.","If the two-condition definition is adopted, the live scientific question shifts from whether these searches are model independent to whether they are reliable enough; an autoencoder that meets the definition could still be blind to a large class of non-resonant new physics, and the paper's own LHC Olympics evidence shows this risk is real.","The target/background distinction could generalize beyond collider physics: any data-driven search that defines a well-characterized null distribution and refuses to specify alternatives, such as outlier screens in astronomy or genomics, would inherit the same model-independence status."],"forward_implications":["If the two conditions are accepted, precision measurements, SMEFT parameter scans, and deep-learning anomaly searches form one family, while simplified-model searches count only as model agnostic.","Autoencoder searches at the LHC can be described as model independent today rather than as a promissory note, which helps clarify the growing terminology of 'model independent' versus 'model agnostic'.","Because independence is graded, future anomaly detectors can be compared by how many and how varied their benchmarking scenarios are, without losing the label at the threshold.","Deep learning contributes to the search for new physics not only through improved performance and computational efficiency but by reducing target-model bias, giving it a distinct methodological advantage."],"supporting_citations":[{"why":"Supplies the LHC Olympics results showing unsupervised methods can flag blinded resonant anomalies but that none found anomalies from raw four-vectors without model-inspired dimensionality reduction, forming central evidence for both capability and limits.","marker":"Kasieczka et al., 2021"},{"why":"Provides the Dark Machines benchmark competition for unsupervised searches, showing that most algorithms improve significance on several BSM scenarios but none works for all, supporting the gradable view of model independence.","marker":"Aarrestad et al., 2022"},{"why":"Demonstrates that an autoencoder can improve signal-to-background ratio without a target signal, serving as the core example of deep-learning anomaly detection with no target model.","marker":"Farina et al., 2020"},{"why":"Supplies the framing that machine-learning searches can be 'liberated from model dependence compared with traditional searches' and describes the hypervariate analysis of full phase space.","marker":"Karagiorgi et al., 2021"},{"why":"Cited as a representative use of 'model agnostic' for target-less methods, providing the contrast that motivates the paper's terminological schema.","marker":"Dillon et al., 2023"},{"why":"Gives the philosophical account of simplified models as a model-agnostic approach, anchoring the middle of the model-dependence spectrum.","marker":"McCoy and Massimi, 2018"},{"why":"Offers the SMEFT case study of bottom-up parametrization, placing precision-measurement-style searches at the model-independent end alongside deep learning.","marker":"Bechtle et al., 2022"},{"why":"Differentiates the two basic anomaly-search methods and compares loss metrics for autoencoders, supplying the technical framework for anomaly scoring.","marker":"Fraser et al., 2022"}],"fun_headline_variants":["Autoencoders earn true model independence","Model independence is a degree, not a binary","No target, defined background: the key to model independence","Deep learning meets the bar for model independence","Autoencoders: model independence without a target model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The practical payoff rests on the belief that a network trained to flag known new-physics scenarios will also flag genuinely new and unexpected ones; if this generalization fails, deep-learning anomaly searches would be model independent by definition but not in practice, and the paper's main benefit for physics would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Autoencoders earn true model independence","Model independence is a degree, not a binary","No target, defined background: the key to model independence","Deep learning meets the bar for model independence","Autoencoders: model independence without a target model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00074,"raw_usage":{"total_tokens":3244,"prompt_tokens":823,"completion_tokens":2421,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":2350}},"tokens_in":439,"tokens_out":2421,"duration_ms":19713,"temperature":1.0,"reasoning_tokens":2350,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:09:21.786920+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a blinded anomaly search on a collider dataset in which a non-resonant or multi-pronged new-physics signal has been injected, using an autoencoder that has strong benchmark scores on several resonant BSM scenarios; if the network's anomaly scores show no significant excess where the injected signal is known to be, the paper's generalization assumption is refuted in that regime.","supporting_citations":[{"cited_title":"H., Dai, B., De Freitas, F","cited_arxiv_id":null,"evidence_quote":"Supplies the LHC Olympics results showing unsupervised methods can flag blinded resonant anomalies but that none found anomalies from raw four-vectors without model-inspired dimensionality reduction, forming central evidence for both capability and limits."},{"cited_title":"D., Doglioni, C., Duarte, J","cited_arxiv_id":null,"evidence_quote":"Provides the Dark Machines benchmark competition for unsupervised searches, showing that most algorithms improve significance on several BSM scenarios but none works for all, supporting the gradable view of model independence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that an autoencoder can improve signal-to-background ratio without a target signal, serving as the core example of deep-learning anomaly detection with no target model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the framing that machine-learning searches can be 'liberated from model dependence compared with traditional searches' and describes the hypervariate analysis of full phase space."},{"cited_title":"M., Favaro, L., Feiden, F., Modak, T., and Plehn, T","cited_arxiv_id":null,"evidence_quote":"Cited as a representative use of 'model agnostic' for target-less methods, providing the contrast that motivates the paper's terminological schema."},{"cited_title":"and Massimi, M","cited_arxiv_id":null,"evidence_quote":"Gives the philosophical account of simplified models as a model-agnostic approach, anchoring the middle of the model-dependence spectrum."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers the SMEFT case study of bottom-up parametrization, placing precision-measurement-style searches at the model-independent end alongside deep learning."},{"cited_title":"K., Ostdiek, B., and Schwartz, M","cited_arxiv_id":null,"evidence_quote":"Differentiates the two basic anomaly-search methods and compares loss metrics for autoencoders, supplying the technical framework for anomaly scoring."}],"review_version":1}